MixMAE: Mixed and Masked Autoencoder for Efficient Pretraining of Hierarchical Vision Transformers

Jihao LiuXin HuangJinliang ZhengYu LiuHongsheng Li

article2023CVPR95 citations

Proposes a dual-reconstruction pretraining framework that replaces masked tokens in one image with visible patches from another, eliminating uninformative mask tokens and improving efficiency across hierarchical Vision Transformers.

Listen

Training computer vision models on massive unlabelled image datasets allows systems to learn general visual representations, but standard masked image modelling methods face severe computational inefficiencies. Traditional approaches replace large portions of an image with artificial mask placeholder tokens, which wastes substantial computing power on uninformative tokens and introduces a mismatch during downstream fine-tuning. While standard vision transformers can bypass these tokens by processing incomplete patch sequences, hierarchical multi-scale architectures—which are vital for dense visual tasks like object detection and segmentation—require structured two-dimensional inputs and cannot use that approach.

The article introduces and evaluates Mixed and Masked Autoencoder (MixMAE), an efficient self-supervised pretraining framework designed specifically for hierarchical vision transformers. The core objective is to demonstrate that blending patches from multiple images and performing dual reconstruction removes the need for mask tokens in the encoder, thereby accelerating training speed and improving transfer performance across downstream vision tasks.

The researchers implemented MixMAE by combining visible image patches from different images into a single mixed input. A hierarchical encoder (primarily evaluated with Swin Transformer variants) processes this mixed image using masked attention to prevent cross-image contamination. The resulting representations are then separated, appended with mask tokens, and processed by a lightweight decoder tasked with reconstructing the original uncorrupted images. The method was evaluated through extensive self-supervised pretraining on ImageNet-1K, followed by fine-tuning across seven standard vision benchmarks covering image classification, object detection, instance segmentation, and semantic segmentation.

The evaluation produced several key findings: First, MixMAE pretraining on a base Swin Transformer achieved 85.1% top-1 accuracy on ImageNet-1K with 600 epochs of pretraining, outperforming competing masked image modelling frameworks while requiring 25% to 62% fewer pretraining epochs. Second, the architecture scales effectively up to 600 million parameters, achieving 86.9% accuracy on the huge model configuration while cutting required inference compute by 30% to 40% compared to standard vision transformer baselines. Third, on downstream dense prediction benchmarks, MixMAE established superior performance tradeoffs: on COCO object detection it reached 52.7 box average precision (exceeding standard baselines by 2.9 points), and on ADE20K semantic segmentation it reached 51.1 mean Intersection over Union using only half the inference computation of prior architectures. Finally, ablation studies confirmed that using mixed real image tokens significantly outperformed alternative strategies such as filling masked regions with zeros, learnable vectors, or shuffled patches.

These findings indicate that organisations deploying high-capacity visual perception systems can achieve superior accuracy with substantially lower training and operational compute costs. The compatibility with hierarchical architectures allows teams to deploy a single unified pretraining framework directly to complex spatial reasoning tasks, reducing the engineering overhead of redesigning architectures for multi-scale dense prediction.

Organisations investing in large-scale computer vision models should adopt mixed-token pretraining strategies over traditional mask-token approaches when training hierarchical backbones. The article recommends scaling to 75% masking ratios by mixing four images to maximise pretraining efficiency. Before extending this method to production, technical teams should conduct pilot studies on their domain-specific visual distributions. While the empirical results are robust across several public benchmarks, the article's experiments remain limited to standard natural image datasets, and further validation is required to assess whether the dual-reconstruction benefits translate directly to other modalities such as video, audio, or text.

Cover for MixMAE: Mixed and Masked Autoencoder for Efficient Pretraining of Hierarchical Vision Transformers

Abstract

In this paper, we propose Mixed and Masked AutoEncoder (MixMAE), a simple but efficient pretraining method that is applicable to various hierarchical Vision Transformers. Existing masked image modeling (MIM) methods for hierarchical Vision Transformers replace a random subset of input tokens with a special [MASK] symbol and aim at reconstructing original image tokens from the corrupted image. However, we find that using the [MASK] symbol greatly slows down the training and causes pretraining-finetuning inconsistency, due to the large masking ratio (e.g., 60% in SimMIM). On the other hand, MAE does not introduce [MASK] tokens at its encoder at all but is not applicable for hierarchical Vision Transformers. To solve the issue and accelerate the pretraining of hierarchical models, we replace the masked tokens of one image with visible tokens of another image, i.e., creating a mixed image. We then conduct dual reconstruction to reconstruct the two original images from the mixed input, which significantly improves efficiency. While MixMAE can be applied to various hierarchical Transformers, this paper explores using Swin Transformer with a large window size and scales up to huge model size (to reach 600M parameters). Empirical results demonstrate that MixMAE can learn high-quality visual representations efficiently. Notably, MixMAE with Swin-B/W14 achieves 85.1% top-1 accuracy on ImageNet-1K by pretraining for 600 epochs. Besides, its transfer performances on the other 6 datasets show that MixMAE has better FLOPs / performance tradeoff than previous popular MIM methods.

Table of Contents

  • 1. Introduction
  • 2. Related Works
  • 3. Methodology
  • 3.1. A Revisit of Masked Image Modeling
  • 3.2. Mixed and Masked Autoencoder (MixMAE)
  • 3.3. Reducing the Difficulty of the Pretext Task
  • 4. Experimental Setup
  • 5. Main Results
  • 5.1. Results on ImageNet-1K
  • 5.2. Results of Transferring to Downstream Tasks
  • 6. Ablation Studies
  • 7. Discussion and Conclusion
  • Acknowledgement
  • References

Knowls

  1. Knowl 1 — MixMAE Pretraining Framework

    model/method

    Mixed and Masked Autoencoder (MixMAE) is a self-supervised masked image modeling framework designed for hierarchical Vision Transformers. Traditional masked image modeling methods for hierarchical vision transformers replace masked patches with special [MASK] tokens, causing pretrain-finetuning inconsistency and redundant computation in the encoder. In contrast, plain masked autoencoders (MAE) completely remove masked tokens from the encoder, which is incompatible with hierarchical architectures that require a regular 2D grid of tokens.

    MixMAE addresses both issues by synthesizing a mixed 2D image from two (or more) unlabelled images without introducing [MASK] symbols into the encoder input. Given two images divided into non-overlapping patch sets x1px_1^p and x2px_2^p, a binary mixing mask M∈{0,1}N\mathrm{M} \in \{0, 1\}^N selects visible patches from either image to form a full 2D mixed image:

    x^mp=x1p⊙M+x2p⊙(1−M)\hat{x}^p_m = x^p_1 \odot \mathrm{M} + x^p_2 \odot (1 - \mathrm{M})

    where ⊙\odot represents element-wise multiplication.

    The complete mixed image is processed by a hierarchical Vision Transformer encoder. Prior to decoding, the output latent token embeddings are separated (unmixed) into two distinct groups according to M\mathrm{M}. Learnable [MASK] tokens are inserted into the missing positions of each stream, and a shared lightweight Transformer decoder reconstructs the original pixel patches of both input images simultaneously.

  2. Knowl 2 — Dual Reconstruction Loss Formulation

    equation

    MixMAE trains its hierarchical encoder and lightweight decoder using a dual reconstruction loss computed across the masked patches of both constituent images in the mixed input. Let x1p,x2p∈RN×Dx_1^p, x_2^p \in \mathbb{R}^{N \times D} denote the original patch sequences of two images, M∈{0,1}N\mathrm{M} \in \{0, 1\}^N the binary patch selection mask where M=1\mathrm{M}=1 denotes patches drawn from x1px_1^p and M=0\mathrm{M}=0 from x2px_2^p, and y1p,y2p∈RN×Dy_1^p, y_2^p \in \mathbb{R}^{N \times D} the corresponding pixel patch predictions from the decoder. The reconstruction loss Lrec\mathcal{L}_{rec} is formulated as:

    Lrec=∥(y1p−x1p)⊙(1−M)∥22+∥(y2p−x2p)⊙M∥22\mathcal{L}_{rec} = \|(y_1^p - x_1^p) \odot (1-\mathrm{M})\|_2^2 + \|(y_2^p - x_2^p) \odot \mathrm{M}\|_2^2

    where ⊙\odot represents element-wise multiplication and ∥⋅∥22\|\cdot\|_2^2 is the squared L2L_2 norm. The loss is computed only on the masked locations for each image using normalized raw pixel targets.

  3. Knowl 3 — Masked Self-Attention for Pretext Task Difficulty Reduction

    model/method

    Reconstructing two distinct images from a single mixed token sequence presents an optimization challenge due to semantic interference between tokens of different source images. MixMAE resolves this difficulty by enforcing masked self-attention within the hierarchical encoder.

    Under masked self-attention, attention fields are explicitly constrained such that each visual token is only permitted to attend to tokens originating from the same source image. The spatial mixing mask M∈{0,1}N\mathrm{M} \in \{0, 1\}^N is resized via nearest-neighbor interpolation to match the feature map spatial resolution at each stage of the hierarchical encoder. This binary mask is applied directly to the self-attention affinity matrices to mask out cross-image attention weights. This mechanism incurs negligible computational and parameter overhead, accelerates training convergence, and outperforms implicit image separation via stage-wise mix embeddings.

  4. Knowl 4 — Hierarchical Vision Transformer Encoder Architecture in MixMAE

    model/method

    MixMAE utilizes a modified Swin Transformer as its default hierarchical encoder backbone. The input image (or mixed image) is projected into non-overlapping patches and processed through 4 sequential stages of Transformer blocks with downsampling layers between stages, reaching a total downsampling ratio of 1/321/32. Unlike standard Swin Transformer, MixMAE removes shifted window attention and instead uses an enlarged window size (14×1414 \times 14 for stages 1, 2, and 3; 7×77 \times 7 for stage 4 on 224×224224 \times 224 inputs), which effectively conducts global self-attention in stages 3 and 4 to capture global context under high masking ratios.

    The encoder architectures are parameterized by channel dimensions CC, head counts HH, and block depths BB per stage:

    • Base (Swin-B): C=(128,256,512,1024)C = (128, 256, 512, 1024), H=(4,8,16,32)H = (4, 8, 16, 32), B=(2,2,18,2)B = (2, 2, 18, 2) (~88M parameters).
    • Large (Swin-L): C=(192,384,768,1536)C = (192, 384, 768, 1536), H=(6,12,24,48)H = (6, 12, 24, 48), B=(2,2,18,2)B = (2, 2, 18, 2) (~235M parameters).
    • Huge (Swin-H): C=(352,704,1408,2816)C = (352, 704, 1408, 2816), H=(11,22,44,88)H = (11, 22, 44, 88), B=(2,2,18,2)B = (2, 2, 18, 2) (~600M parameters).

    A linear projection maps the final encoder output dimension to 512 before feeding into an 8-block lightweight Transformer decoder.

  5. Knowl 5 — ImageNet-1K Pretraining and Finetuning Comparison

    data/table

    MixMAE achieves superior ImageNet-1K fine-tuning and linear probing performance compared to existing masked image modeling methods using comparable model sizes while requiring substantially fewer pretraining epochs.

    Method Backbone FLOPs (G) Param. (M) Pretrain Epochs FT Top-1 (%) LIN Top-1 (%)
    ViT ViT-B 17.5 86 14 79.9 -
    BEiT ViT-B 17.6 87 800 83.2 37.6
    CAE ViT-B 17.5 86 800 83.6 68.6
    MaskFeat ViT-B 17.5 86 300 83.6 -
    data2vec ViT-B 17.5 86 800 84.2 -
    iBOT ViT-B 17.5 86 1600 84.0 79.5
    PeCo ViT-B 17.5 86 800 84.5 -
    MAE ViT-B 17.5 86 1600 83.6 67.8
    MAE Swin-B/W14 16.3 88 600 84.4 61.0
    EsViT Swin-B/W14 16.3 87 300 83.9 81.3
    SimMIM ViT-B 17.5 86 800 83.8 56.7
    SimMIM Swin-B 15.6 88 800 84.0 -
    SimMIM Swin-B/W14 16.3 88 300 84.1 20.2
    GreenMIM Swin-B 15.6 88 800 83.8 -
    GreenMIM Swin-B/W14 16.3 88 800 84.1 -
    MixMAE Swin-B 15.6 88 600 84.6 61.2
    MixMAE Swin-B/W14 16.3 88 300 84.8 63.8
    MixMAE Swin-B/W14 16.3 88 600 85.1 71.0

    MixMAE with Swin-B/W14 reaches 84.8% top-1 accuracy in 300 epochs and 85.1% in 600 epochs, outperforming BEiT (+1.9% at 600 epochs) and SimMIM (+1.0% vs. 800-epoch SimMIM).

  6. Knowl 6 — Generalization Across Hierarchical Vision Backbones

    data/table

    MixMAE applies across multiple hierarchical Vision Transformer architectures, consistently outperforming SimMIM, GreenMIM, and supervised baselines on ImageNet-1K classification.

    Pretrain Method Backbone Pretrain Epochs Finetune Epochs Top-1 Acc. (%)
    Supervised Swin-B - 300 83.5
    SimMIM Swin-B 800 100 84.0
    GreenMIM Swin-B 800 100 83.8
    MixMAE Swin-B 600 100 84.6
    Supervised Swin-L - 300 83.5
    SimMIM Swin-L 800 100 85.4
    GreenMIM Swin-L 800 100 85.1
    MixMAE Swin-L 600 50 85.9
    Supervised PVT-L - 300 81.7
    SimMIM PVT-L 800 100 82.0
    MixMAE PVT-L 600 100 83.4
    Supervised Twins-SVT-L - 300 83.7
    SimMIM Twins-SVT-L 800 100 83.3
    GreenMIM Twins-SVT-L 800 100 83.9
    MixMAE Twins-SVT-L 600 100 83.9

    MixMAE outperforms SimMIM on Swin-B (+0.6%), Swin-L (+0.5% with half the finetuning epochs), and PVT-L (+1.4%), while reducing pretraining epochs from 800 to 600.

  7. Knowl 7 — Downstream Transfer to Object Detection, Instance Segmentation, and Semantic Segmentation

    data/table

    MixMAE transfers effectively to dense downstream tasks, including object detection/instance segmentation on COCO (using Mask R-CNN) and semantic segmentation on ADE20K (using UperNet), achieving higher accuracy and lower inference computational complexity than plain ViT methods.

    Method Backbone Pretrain Ep. COCO FLOPs(G) COCO APboxAP^{box} COCO APmaskAP^{mask} ADE FLOPs(G) ADE Params(M) ADE mIoU
    MoCo v3 ViT-B 300 853 47.9 42.9 606 164 47.3
    BEiT ViT-B 800 853 49.8 44.4 606 164 47.1
    MAE ViT-B 1600 853 50.3 44.9 606 164 48.1
    iBOT ViT-B 1600 - 51.2 44.2 - - 50.0
    EsViT Swin-B 300 - - - - - 47.3
    SimMIM Swin-B 800 - 52.3 - - - 52.8
    SimMIM Swin-B/W14 300 701 51.1 45.4 302 122 48.9
    GreenMIM Swin-B 800 - 50.0 44.1 - - -
    MixMAE Swin-B/W14 300 701 52.3 46.4 302 122 49.9
    MixMAE Swin-B/W14 600 701 52.7 47.0 302 122 51.1
    MoCo v3 ViT-L 300 1907 49.3 43.9 877 392 49.1
    BEiT ViT-L 800 1907 53.3 47.1 877 392 53.3
    MAE ViT-L 1600 1907 53.3 47.2 877 392 53.6
    SimMIM Swin-L 800 - 53.8 - - - 53.5
    MixMAE Swin-L 600 1119 54.3 48.2 460 236 53.8

    MixMAE Swin-L achieves 54.3 APboxAP^{box} and 48.2 APmaskAP^{mask} on COCO and 53.8 mIoU on ADE20K, outperforming 1600-epoch MAE ViT-L with 41% fewer inference FLOPs on COCO and 48% fewer inference FLOPs on ADE20K.

  8. Knowl 8 — Transfer Evaluation on Fine-Grained and Scene Classification

    data/table

    MixMAE evaluated on fine-grained and scene recognition benchmarks (iNaturalist 2018, iNaturalist 2019, Places205, and Places365) demonstrates consistent performance gains over supervised and self-supervised baselines.

    Method Backbone FLOPs (G) Param. (M) INat2018 INat2019 Places205 Places365 Average
    DINO ViT-B 17.5 86 72.6 78.2 - - -
    MAE ViT-B 17.5 86 75.4 80.5 63.9 57.9 69.4
    MixMAE Swin-B/W14 16.3 88 78.2 83.3 68.6 59.0 72.3
    MAE ViT-L 61.3 304 80.1 83.4 65.8 59.4 72.1
    MixMAE Swin-L 35.8 235 80.6 84.4 69.3 59.6 73.5

    MixMAE Swin-B/W14 improves average accuracy across the 4 classification datasets by +2.9% over MAE ViT-B. MixMAE Swin-L exceeds MAE ViT-L by +1.4% on average while requiring 42% fewer inference FLOPs (35.8G vs. 61.3G).

  9. Knowl 9 — Ablation of Token Filling Strategies

    data/table

    An ablation study comparing strategies for populating masked patch locations in the input demonstrates that filling with tokens from another real image (Mix) yields the best representation quality.

    Filling Type ImageNet-1K Top-1 Acc. (%) ADE20K mIoU
    Mix (Default) 84.6 49.9
    Zero 84.1 48.0
    Learnable 84.1 48.9
    Shuffle 82.6 43.0
    Zoomin 83.5 44.9

    All models use Swin-B/W14 pretrained for 300 epochs at a 50% masking ratio. 'Zero' fills masked tokens with 0s; 'Learnable' uses a shared learnable vector; 'Shuffle' fills masked positions with randomly shuffled tokens from the same image; 'Zoomin' fills masked positions with patches from a zoomed-in crop of the same image. Mix outperforms all alternative filling contents on both ImageNet-1K classification (+0.5% over Zero/Learnable) and ADE20K semantic segmentation (+1.0 to +1.9 mIoU).

  10. Knowl 10 — Ablation of Masking Ratio and Number of Mixed Images

    data/table

    Evaluating the number of mixed images used to construct the input sequence shows that higher masking ratios achieved by mixing more images provide superior representation learning compared to adding standard [MASK] tokens.

    Number of Images Masking Ratio ImageNet-1K Top-1 Acc. (%) ADE20K mIoU
    2 0.50 84.6 49.9
    2 w/ [MASK] 0.75 84.4 49.0
    3 0.67 84.7 49.9
    4 (Default) 0.75 84.8 49.9
    5 0.80 84.5 49.5

    All entries are pretrained on Swin-B/W14 for 300 epochs. A 75% masking ratio achieved by mixing 4 images yields the highest performance (84.8% top-1). In contrast, achieving a 75% masking ratio by combining 2 mixed images with [MASK] tokens degrades accuracy to 84.4% and mIoU to 49.0, demonstrating that replacing [MASK] tokens with real image tokens is crucial.

  11. Knowl 11 — Ablations on Dual Reconstruction and Difficulty Reduction Methods

    data/table

    Ablation experiments show that dual reconstruction and masked self-attention are essential for the performance and convergence of MixMAE.

    Dual Reconstruction (300 Epochs, Swin-B/W14):

    Dual Reconstruction ImageNet-1K Top-1 Acc. (%) ADE20K mIoU
    Yes (Dual) 84.6 49.9
    No (Single) 84.0 47.3

    Reconstructing both images simultaneously provides a +0.6% gain on ImageNet-1K and a +2.6 gain in ADE20K mIoU over reconstructing only a single image.

    Difficulty Reduction Across Pretraining Epochs (ImageNet-1K Top-1 Acc. %):

    Approach 300 Epochs 600 Epochs 900 Epochs
    Without unmixing / separation 84.4 84.4 84.4
    Mix embedding 84.4 84.6 84.8
    Masked self-attention (Default) 84.6 85.1 85.1

    Direct reconstruction without unmixing saturates early at 84.4%. Mix embedding improves to 84.8% with 900 epochs, while masked self-attention converges faster, reaching 85.1% top-1 accuracy by 600 epochs.

Coverage note — None was omitted; all key architectural components, objectives, benchmarks, and ablation studies from the paper are represented.

References

  1. 1.Jean-Baptiste Alayrac, Joao Carreira, and Andrew Zisserman. The visual centrifuge: Model-free layered video representations. In CVPR, 2019.
  2. 2.Sara Atito, Muhammad Awais, and Josef Kittler. Sit: Self-supervised vision transformer. arXiv:2104.03602, 2021.
  3. 3.Alexei Baevski, Wei-Ning Hsu, Qiantong Xu, Arun Babu, Jiatao Gu, and Michael Auli. Data2vec: A general framework for self-supervised learning in speech, vision and language. arXiv:2202.03555, 2022.
  4. 4.Hangbo Bao, Li Dong, and Furu Wei. Beit: Bert pre-training of image transformers. In ICLR, 2021.
  5. 5.Mathilde Caron, Hugo Touvron, Ishan Misra, Herve J egou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. Emerging properties in self-supervised vision transformers. In ICCV, 2021.
  6. 6.Xiaokang Chen, Mingyu Ding, Xiaodi Wang, Ying Xin, Shentong Mo, Yunhao Wang, Shumin Han, Ping Luo, Gang Zeng, and Jingdong Wang. Context autoencoder for self-supervised representation learning. arXiv:2202.03026, 2022.
  7. 7.Xinlei Chen, Saining Xie, and Kaiming He. An empirical study of training self-supervised vision transformers. In ICCV, 2021.
  8. 8.Xiangxiang Chu, Zhi Tian, Yuqing Wang, Bo Zhang, Haibing Ren, Xiaolin Wei, Huaxia Xia, and Chunhua Shen. Twins: Revisiting the design of spatial attention in vision transformers. NeurIPS, 34:9355–9366, 2021.
  9. 9.Kevin Clark, Minh-Thang Luong, Quoc V Le, and Christopher D Manning. Electra: Pre-training text encoders as discriminators rather than generators. In ICLR, 2020.
  10. 10.Zihang Dai, Hanxiao Liu, Quoc V Le, and Mingxing Tan. Coatnet: Marrying convolution and attention for all data sizes. NeurIPS, 34:3965–3977, 2021.
  11. 11.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, K. Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In CVPR, 2009.
  12. 12.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. In NAACL, 2019.
  13. 13.Xiaoyi Dong, Jianmin Bao, Ting Zhang, Dongdong Chen, Weiming Zhang, Lu Yuan, Dong Chen, Fang Wen, and Nenghai Yu. Peco: Perceptual codebook for bert pre-training of vision transformers. arXiv:2111.12710, 2021.
  14. 14.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition at scale. In ICLR, 2021.
  15. 15.Raghav Goyal, Samira Ebrahimi Kahou, Vincent Michalski, Joanna Materzynska, Susanne Westphal, Heuna Kim, Valentin Haenel, Ingo Frund, Peter N. Yianilos, Moritz Mueller-Freitag, Florian Hoppe, Christian Thurau, Ingo Bax, and Roland Memisevic. The “something something” video database for learning and evaluating visual common sense. In ICCV, 2017.
  16. 16.Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Doll’ar, and Ross B. Girshick. Masked autoencoders are scalable vision learners. In CVPR, 2021.
  17. 17.Kaiming He, Georgia Gkioxari, Piotr Dollar, and Ross Gir- shick. Mask r-cnn. In ICCV, 2017.
  18. 18.Grant Van Horn, Oisin Mac Aodha, Yang Song, Yin Cui, Chen Sun, Alexander Shepard, Hartwig Adam, Pietro Perona, and Serge J. Belongie. The inaturalist species classification and detection dataset. In CVPR, 2018.
  19. 19.Gao Huang, Yu Sun, Zhuang Liu, Daniel Sedra, and Kilian Q Weinberger. Deep networks with stochastic depth. In ECCV, 2016.
  20. 20.Lang Huang, Shan You, Mingkai Zheng, Fei Wang, Chen Qian, and Toshihiko Yamasaki. Green hierarchical vision transformer for masked image modeling. arXiv preprint arXiv:2205.13515, 2022.
  21. 21.Chunyuan Li, Jianwei Yang, Pengchuan Zhang, Mei Gao, Bin Xiao, Xiyang Dai, Lu Yuan, and Jianfeng Gao. Efficient self-supervised vision transformers for representation learning. In ICLR, 2022.
  22. 22.Xiang Li, Wenhai Wang, Lingfeng Yang, and Jian Yang. Uniform masking: Enabling mae pre-training for pyramid-based vision transformers with locality. arXiv:2205.10063, 2022.
  23. 23.Yanghao Li, Saining Xie, Xinlei Chen, Piotr Dollar, Kaiming He, and Ross Girshick. Benchmarking detection transfer learning with vision transformers. arXiv:2111.11429, 2021.
  24. 24.Tsung-Yi Lin, Piotr Dollar, Ross Girshick, Kaiming He, Bharath Hariharan, and Serge Belongie. Feature pyramid networks for object detection. In CVPR, 2017.
  25. 25.Tsung-Yi Lin, Michael Maire, Serge J. Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollar, and C. Lawrence Zitnick. Microsoft coco: Common objects in context. In ECCV, 2014.
  26. 26.Jihao Liu, Xin Huang, Guanglu Song, Hongsheng Li, and Yu Liu. Uninet: Unified architecture search with convolution, transformer, and mlp. In ECCV, pages 33–49. Springer, 2022.
  27. 27.Ze Liu, Han Hu, Yutong Lin, Zhuliang Yao, Zhenda Xie, Yixuan Wei, Jia Ning, Yue Cao, Zheng Zhang, Li Dong, et al. Swin transformer v2: Scaling up capacity and resolution. In CVPR, pages 12009–12019, 2022.
  28. 28.Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. In ICCV, 2021.
  29. 29.Mehdi Noroozi and Paolo Favaro. Unsupervised learning of visual representations by solving jigsaw puzzles. In ECCV, 2016.
  30. 30.Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever. Zero-shot text-to-image generation. In ICML, 2021.
  31. 31.Jason Tyler Rolfe. Discrete variational autoencoders. In ICLR, 2016.
  32. 32.Chen Sun, Abhinav Shrivastava, Saurabh Singh, and Abhinav Gupta. Revisiting unreasonable effectiveness of data in deep learning era. In ICCV, 2017.
  33. 33.Zhan Tong, Yibing Song, Jue Wang, and Limin Wang. Video-mae: Masked autoencoders are data-efficient learners for self-supervised video pre-training. arXiv:2203.12602, 2022.
  34. 34.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. NeurIPS, 30, 2017.
  35. 35.Wenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan, Kaitao Song, Ding Liang, Tong Lu, Ping Luo, and Ling Shao. Pyramid vision transformer: A versatile backbone for dense prediction without convolutions. In ICCV, 2021.
  36. 36.Chen Wei, Haoqi Fan, Saining Xie, Chao-Yuan Wu, Alan Yuille, and Christoph Feichtenhofer. Masked feature prediction for self-supervised visual pre-training. arXiv:2112.09133, 2021.
  37. 37.Yixuan Wei, Han Hu, Zhenda Xie, Zheng Zhang, Yue Cao, Jianmin Bao, Dong Chen, and Baining Guo. Contrastive learning rivals masked image modeling in fine-tuning via feature distillation. arXiv:2205.14141, 2022.
  38. 38.Tete Xiao, Yingcheng Liu, Bolei Zhou, Yuning Jiang, and Jian Sun. Unified perceptual parsing for scene understanding. In ECCV, 2018.
  39. 39.Tete Xiao, Ilija Radosavovic, Trevor Darrell, and Jitendra Malik. Masked visual pre-training for motor control. arXiv:2111.11429, 2022.
  40. 40.Zhenda Xie, Zheng Zhang, Yue Cao, Yutong Lin, Jianmin Bao, Zhuliang Yao, Qi Dai, and Han Hu. Simmim: A simple framework for masked image modeling. In CVPR, 2021.
  41. 41.Bolei Zhou, Agata Lapedriza, Jianxiong Xiao, Antonio Tor- ralba, and Aude Oliva. Learning deep features for scene recognition using places database. In NeurIPS, 2014.
  42. 42.Bolei Zhou, Hang Zhao, Xavier Puig, Sanja Fidler, Adela Barriuso, and Antonio Torralba. Scene parsing through ade20k dataset. In CVPR, 2017.
  43. 43.Jinghao Zhou, Chen Wei, Huiyu Wang, Wei Shen, Cihang Xie, Alan Yuille, and Tao Kong. ibot: Image bert pre-training with online tokenizer. In ICLR, 2021.

Citation

MLA
Liu, J., et al. “MixMAE: Mixed and Masked Autoencoder for Efficient Pretraining of Hierarchical Vision Transformers”. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023, pp. 6252–61, https://doi.org/10.1109/CVPR52729.2023.00605.
APA
Liu, J., Huang, X., Zheng, J., Liu, Y., & Li, H. (2023). MixMAE: Mixed and Masked Autoencoder for Efficient Pretraining of Hierarchical Vision Transformers. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 6252–6261. https://doi.org/10.1109/CVPR52729.2023.00605
Chicago
Liu, J., X. Huang, J. Zheng, Y. Liu, and H. Li. 2023. “MixMAE: Mixed and Masked Autoencoder for Efficient Pretraining of Hierarchical Vision Transformers”. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 6252–61. https://doi.org/10.1109/CVPR52729.2023.00605.
Harvard
Liu, J. et al. (2023) “MixMAE: Mixed and Masked Autoencoder for Efficient Pretraining of Hierarchical Vision Transformers”, 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp. 6252–6261. Available at: https://doi.org/10.1109/CVPR52729.2023.00605.
Vancouver
1. Liu J, Huang X, Zheng J, Liu Y, Li H (2023) MixMAE: Mixed and Masked Autoencoder for Efficient Pretraining of Hierarchical Vision Transformers. In: 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp 6252–6261

BibTeX

@inproceedings{Liu_2023, title={MixMAE: Mixed and Masked Autoencoder for Efficient Pretraining of Hierarchical Vision Transformers}, url={http://dx.doi.org/10.1109/CVPR52729.2023.00605}, DOI={10.1109/cvpr52729.2023.00605}, booktitle={2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)}, publisher={IEEE}, author={Liu, Jihao and Huang, Xin and Zheng, Jinliang and Liu, Yu and Li, Hongsheng}, year={2023}, month=June, pages={6252–6261} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE