EVA: Exploring the Limits of Masked Visual Representation Learning at Scale

Yuxin FangWen WangBinhui XieQuan SunLedell WuXinggang WangTiejun HuangXinlong WangYue Cao

article2023CVPR1,042 citations

Demonstrates that pre-training a one-billion-parameter Vision Transformer to reconstruct image-text aligned features using only public data establishes state-of-the-art transfer performance across major vision tasks and efficiently stabilizes the training of large multimodal models.

Listen

Scaling deep learning models using self-supervised masked pre-training has transformed natural language processing, but vision foundation models have lagged behind, frequently depending on proprietary datasets and heavy supervised training. Natural images are raw and sparse, making it difficult for models to capture high-level visual meaning solely from low-level pixel reconstruction. The article introduces EVA, a vision-centric foundation model designed to explore the limits of visual representation learning at scale using only publicly available data.

The article demonstrates the training of a vanilla Vision Transformer with one billion parameters using an efficient masked image modeling pretext task. The approach involves conditioning the model on visible image patches to directly reconstruct masked-out visual features derived from an image-text aligned teacher model, OpenAI CLIP-L/14. The pre-training relies entirely on 29.6 million publicly accessible unlabeled images across multiple standard datasets, without requiring complex semantic tokenization or paired image-text captions.

EVA achieves state-of-the-art results across several major vision benchmarks. For image classification, the model achieves 89.7% top-1 accuracy on ImageNet-1K with minimal supervised fine-tuning, while demonstrating superior out-of-distribution robustness with an average performance gap of only 5.6% across six robustness variants. In dense object-level tasks, EVA establishes new records on COCO and demonstrates an emergent capability on the complex LVIS benchmark, achieving an identical 55.0 mask average precision on both datasets despite LVIS having over 1,200 categories compared to COCO's 80. In video understanding, EVA achieves 89.7% accuracy on Kinetics-400 and 82.9% on Kinetics-700. When serving as the vision tower for a giant 1.1-billion-parameter CLIP model, EVA achieves an average zero-shot classification accuracy of 75.7% across 12 benchmarks, outperforming larger models.

These findings indicate that masked visual feature reconstruction bridges low-level geometric structures and high-level visual semantics without requiring proprietary datasets. In multi-modal foundation models, using EVA to initialize the vision tower stabilizes training and drastically cuts resource costs, enabling 16-bit floating-point optimization on 256 GPUs with about one-third of the hardware and data requirements of competing approaches. Decision-makers can leverage EVA as a foundational architecture to lower training costs, eliminate dependencies on private datasets, and accelerate the development of large-scale vision and multi-modal systems.

Future efforts should explore applying this masked pre-training strategy across even broader multi-modal workflows and scaling beyond one billion parameters. Hardware memory constraints limited full model adaptation on dense prediction benchmarks, such as ADE20K semantic segmentation, where the model used fewer decoders and slightly trailed competing systems. Despite these boundary constraints, the extensive validation across multiple public benchmarks provides strong confidence in the scalability, transferability, and stability of the EVA framework.

No sufficiently relevant recommendations were found.

Cover for EVA: Exploring the Limits of Masked Visual Representation Learning at Scale

Abstract

We launch EVA, a vision-centric foundation model to explore the limits of Visual representation at scAle using only publicly accessible data. EVA is a vanilla ViT pre-trained to reconstruct the masked out image-text aligned vision features conditioned on visible image patches. Via this pretext task, we can efficiently scale up EVA to one billion parameters, and sets new records on a broad range of representative vision downstream tasks, such as image recognition, video action recognition, object detection, instance segmentation and semantic segmentation without heavy supervised training. Moreover, we observe quantitative changes in scaling EVA result in qualitative changes in transfer learning performance that are not present in other models. For instance, EVA takes a great leap in the challenging large vocabulary instance segmentation task: our model achieves almost the same state-of-the-art performance on LVIS dataset with over a thousand categories and COCO dataset with only eighty categories. Beyond a pure vision encoder, EVA can also serve as a vision-centric, multi-modal pivot to connect images and text. We find initializing the vision tower of a giant CLIP from EVA can greatly stabilize the training and outperform the training from scratch counterpart with much fewer samples and less compute, providing a new direction for scaling up and accelerating the costly training of multi-modal foundation models.

Table of Contents

  • 1 Introduction
  • 2 Fly EVA to the Moon
  • 2.1 The Feature Instrumentality Project
  • 2.2 Pre-training
  • 2.3 Evaluation on Downstream Tasks
  • 2.3.1 Image Classification
  • 2.3.2 Video Action Recognition
  • 2.3.3 Object Detection & Instance Segmentation
  • 2.3.4 Semantic Segmentation
  • 2.3.5 Contrastive Language-Image Pre-training with Zero-shot Classification Evaluation
  • 3 Related Work
  • 4 Conclusion
  • A Appendix
  • A.1 Image Classification
  • A.2 Video Action Classification
  • A.3 Object Detection & Instance Segmentation
  • A.4 Semantic Segmentation
  • References

Knowls

  1. Knowl 1 — Masked prediction of image–text-aligned visual features

    model/method

    EVA pre-trains a vision transformer by predicting the image–text-aligned visual features of masked image regions from the visible image patches. Its targets are features from the publicly available OpenAI CLIP-L/14 vision encoder, trained on 224 × 224 pixel images. EVA normalizes its output features, maps them through a linear projection to the CLIP feature dimension, and minimizes negative cosine similarity between predictions and targets. This feature-regression objective combines semantic targets from image–text learning with masked-image context, without requiring a visual tokenizer or paired image–text data in EVA's own pre-training corpus.

  2. Knowl 2 — EVA’s billion-parameter architecture and pre-training recipe

    experimental setup

    EVA is a vanilla ViT with 1,011 million parameters: 14 × 14 pixel patches, 40 layers, hidden dimension 1,408, MLP dimension 6,144, and 16 attention heads. It was pre-trained for 150 epochs at 224 × 224 resolution with block-wise masking of 40% of the patches. The 29.6 million publicly accessible pre-training images came from ImageNet-21K, CC12M, CC3M, Objects365, COCO, and ADE20K; CC12M and CC3M captions were not used, and only the training splits of COCO and ADE20K were used. Optimization used AdamW with peak learning rate 0.001, betas (0.9, 0.98), weight decay 0.05, and cosine learning-rate decay. Stochastic depth was 0.1, and RandResizeCrop used a scale range of 0.2–1.0. Training used fp16, ZeRO stage 1, and 128 NVIDIA A100 40GB GPUs; the reported throughput was about 3,150 samples per second, peak memory about 26.5 GB, and training duration about 14.5 days. EVA used neither relative positional embeddings nor layer-scale during pre-training.

  3. Knowl 3 — Pilot study favors direct feature regression over tokenization or distillation

    empirical result

    In pilot experiments with ViT-B, the authors compared masked prediction of CLIP visual features with and without additional feature tokenization, and feature distillation, evaluating ImageNet-1K top-1 accuracy and ADE20K single-scale mIoU. Without tokenization, 800 pre-training epochs produced 85.5% ImageNet-1K accuracy and 53.3 ADE20K mIoU; tokenized-feature prediction at 1,600 epochs reached 85.5% and 53.1, respectively. For distillation, 800 epochs with distillation reached 85.1% and 52.7, whereas 800 epochs without distillation reached 85.5% and 53.3. The shorter tokenized and distilled runs also did not show a consistent advantage over direct fine-tuning of the CLIP vision encoder. These pilot results motivated scaling the simpler direct masked-feature regression approach.

  4. Knowl 4 — ImageNet accuracy and robustness under distribution shifts

    empirical result

    After intermediate fine-tuning on ImageNet-21K for 60 epochs at 224 × 224 resolution and ImageNet-1K fine-tuning for 10 epochs, EVA achieved 89.6% top-1 accuracy on the ImageNet-1K validation set with 336 × 336 inputs, and 89.7% with 560 × 560 inputs. The ImageNet-1K classifier was a linear layer. In evaluations using one fine-tuned model without specialized fine-tuning on six ImageNet validation sets, EVA achieved top-1 accuracies of 89.6% on ImageNet-1K, 81.6% on ImageNet-V2, 90.8% on ImageNet-ReaL, 86.2% on ImageNet-Adversarial, 88.3% on ImageNet-Rendition, and 67.7% on ImageNet-Sketch. Its average across those six sets was 84.0%, and the reported gap between that average and its ImageNet-1K accuracy was 5.6 percentage points—the smallest such gap among the compared models.

  5. Knowl 5 — EVA sharply narrows the LVIS–COCO instance-segmentation gap

    empirical result

    In a controlled comparison using the official COCO and LVIS validation annotations, EVA’s single-scale results were 64.1 box AP and 55.0 mask AP on COCO, and 62.2 box AP and 55.0 mask AP on LVIS. The COCO–LVIS gaps were therefore 1.9 AP for boxes and 0.0 AP for masks. For comparison, the best individual prior results in the paper had gaps of 9.8 box AP and 5.3 mask AP. The authors also evaluated both datasets on their shared LVIS val-5K images, using LVIS annotations and the corresponding 80-category COCO subset: EVA reached 69.6 versus 68.3 box AP and 59.6 versus 59.8 mask AP on COCO and LVIS, respectively. The pre-training corpus included 15,000 images that were also in the 20,000-image LVIS validation set; the shared-image evaluation was reported as a check against potential contamination, and the authors said the gap-reduction conclusion remained unchanged.

  6. Knowl 6 — COCO detection and instance segmentation with Cascade Mask R-CNN

    empirical result

    EVA was evaluated with Cascade Mask R-CNN, using Objects365 for intermediate detector fine-tuning and then fine-tuning on COCO. On COCO validation, the no-test-time-augmentation result was 64.2 box AP and 55.0 mask AP; with test-time augmentation, box AP was 64.5. On COCO test-dev, the no-augmentation result was 64.4 box AP and 55.5 mask AP; with test-time augmentation, box AP was 64.7. The test-time-augmented rows did not report mask AP. These results used the classic R-CNN-family detector rather than a DETR-style detector, and were reported as state-of-the-art results in the paper’s November 2022 comparison.

  7. Knowl 7 — EVA initializes a more resource-efficient billion-scale CLIP model

    model/method

    For EVA CLIP, the 1.0-billion-parameter EVA encoder initializes the vision tower, while the language tower is initialized from OpenAI CLIP-L. The resulting model has 1.1 billion parameters in total, including a 124-million-parameter text tower, and is trained contrastively on LAION-400M. The reported configuration used fp16 with dynamic loss scaling, a batch size of 41,000, 256 NVIDIA A100 40GB GPUs, and 11 billion samples seen. The authors report that this training was stable without bfloat16. Relative to the compared OpenCLIP-H and OpenCLIP-g configurations, EVA CLIP used a smaller image–text data pool and fewer GPUs; the paper presents EVA initialization as a way to improve training stability and reduce the resources needed to train a large CLIP model.

  8. Knowl 8 — EVA CLIP’s zero-shot transfer across image and video benchmarks

    empirical result

    Across the 12 zero-shot image and video classification benchmarks reported in the paper, EVA CLIP achieved the highest average performance and was best on 10 benchmarks. Its tabulated image top-1 accuracies were 78.5% on ImageNet-1K, 71.5% on ImageNet-V2, 73.6% on ImageNet-Adversarial, 92.5% on ImageNet-Rendition, 67.3% on ImageNet-Sketch, 72.3% on ObjectNet, 98.3% on CIFAR-10, and 88.7% on CIFAR-100. Its video accuracies were 76.1% on UCF-101, 65.2% on Kinetics-400, 64.4% on Kinetics-600, and 58.4% on Kinetics-700. For the six image benchmarks used to assess natural distribution shifts, the reported accuracy gap from ImageNet-1K was 2.5 percentage points, the smallest among the compared CLIP models. On ImageNet-1K, the EVA CLIP vision encoder also reached 78.5% zero-shot, 86.5% with linear probing, and 89.4% after fine-tuning; the compared prior best results were 78.0%, 82.3%, and 89.1%, respectively.

  9. Knowl 9 — Video recognition transfer with spatial–temporal attention

    empirical result

    EVA was adapted to video using spatial–temporal attention without a video-specific architectural redesign. The authors merged Kinetics-400, Kinetics-600, and Kinetics-700 into a deduplicated 722-class training set containing 0.63 million videos, trained for 40 epochs using 8 frames at 224 × 224 resolution, and then fine-tuned on each target dataset for 1–2 epochs. Fine-tuning and evaluation used 16 frames × 3 crops × 4 clips at 224 × 224. EVA achieved 89.7% top-1 accuracy on Kinetics-400, 89.8% on Kinetics-600, and 82.9% on Kinetics-700. Even without the merged-dataset intermediate fine-tuning, adapting image-pre-trained EVA to Kinetics-400 achieved 88.4% top-1 accuracy.

  10. Knowl 10 — Semantic-segmentation performance and adaptation constraint

    empirical result

    Using a ViT-Adapter plus Mask2Former transfer pipeline, EVA achieved 61.5 single-scale mIoU and 62.3 multi-scale mIoU on ADE20K, and 53.4 single-scale mIoU on COCO-Stuff-164K. The segmentation adaptation was constrained by 40 GB of GPU memory: relative position biases were omitted, the Mask2Former head used eight decoders rather than nine, and its feature dimension was about 0.6 times the EVA encoder dimension. The authors noted that EVA’s ADE20K result was slightly below BEiT-3’s and suggested that the weakened segmentation-head configuration may have contributed.

Coverage note — No substantial contributed result was deliberately omitted; exhaustive comparator rows and ancillary per-task implementation details were left out where they did not constitute separate contributions.

References

  1. 1.Clip: Connecting text and images. https://openai.com/blog/clip/. 7
  2. 2.Large scale openclip: L/14, h/14 and g/14 trained on laion-2b. https://laion.ai/blog/large-openclip. 2, 7
  3. 3.Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton. Layer normalization. arXiv preprint arXiv:1607.06450, 2016. 3
  4. 4.Alexei Baevski, Wei-Ning Hsu, Qiantong Xu, Arun Babu, Jiatao Gu, and Michael Auli. Data2vec: A general framework for self-supervised learning in speech, vision and language. arXiv preprint arXiv:2202.03555, 2022. 8
  5. 5.Hangbo Bao, Li Dong, and Furu Wei. Beit: Bert pre-training of image transformers. arXiv preprint arXiv:2106.08254, 2021. 1, 3, 4, 5, 8
  6. 6.Andrei Barbu, David Mayo, Julian Alverio, William Luo, Christopher Wang, Dan Gutfreund, Josh Tenenbaum, and Boris Katz. Objectnet: A large-scale bias-controlled dataset for pushing the limits of object recognition models. In NeurIPS, 2019. 7, 8
  7. 7.Lucas Beyer, Olivier J Henaff, Alexander Kolesnikov, Xiaohua Zhai, and Aaron van den Oord. Are we done with imagenet? arXiv preprint arXiv:2006.07159, 2020. 4
  8. 8.Navaneeth Bodla, Bharat Singh, Rama Chellappa, and Larry S Davis. Soft-nms–improving object detection with one line of code. In ICCV, 2017. 5
  9. 9.Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. arXiv preprint arXiv:2005.14165, 2020. 1
  10. 10.Minwoo Byeon, Beomhee Park, Haecheon Kim, Sungjun Lee, Woonhyuk Baek, and Saehoon Kim. Coyo-700m: Image-text pair dataset. https://github.com/kakaobrain/coyo-dataset, 2022. 7
  11. 11.Holger Caesar, Jasper Uijlings, and Vittorio Ferrari. Coco-stuff: Thing and stuff classes in context. In CVPR, 2018. 2, 6
  12. 12.Zhaowei Cai and Nuno Vasconcelos. Cascade r-cnn: high quality object detection and instance segmentation. TPAMI, 2019. 5, 6
  13. 13.Mathilde Caron, Hugo Touvron, Ishan Misra, Herve Jégou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. Emerging properties in self-supervised vision transformers. arXiv preprint arXiv:2104.14294, 2021. 1
  14. 14.Joao Carreira, Eric Noland, Andras Banki-Horvath, Chloe Hillier, and Andrew Zisserman. A short note about kinetics-600. arXiv preprint arXiv:1808.01340, 2018. 2, 4, 7, 8
  15. 15.Joao Carreira, Eric Noland, Chloe Hillier, and Andrew Zisserman. A short note on the kinetics-700 human action dataset. arXiv preprint arXiv:1907.06987, 2019. 2, 4, 7, 8
  16. 16.Soravit Changpinyo, Piyush Sharma, Nan Ding, and Radu Soricut. Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts. In CVPR, 2021. 3
  17. 17.Kai Chen, Jiangmiao Pang, Jiaqi Wang, Yu Xiong, Xiaoxiao Li, Shuyang Sun, Wansen Feng, Ziwei Liu, Jianping Shi, Wanli Ouyang, et al. Hybrid task cascade for instance segmentation. In CVPR, 2019. 5, 6
  18. 18.Mark Chen, Alec Radford, Rewon Child, Jeffrey Wu, Heewoo Jun, David Luan, and Ilya Sutskever. Generative pretraining from pixels. In ICML, 2020. 8
  19. 19.Qiang Chen, Jian Wang, Chuchu Han, Shan Zhang, Zexian Li, Xiaokang Chen, Jiahui Chen, Xiaodi Wang, Shuming Han, Gang Zhang, et al. Group detr v2: Strong object detector with encoder-decoder pretraining. arXiv preprint arXiv:2211.03594, 2022. 2, 5, 6
  20. 20.Tianqi Chen, Bing Xu, Chiyuan Zhang, and Carlos Guestrin. Training deep nets with sublinear memory cost. arXiv preprint arXiv:1604.06174, 2016. 3
  21. 21.Xiaokang Chen, Mingyu Ding, Xiaodi Wang, Ying Xin, Shentong Mo, Yunhao Wang, Shumin Han, Ping Luo, Gang Zeng, and Jingdong Wang. Context autoencoder for self-supervised representation learning. arXiv preprint arXiv:2202.03026, 2022. 8
  22. 22.Xinlei Chen, Saining Xie, and Kaiming He. An empirical study of training self-supervised vision transformers. arXiv preprint arXiv:2104.02057, 2021. 1
  23. 23.Zhe Chen, Yuchen Duan, Wenhai Wang, Junjun He, Tong Lu, Jifeng Dai, and Yu Qiao. Vision transformer adapter for dense predictions. arXiv preprint arXiv:2205.08534, 2022. 2, 6
  24. 24.Bowen Cheng, Ishan Misra, Alexander G Schwing, Alexander Kirillov, and Rohit Girdhar. Masked-attention mask transformer for universal image segmentation. arXiv preprint arXiv:2112.01527, 2021. 6
  25. 25.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. Palm: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311, 2022. 1
  26. 26.Xiyang Dai, Yinpeng Chen, Bin Xiao, Dongdong Chen, Mengchen Liu, Lu Yuan, and Lei Zhang. Dynamic head: Unifying object detection heads with attentions. In CVPR, 2021. 6
  27. 27.Zihang Dai, Hanxiao Liu, Quoc V Le, and Mingxing Tan. Coatnet: Marrying convolution and attention for all data sizes. NeurIPS, 2021. 4
  28. 28.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In CVPR, 2009. 2, 3, 4, 7, 8
  29. 29.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018. 1
  30. 30.Xiaoyi Dong, Jianmin Bao, Ting Zhang, Dongdong Chen, Weiming Zhang, Lu Yuan, Dong Chen, Fang Wen, and Nenghai Yu. Peco: Perceptual codebook for bert pre-training of vision transformers. arXiv preprint arXiv:2111.12710, 2021. 8
  31. 31.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition at scale. In ICLR, 2021. 1, 2, 3, 4, 8
  32. 32.Yuxin Fang, Li Dong, Hangbo Bao, Xinggang Wang, and Furu Wei. Corrupted image modeling for self-supervised visual pre-training. arXiv preprint arXiv:2202.03382, 2022. 8
  33. 33.Yuxin Fang, Quan Sun, Xinggang Wang, Tiejun Huang, Xinlong Wang, and Yue Cao. Eva-02: A visual representation for neon genesis. arXiv preprint arXiv:2303.11331, 2023. 6
  34. 34.Christoph Feichtenhofer, Haoqi Fan, Yanghao Li, and Kaiming He. Masked autoencoders as spatiotemporal learners. arXiv preprint arXiv:2205.09113, 2022. 5
  35. 35.WeiFu Fu, CongChong Nie, Ting Sun, Jun Liu, TianLiang Zhang, and Yong Liu. Lvis challenge track technical report 1st place solution: Distribution balanced and boundary refinement for large vocabulary instance segmentation. arXiv preprint arXiv:2111.02668, 2021. 2, 6
  36. 36.Golnaz Ghiasi, Yin Cui, Aravind Srinivas, Rui Qian, Tsung-Yi Lin, Ekin D. Cubuk, Quoc V. Le, and Barret Zoph. Simple copy-paste is a strong data augmentation method for instance segmentation. In CVPR, 2021. 5, 6
  37. 37.Ross Girshick, Jeff Donahue, Trevor Darrell, and Jitendra Malik. Rich feature hierarchies for accurate object detection and semantic segmentation. In CVPR, 2014. 5, 6
  38. 38.Agrim Gupta, Piotr Dollar, and Ross Girshick. Lvis: A dataset for large vocabulary instance segmentation. In CVPR, 2019. 2, 5
  39. 39.Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollar, and Ross Girshick. Masked autoencoders are scalable vision learners. arXiv preprint arXiv:2111.06377, 2021. 1, 4, 8
  40. 40.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In CVPR, 2016. 8
  41. 41.Dan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath, Frank Wang, Evan Dorundo, Rahul Desai, Tyler Zhu, Samyak Parajuli, Mike Guo, et al. The many faces of robustness: A critical analysis of out-of-distribution generalization. In CVPR, 2021. 4, 7, 8
  42. 42.Dan Hendrycks, Kevin Zhao, Steven Basart, Jacob Steinhardt, and Dawn Song. Natural adversarial examples. In CVPR, 2021. 4, 7, 8
  43. 43.Zejiang Hou, Fei Sun, Yen-Kuang Chen, Yuan Xie, and Sun-Yuan Kung. Milan: Masked image pretraining on language assisted representation. arXiv preprint arXiv:2208.06049, 2022. 1, 3, 8
  44. 44.Gao Huang, Zhuang Liu, Laurens Van Der Maaten, and Kilian Q Weinberger. Densely connected convolutional networks. In CVPR, 2017. 8
  45. 45.Gao Huang, Yu Sun, Zhuang Liu, Daniel Sedra, and Kilian Q Weinberger. Deep networks with stochastic depth. In ECCV, 2016. 3
  46. 46.Zhaojin Huang, Lichao Huang, Yongchao Gong, Chang Huang, and Xinggang Wang. Mask scoring r-cnn. In CVPR, 2019. 5
  47. 47.Gabriel Ilharco, Mitchell Wortsman, Ross Wightman, Cade Gordon, Nicholas Carlini, Rohan Taori, Achal Dave, Vaishaal Shankar, Hongseok Namkoong, John Miller, Hannaneh Hajishirzi, Ali Farhadi, and Ludwig Schmidt. Openclip. https://github.com/mlfoundations/open_clip, 2021. 2, 7, 8
  48. 48.Jitesh Jain, Anukriti Singh, Nikita Orlov, Zilong Huang, Jiachen Li, Steven Walton, and Humphrey Shi. Semask: Semantically masked transformers for semantic segmentation. arXiv preprint arXiv:2112.12782, 2021. 6
  49. 49.Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig. Scaling up visual and vision-language representation learning with noisy text supervision. In ICML, 2021. 7
  50. 50.Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. Scaling laws for neural language models. arXiv preprint arXiv:2001.08361, 2020. 1
  51. 51.Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, et al. The kinetics human action video dataset. arXiv preprint arXiv:1705.06950, 2017. 2, 4, 7, 8
  52. 52.Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014. 3
  53. 53.Alexander Kirillov, Yuxin Wu, Kaiming He, and Ross Girshick. Pointrend: Image segmentation as rendering. In CVPR, 2020. 6
  54. 54.Alex Krizhevsky, Geoffrey Hinton, et al. Learning multiple layers of features from tiny images. 2009. 7, 8
  55. 55.Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. Imagenet classification with deep convolutional neural networks. In NeurIPS, 2012. 8
  56. 56.Yann LeCun, Bernhard Boser, John S Denker, Donnie Henderson, Richard E Howard, Wayne Hubbard, and Lawrence D Jackel. Back-propagation applied to handwritten zip code recognition. Neural computation, 1989. 8
  57. 57.Feng Li, Hao Zhang, Shilong Liu, Lei Zhang, Lionel M Ni, Heung-Yeung Shum, et al. Mask dino: Towards a unified transformer-based framework for object detection and segmentation. arXiv preprint arXiv:2206.02777, 2022. 2, 6
  58. 58.Kunchang Li, Yali Wang, Yinan He, Yizhuo Li, Yi Wang, Limin Wang, and Yu Qiao. Uniformerv2: Spatiotemporal learning by arming image vits with video uniformer. arXiv preprint arXiv:2211.09552, 2022. 5
  59. 59.Liunian Harold Li, Pengchuan Zhang, Haotian Zhang, Jianwei Yang, Chunyuan Li, Yiwu Zhong, Lijuan Wang, Lu Yuan, Lei Zhang, Jenq-Neng Hwang, et al. Grounded language-image pre-training. In CVPR, 2022. 6
  60. 60.Yanghao Li, Hanzi Mao, Ross Girshick, and Kaiming He. Exploring plain vision transformer backbones for object detection. arXiv preprint arXiv:2203.16527, 2022. 5, 6
  61. 61.Yanghao Li, Chao-Yuan Wu, Haoqi Fan, Karttikeya Mangalam, Bo Xiong, Jitendra Malik, and Christoph Feichtenhofer. Improved multiscale vision transformers for classification and detection. arXiv preprint arXiv:2112.01526, 2021. 4
  62. 62.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollar, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In European conference on computer vision, pages 740–755. Springer, 2014. 2, 3, 5
  63. 63.Xingbin Liu, Jinghao Zhou, Tao Kong, Xianming Lin, and Rongrong Ji. Exploring target representations for masked autoencoders. arXiv preprint arXiv:2209.03917, 2022. 8
  64. 64.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692, 2019. 1
  65. 65.Ze Liu, Han Hu, Yutong Lin, Zhuliang Yao, Zhenda Xie, Yixuan Wei, Jia Ning, Yue Cao, Zheng Zhang, Li Dong, et al. Swin transformer v2: Scaling up capacity and resolution. In CVPR, 2022. 1, 2, 4, 5, 6, 8
  66. 66.Zhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer, Trevor Darrell, and Saining Xie. A convnet for the 2020s. In CVPR, 2022. 4, 8
  67. 67.Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. In ICLR, 2019. 3
  68. 68.Bolin Ni, Houwen Peng, Minghao Chen, Songyang Zhang, Gaofeng Meng, Jianlong Fu, Shiming Xiang, and Haibin Ling. Expanding language-image pretrained models for general video recognition. In ECCV, 2022. 5
  69. 69.Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al. Pytorch: An imperative style, high-performance deep learning library. NeurIPS, 2019. 3
  70. 70.Zhiliang Peng, Li Dong, Hangbo Bao, Qixiang Ye, and Furu Wei. Beit v2: Masked image modeling with vector-quantized visual tokenizers. arXiv preprint arXiv:2208.06366, 2022. 1, 2, 3, 4, 5, 7, 8
  71. 71.Hieu Pham, Zihang Dai, Golnaz Ghiasi, Kenji Kawaguchi, Hanxiao Liu, Adams Wei Yu, Jiahui Yu, Yi-Ting Chen, Minh-Thang Luong, Yonghui Wu, et al. Combined scaling for open-vocabulary image classification. arXiv preprint arXiv: 2111.10050, 2021. 1
  72. 72.Hieu Pham, Zihang Dai, Golnaz Ghiasi, Hanxiao Liu, Adams Wei Yu, Minh-Thang Luong, Mingxing Tan, and Quoc V Le. Combined scaling for zero-shot transfer learning. arXiv preprint arXiv:2111.10050, 2021. 7
  73. 73.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In ICML, 2021. 1, 2, 3, 7
  74. 74.Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al. Improving language understanding by generative pre-training. 2018. 1
  75. 75.Jack W Rae, Sebastian Borgeaud, Trevor Cai, Katie Millican, Jordan Hoffmann, Francis Song, John Aslanides, Sarah Henderson, Roman Ring, Susannah Young, et al. Scaling language models: Methods, analysis & insights from training gopher. arXiv preprint arXiv:2112.11446, 2021. 1
  76. 76.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, Peter J Liu, et al. Exploring the limits of transfer learning with a unified text-to-text transformer. JMLR, 2020. 1
  77. 77.Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He. Zero: Memory optimizations toward training trillion parameter models. In SC20, 2020. 3, 8
  78. 78.Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. Hierarchical text-conditional image generation with clip latents. arXiv preprint arXiv:2204.06125, 2022. 7
  79. 79.Yongming Rao, Wenliang Zhao, Yansong Tang, Jie Zhou, Ser-Nam Lim, and Jiwen Lu. Hornet: Efficient high-order spatial interactions with recursive gated convolutions. arXiv preprint arXiv:2207.14284, 2022. 6
  80. 80.Jeff Rasley, Samyam Rajbhandari, Olatunji Ruwase, and Yuxiong He. Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters. In KDD, 2020. 3, 8
  81. 81.Benjamin Recht, Rebecca Roelofs, Ludwig Schmidt, and Vaishaal Shankar. Do imagenet classifiers generalize to imagenet? In ICML, 2019. 4, 8
  82. 82.Benjamin Recht, Rebecca Roelofs, Ludwig Schmidt, and Vaishaal Shankar. Do imagenet classifiers generalize to imagenet?, 2019. 7
  83. 83.Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bjorn Ommer. High-resolution image synthesis with latent diffusion models. In CVPR, 2022. 7
  84. 84.Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily Denton, Seyed Kamyar Seyed Ghasemipour, Burcu Karagol Ayan, S Sara Mahdavi, Rapha Gontijo Lopes, et al. Photorealistic text-to-image diffusion models with deep language understanding. arXiv preprint arXiv:2205.11487, 2022. 7
  85. 85.Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Wortsman, et al. Laion-5b: An open large-scale dataset for training next generation image-text models. arXiv preprint arXiv:2210.08402, 2022. 7
  86. 86.Christoph Schuhmann, Richard Vencu, Romain Beaumont, Robert Kaczmarczyk, Clayton Mullis, Aarush Katta, Theo Coombes, Jenia Jitsev, and Aran Komatsuzaki. Laion-400m: Open dataset of clip-filtered 400 million image-text pairs. arXiv preprint arXiv:2111.02114, 2021. 7
  87. 87.Shuai Shao, Zeming Li, Tianyuan Zhang, Chao Peng, Gang Yu, Xiangyu Zhang, Jing Li, and Jian Sun. Objects365: A large-scale, high-quality dataset for object detection. In ICCV, 2019. 3, 5
  88. 88.Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut. Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning. In ACL, 2018. 3
  89. 89.Peter Shaw, Jakob Uszkoreit, and Ashish Vaswani. Self-attention with relative position representations. arXiv preprint arXiv:1803.02155, 2018. 3
  90. 90.Peter Shaw, Jakob Uszkoreit, and Ashish Vaswani. Self-attention with relative position representations. arXiv preprint arXiv:1803.02155, 2018. 7
  91. 91.Karen Simonyan and Andrew Zisserman. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556, 2014. 8
  92. 92.Khurram Soomro, Amir Roshan Zamir, and Mubarak Shah. Ucf101: A dataset of 101 human actions classes from videos in the wild. arXiv preprint arXiv:1212.0402, 2012. 7, 8
  93. 93.Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich. Going deeper with convolutions. In CVPR, 2015. 8
  94. 94.Mingxing Tan and Quoc Le. Efficientnet: Rethinking model scaling for convolutional neural networks. In ICML, 2019. 8
  95. 95.Zhan Tong, Yibing Song, Jue Wang, and Limin Wang. Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training. arXiv preprint arXiv:2203.12602, 2022. 5
  96. 96.Hugo Touvron, Matthieu Cord, and Herve Jégou. DeiT iii: Revenge of the vit. arXiv preprint arXiv:2204.07118, 2022. 4
  97. 97.Hugo Touvron, Matthieu Cord, Alexandre Sablayrolles, Gabriel Synnaeve, and Herve Jégou. Going deeper with image transformers. In ICCV, 2021. 3
  98. 98.Zhengzhong Tu, Hossein Talebi, Han Zhang, Feng Yang, Peyman Milanfar, Alan Bovik, and Yinxiao Li. Maxvit: Multi-axis vision transformer. arXiv preprint arXiv:2204.01697, 2022. 4
  99. 99.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. NeurIPS, 30, 2017. 1
  100. 100.Haohan Wang, Songwei Ge, Zachary Lipton, and Eric P Xing. Learning robust global representations by penalizing local predictive power. NeurIPS, 2019. 4, 7, 8
  101. 101.Wenhui Wang, Hangbo Bao, Li Dong, Johan Bjorck, Zhiliang Peng, Qiang Liu, Kriti Aggarwal, Owais Khan Mohammed, Saksham Singhal, Subhojit Som, et al. Image as a foreign language: Beit pretraining for all vision and vision-language tasks. arXiv preprint arXiv:2208.10442, 2022. 1, 2, 3, 4, 5, 6, 7, 8
  102. 102.Xinlong Wang, Rufeng Zhang, Chunhua Shen, Tao Kong, and Lei Li. Solo: A simple framework for instance segmentation. TPAMI, 2021. 5
  103. 103.Chen Wei, Haoqi Fan, Saining Xie, Chao-Yuan Wu, Alan Yuille, and Christoph Feichtenhofer. Masked feature prediction for self-supervised visual pre-training. In CVPR, 2022. 2, 5, 8
  104. 104.Jason Wei, Maarten Bosma, Vincent Y Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M Dai, and Quoc V Le. Finetuned language models are zero-shot learners. arXiv preprint arXiv:2109.01652, 2021. 1
  105. 105.Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, et al. Emergent abilities of large language models. arXiv preprint arXiv:2206.07682, 2022. 1, 2
  106. 106.Longhui Wei, Lingxi Xie, Wengang Zhou, Houqiang Li, and Qi Tian. Mvp: Multimodality-guided visual pre-training. arXiv preprint arXiv:2203.05175, 2022. 1, 3, 8
  107. 107.Yixuan Wei, Han Hu, Zhenda Xie, Zheng Zhang, Yue Cao, Jianmin Bao, Dong Chen, and Baining Guo. Contrastive learning rivals masked image modeling in fine-tuning via feature distillation. arXiv preprint arXiv:2205.14141, 2022. 2, 3, 4, 5, 6
  108. 108.Ross Wightman. Pytorch image models. https://github.com/rwightman/pytorch-image-models, 2019. 4
  109. 109.Wenhao Wu, Zhun Sun, and Wanli Ouyang. Transferring textual knowledge for visual recognition. arXiv preprint arXiv:2207.01297, 2022. 2
  110. 110.Qizhe Xie, Minh-Thang Luong, Eduard Hovy, and Quoc V Le. Self-training with noisy student improves imagenet classification. In CVPR, 2020. 4
  111. 111.Saining Xie, Ross Girshick, Piotr Dollar, Zhuowen Tu, and Kaiming He. Aggregated residual transformations for deep neural networks. In CVPR, 2017. 8
  112. 112.Zhenda Xie, Zigang Geng, Jingcheng Hu, Zheng Zhang, Han Hu, and Yue Cao. Revealing the dark secrets of masked image modeling. arXiv preprint arXiv:2205.13543, 2022. 1
  113. 113.Zhenda Xie, Zheng Zhang, Yue Cao, Yutong Lin, Jianmin Bao, Zhuliang Yao, Qi Dai, and Han Hu. Simmim: A simple framework for masked image modeling. arXiv preprint arXiv:2111.09886, 2021. 1, 8
  114. 114.Zhenda Xie, Zheng Zhang, Yue Cao, Yutong Lin, Yixuan Wei, Qi Dai, and Han Hu. On data scaling in masked image modeling. arXiv preprint arXiv:2206.04664, 2022. 1
  115. 115.Mengde Xu, Zheng Zhang, Han Hu, Jianfeng Wang, Lijuan Wang, Fangyun Wei, Xiang Bai, and Zicheng Liu. End-to-end semi-supervised object detection with soft teacher. In ICCV, 2021. 6
  116. 116.Jianwei Yang, Chunyuan Li, and Jianfeng Gao. Focal modulation networks. arXiv preprint arXiv:2203.11926, 2022. 2, 5, 6
  117. 117.Jiahui Yu, Zirui Wang, Vijay Vasudevan, Legg Yeung, Mojtaba Seyedhosseini, and Yonghui Wu. Coca: Contrastive captioners are image-text foundation models. arXiv preprint arXiv:2205.01917, 2022. 4, 5
  118. 118.Lu Yuan, Dongdong Chen, Yi-Ling Chen, Noel Codella, Xiyang Dai, Jianfeng Gao, Houdong Hu, Xuedong Huang, Boxin Li, Chunyuan Li, et al. Florence: A new foundation model for computer vision. arXiv preprint arXiv:2111.11432, 2021. 5, 6
  119. 119.Xiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, and Lucas Beyer. Scaling vision transformers. In CVPR, 2022. 1, 2, 3, 4, 8
  120. 120.Bowen Zhang, Jiahui Yu, Christopher Fifty, Wei Han, Andrew M Dai, Ruoming Pang, and Fei Sha. Co-training transformer with videos and images improves action recognition. arXiv preprint arXiv:2112.07175, 2021. 5
  121. 121.Hao Zhang, Feng Li, Shilong Liu, Lei Zhang, Hang Su, Jun Zhu, Lionel M Ni, and Heung-Yeung Shum. Dino: Detr with improved denoising anchor boxes for end-to-end object detection. arXiv preprint arXiv:2203.03605, 2022. 5, 6
  122. 122.Haotian Zhang, Pengchuan Zhang, Xiaowei Hu, Yen-Chun Chen, Liunian Harold Li, Xiyang Dai, Lijuan Wang, Lu Yuan, Jenq-Neng Hwang, and Jianfeng Gao. Glipv2: Unifying localization and vision-language understanding. arXiv preprint arXiv:2206.05836, 2022. 6
  123. 123.Bolei Zhou, Hang Zhao, Xavier Puig, Tete Xiao, Sanja Fidler, Adela Barriuso, and Antonio Torralba. Semantic understanding of scenes through the ade20k dataset. IJCV, 2018. 2, 3, 6
  124. 124.Jinghao Zhou, Chen Wei, Huiyu Wang, Wei Shen, Cihang Xie, Alan Yuille, and Tao Kong. ibot: Image bert pre-training with online tokenizer. arXiv preprint arXiv:2111.07832, 2021. 1, 2, 8
  125. 125.Xingyi Zhou, Vladlen Koltun, and Philipp Krahenbühl. Probabilistic two-stage detection. arXiv preprint arXiv:2103.07461, 2021. 5

Citation

MLA
Fang, Y., et al. “EVA: Exploring the Limits of Masked Visual Representation Learning at Scale”. arXiv, 2022, http://arxiv.org/abs/2211.07636v2.
APA
Fang, Y., Wang, W., Xie, B., Sun, Q., Wu, L., Wang, X., Huang, T., Wang, X., & Cao, Y. (2022). EVA: Exploring the Limits of Masked Visual Representation Learning at Scale. arXiv. http://arxiv.org/abs/2211.07636v2
Chicago
Fang, Y., W. Wang, B. Xie, et al. 2022. “EVA: Exploring the Limits of Masked Visual Representation Learning at Scale”. arXiv. http://arxiv.org/abs/2211.07636v2.
Harvard
Fang, Y. et al. (2022) “EVA: Exploring the Limits of Masked Visual Representation Learning at Scale”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2211.07636v2.
Vancouver
1. Fang Y, Wang W, Xie B, Sun Q, Wu L, Wang X, Huang T, Wang X, Cao Y (2022) EVA: Exploring the Limits of Masked Visual Representation Learning at Scale. arXiv

BibTeX

@article{fang2022eva,
  title = {EVA: Exploring the Limits of Masked Visual Representation Learning at Scale},
  author = {Fang, Yuxin and Wang, Wen and Xie, Binhui and Sun, Quan and Wu, Ledell and Wang, Xinggang and Huang, Tiejun and Wang, Xinlong and Cao, Yue},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2211.07636v2},
  eprint = {2211.07636}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE