Masked Feature Prediction for Self-Supervised Visual Pre-Training

Chen WeiHaoqi FanSaining XieChao-Yuan WuAlan L. YuilleChristoph Feichtenhofer

article2022CVPR895 citations

Demonstrates that regressing simple Histograms of Oriented Gradients for masked video patches provides an efficient, highly effective self-supervised pre-training objective that sets state-of-the-art performance across major video action recognition benchmarks without relying on external tokenizers or extra supervision.

Listen

Modern computer vision models built on Transformer architectures require massive amounts of training data. Unlike natural language processing models, which excel by predicting masked words from unlabelled text, visual models have historically depended on expensive, manually labeled image datasets or complex multi-stage pipelines to avoid severe performance issues. In video recognition, this dependency is especially problematic because the gap between models trained on labeled datasets versus models trained from scratch on raw video has exceeded 5% in accuracy.

The article introduces and evaluates Masked Feature Prediction (MaskFeat), a self-supervised pre-training method designed directly for unlabeled video and image data. Its objective is to demonstrate that directly predicting simple visual feature descriptors in masked regions enables visual Transformers to achieve superior recognition performance without external labels or complex tokenizer networks.

MaskFeat operates by dividing a video or image into small space-time blocks, randomly concealing approximately 40% of them, and training a single neural network to predict feature representations of the hidden regions using only the visible surrounding content. The authors systematically evaluated five distinct prediction targets: raw pixel colors, Histograms of Oriented Gradients (HOG—a classical hand-crafted edge and texture descriptor), discrete visual tokens requiring extra pre-trained neural networks, continuous deep network features, and pseudo-class labels. These variants were tested across standard benchmark datasets, including Kinetics-400, Kinetics-600, AVA action detection, Something-Something v2, and ImageNet-1K.

The investigation produced several key findings. First, predicting standard HOG descriptors proved to be the most effective strategy, balancing high performance with computational efficiency while avoiding the need for secondary teacher networks or discrete visual vocabularies. Second, on the Kinetics-400 video benchmark, a MaskFeat-trained model achieved 86.7% top-1 accuracy using no external data, surpassing the best previous scratch-trained baseline by 5.2 percentage points and matching models trained on hundreds of millions of labeled images. Third, the pre-trained models transferred exceptionally well to fine-grained downstream tasks, reaching state-of-the-art results of 38.8 mAP on action localization in AVA and 75.0% accuracy on human-object interaction in Something-Something v2, outperforming models supervised on large image sets. Finally, experiments on single images showed that ViT models trained with MaskFeat on ImageNet-1K attained 85.7% accuracy, outperforming supervised models trained on datasets ten times larger.

These findings demonstrate that pre-training directly on raw, unlabeled video effectively replaces the costly practice of pre-training on massive supervised image collections. By using standard gradient descriptors with local contrast normalization, the model learns essential structural and motion patterns while ignoring irrelevant color and illumination variations. This significantly reduces data annotation costs, pipeline complexity, and the training overhead associated with complex multi-view contrastive setups or multi-stage tokenizers.

Organizations developing video and image understanding systems should adopt direct masked feature prediction with HOG descriptors to reduce reliance on external labeled data. Teams should prioritize spatiotemporal cube masking over single-frame masking when processing video. Further work should explore deploying this framework across longer video sequences, domain-specific industry video archives, and real-time inference constraints. Given the robust empirical results across multiple public benchmarks and diverse network scales, confidence in MaskFeat's core capability to streamline visual pre-training is high.

  • Paper: Masked Autoencoders Are Scalable Vision Learners, Kaiming He et al. (2022). Read this image-based masked-autoencoder foundation first to understand the masking-and-prediction framework MaskFeat adapts for video.
  • Paper: SimMIM: a Simple Framework for Masked Image Modeling, Zhenda Xie et al. (2021). SimMIM establishes direct prediction of masked visual content as a simple alternative to tokenizers, framing MaskFeat’s comparison of prediction targets.
  • Paper: BEiT: BERT Pre-Training of Image Transformers, Hangbo Bao et al. (2022). BEiT shows how masked-language modeling can transfer to vision through discrete visual tokens, one of the target types MaskFeat evaluates against HOG.
  • Paper: ViViT: A Video Vision Transformer, Anurag Arnab et al. (2021). ViViT introduces the video-transformer architectures whose spatiotemporal tokens provide essential context for MaskFeat’s video pre-training design.
Cover for Masked Feature Prediction for Self-Supervised Visual Pre-Training

Abstract

We present Masked Feature Prediction (MaskFeat) for self-supervised pre-training of video models. Our approach first randomly masks out a portion of the input sequence and then predicts the feature of the masked regions. We study five different types of features and find Histograms of Oriented Gradients (HOG), a hand-crafted feature descriptor, works particularly well in terms of both performance and efficiency. We observe that the local contrast normalization in HOG is essential for good results, which is in line with earlier work using HOG for visual recognition. Our approach can learn abundant visual knowledge and drive large-scale Transformer-based models. Without using extra model weights or supervision, MaskFeat pre-trained on unlabeled videos achieves unprecedented results of 86.7% with MViTv2-L on Kinetics-400, 88.3% on Kinetics-600, 80.4% on Kinetics-700, 38.8 mAP on AVA, and 75.0% on SSv2. MaskFeat further generalizes to image input, which can be interpreted as a video with a single frame and obtains competitive results on ImageNet.

Table of Contents

  • 1 Introduction
  • 2 Method
  • 2.1 Masked Feature Prediction
  • 2.2 Target Features
  • 3 Study: Target Features for MaskFeat
  • 4 Experiments: Video Recognition
  • 4.1 Main Results on Kinetics
  • 4.2 Transfer Learning
  • 4.3 Ablations for Video Recognition
  • 5 Experiments: Image Recognition
  • 5.1 Main Results on ImageNet-1K
  • 5.2 Ablations for Image Recognition
  • 6 Related Work
  • 7 Conclusion
  • A Ablations on Image Classification
  • B Implementation Details
  • B.1 ImageNet and Kinetics Experiments
  • B.2 AVA Experiments
  • B.3 SSv2 Experiments
  • C Qualitative Experiments
  • D Acknowledgement
  • E Changes in arXiv v2
  • References

Knowls

  1. Knowl 1 — Masked Feature Prediction trains on features of hidden visual regions

    model/method

    Masked Feature Prediction (MaskFeat) pre-trains a vision Transformer by hiding parts of an image or video and predicting features computed from the intact input. For video, the input is divided into space-time cubes and projected into tokens; for an image, tokens represent spatial patches. A random subset of tokens is replaced by a learnable mask embedding, positional embeddings are added, and the sequence is processed by the Transformer. A linear head predicts the target feature only for masked tokens, and the loss is applied only to those predictions. For a masked video cube, the target is the feature of the spatial patch at the cube’s temporal center. The target feature is extracted from the original, unmasked sample. The paper’s standard video and image experiments mask 40% of the tokens; the method supports different target features, including pixel values, HOG descriptors, discrete tokens, and deep-network features.

  2. Knowl 2 — Dense HOG descriptors are an effective, low-overhead prediction target

    definition

    A Histogram of Oriented Gradients (HOG) descriptor summarizes local edge structure by computing pixel gradients, accumulating gradient magnitudes into orientation bins within spatial cells, and normalizing each cell’s histogram. For MaskFeat, HOG is computed densely over the whole image before the feature map is split into patches; this reduces boundary padding compared with computing descriptors independently on masked patches. The histograms for a masked patch are flattened and concatenated as its prediction target. The default implementation computes gradients separately in the three RGB channels and concatenates their histograms. Unlike a pixel target, HOG summarizes local shape and appearance and is partially invariant to photometric changes through gradient computation and local contrast normalization. Its computation adds negligible overhead and requires no external teacher model.

  3. Knowl 3 — Target-feature comparisons favor HOG among inexpensive targets

    empirical result

    In the video comparison, MViTv2-S, 16×4 was pre-trained for 300 epochs on unlabeled Kinetics-400 (K400) and evaluated by fine-tuning top-1 accuracy. The from-scratch model scored 81.1%; MaskFeat scored 80.7% with RGB targets, 82.2% with HOG, 81.7% with dVAE tokens, 82.5% with DINO features, and 81.9% with supervised MViT features. In the image comparison, ViT-B was pre-trained for 300 epochs on ImageNet-1K (IN-1K), with 100-epoch fine-tuning: scratch scored 81.8%, RGB 82.5%, HOG 83.6%, dVAE tokens 82.8%, unsupervised features 83.6% (MoCo v2), 83.9% (MoCo v3), and 84.0% (DINO), supervised features 82.6% (ResNet-50) or 81.9% (ViT-B), and pseudo-labels 78.8%. Thus HOG improves over scratch by 1.1 percentage points on K400 and 1.8 points on IN-1K without requiring a teacher. Some unsupervised teacher features score higher, but require separate teacher pre-training and feature generation. The authors hypothesize that supervised features and pseudo-labels are less suitable because class-level targets may discard local shapes and textures needed for masked-region modeling.

  4. Knowl 4 — Unlabeled-video MaskFeat scales to strong Kinetics recognition

    empirical result

    MaskFeat pre-training on unlabeled Kinetics videos improves MViTv2 models after supervised fine-tuning for action recognition. On K400, MViTv2-S, 16×4 reached 82.2% top-1 after 300 pre-training epochs, compared with 81.1% from scratch. MViTv2-L, 16×4 reached 84.3% after 800 K400 pre-training epochs, versus 80.5% from scratch and 83.5% with supervised ImageNet-21K pre-training. Pre-training the large model for 300 epochs on Kinetics-600 (K600), which has about 387,000 training videos, raised its K400 fine-tuning result to 85.1%. Larger-input MViTv2-L variants reached 86.7% top-1 on K400 and 87.0% when pre-trained on K600; these results use 40-frame, stride-3 inputs at 352-pixel resolution. The 86.7% K400 result is 5.2 points above the cited best prior result without external data, 81.5%.

  5. Knowl 5 — Video pre-training transfers to action detection and interaction recognition

    empirical result

    Kinetics-pre-trained MViTv2-L, 40×3 models transfer to both spatiotemporal action detection on AVA v2.2 and human-object interaction classification on Something-Something v2 (SSv2). On AVA, the model pre-trained with MaskFeat on K400 obtained 36.3 mAP with center-crop testing and 37.5 mAP at full resolution; K600 pre-training raised these to 37.8 and 38.8 mAP, respectively. The supervised ImageNet-21K-plus-K400 counterpart scored 31.6 mAP with center-crop testing. On SSv2, top-1 accuracy was 73.3% for the supervised counterpart, 74.4% after MaskFeat pre-training on K400, and 75.0% after MaskFeat pre-training on K600. These results show transfer gains from unlabeled Kinetics pre-training on tasks requiring localization or fine-grained motion and interaction recognition.

  6. Knowl 6 — HOG MaskFeat transfers to ImageNet classification without a teacher

    empirical result

    For image recognition, ViT-B and ViT-L were pre-trained for 1,600 epochs at 224×224 resolution on unlabeled IN-1K images using HOG targets, then fine-tuned end to end. Their top-1 accuracies were 84.0% and 85.7%, respectively, compared with scratch results of 81.8% and 81.5%. The ViT-B result matches the reported 84.0% from supervised IN-21K pre-training at 384×384, and the ViT-L result exceeds the corresponding 85.2% supervised result by 0.5 percentage points. MaskFeat uses no external teacher or labeled pre-training data for these results. Fine-tuning schedules were 100 epochs for ViT-B and 50 for ViT-L.

  7. Knowl 7 — Cube masking is the strongest tested video masking strategy

    empirical result

    With MViTv2-S, 16×4 pre-trained for 300 epochs and fine-tuned for 200 epochs on K400, the tested strategies all masked 40% of tokens. Independently masking consecutive frames (“frame” masking) produced 81.0% top-1 accuracy; repeating a 2-D spatial mask across time (“tube” masking) produced 81.9%; and masking spatiotemporal blocks (“cube” masking) produced 82.2%. Cube masks are formed from a spatial block sampled at a time step and extended over a random number of consecutive frames. The result supports using masks that vary across both spatial and temporal dimensions rather than relying only on one.

  8. Knowl 8 — Local contrast normalization is essential to HOG-target pre-training

    empirical result

    In HOG implementation ablations, ViT-B was pre-trained for 300 epochs and fine-tuned on IN-1K; the reported metric is top-1 accuracy. With the default local L2 normalization, accuracy was 83.6%; using L1 normalization gave 82.8%, and omitting normalization gave 82.2%. The default RGB-channel HOG scored 83.6%, compared with 83.2% for grayscale HOG and 83.5% for opponent-color-space HOG. Using 6, 9, or 12 orientation bins gave 83.4%, 83.6%, and 83.5%, respectively. Cell sizes of 4×4, 8×8, and 16×16 pixels gave 83.2%, 83.6%, and 83.2%. These results make local normalization the most consequential tested choice, while also favoring RGB channels and 8×8 cells.

  9. Knowl 9 — HOG predictions are less sensitive than pixel predictions to ambiguous content

    empirical result

    The paper’s qualitative image predictions illustrate two ambiguities in pixel regression that HOG targets can reduce. For a masked balloon, the model predicted a plausible red color even though the original balloon was black, creating a large pixel-wise error despite a reasonable guess. In a texture-rich sea-urchin region, pixel prediction was blurry, whereas HOG captured prominent edge directions. The authors attribute these differences to HOG’s local gradient normalization, which reduces sensitivity to color or illumination ambiguity, and spatial histogramming, which summarizes texture and edge structure instead of requiring exact pixel matches.

Coverage note — No substantial contributed material was omitted; secondary benchmark cost columns and appendix implementation details were not included because they do not change the principal method or findings summarized here.

References

  1. 1.Anurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun, Mario Lucić, and Cordelia Schmid. ViViT: A video vision transformer. In ICCV, 2021. 2, 5
  2. 2.Hangbo Bao, Li Dong, and Furu Wei. BEiT: BERT pre-training of image transformers. arXiv preprint arXiv:2106.08254, 2021. 1, 2, 3, 4, 5, 7, 8
  3. 3.David Bau, Bolei Zhou, Aditya Khosla, Aude Oliva, and Antonio Torralba. Network dissection: Quantifying interpretability of deep visual representations. In CVPR, 2017. 5
  4. 4.Gedas Bertasius, Heng Wang, and Lorenzo Torresani. Is space-time attention all you need for video understanding? In ICML, 2021. 5
  5. 5.Andrew Brock, Soham De, Samuel L Smith, and Karen Simonyan. High-performance large-scale image recognition without normalization. In ICML, 2021. 4
  6. 6.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. Language models are few-shot learners. In NeurIPS, 2020. 1
  7. 7.Mathilde Caron, Piotr Bojanowski, Armand Joulin, and Matthijs Douze. Deep clustering for unsupervised learning of visual features. In ECCV, 2018. 8
  8. 8.Mathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal, Piotr Bojanowski, and Armand Joulin. Unsupervised learning of visual features by contrasting cluster assignments. In NeurIPS, 2020. 8
  9. 9.Mathilde Caron, Hugo Touvron, Ishan Misra, Herve Jégou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. Emerging properties in self-supervised vision transformers. In ICCV, 2021. 3, 4, 7, 8
  10. 10.Joao Carreira, Eric Noland, Andras Banki-Horvath, Chloe Hillier, and Andrew Zisserman. A short note about kinetics-600. arXiv preprint arXiv:1808.01340, 2018. 5
  11. 11.Joao Carreira and Andrew Zisserman. Quo vadis, action recognition? a new model and the kinetics dataset. In CVPR, 2017. 5
  12. 12.Ken Chatfield, Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman. Return of the devil in the details: Delving deep into convolutional nets. In BMVC, 2014. 2, 3, 7
  13. 13.Mark Chen, Alec Radford, Rewon Child, Jeff Wu, Heewoo Jun, Prafulla Dhariwal, David Luan, and Ilya Sutskever. Generative pretraining from pixels. In ICML, 2020. 8
  14. 14.Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. A simple framework for contrastive learning of visual representations. In ICML, 2020. 8
  15. 15.Xinlei Chen, Haoqi Fan, Ross Girshick, and Kaiming He. Improved baselines with momentum contrastive learning. arXiv preprint arXiv:2003.04297, 2020. 4
  16. 16.Xinlei Chen and Kaiming He. Exploring simple siamese representation learning. In CVPR, 2021. 2, 8
  17. 17.Xinlei Chen, Saining Xie, and Kaiming He. An empirical study of training self-supervised vision transformers. In ICCV, 2021. 4, 7
  18. 18.Navneet Dalal and Bill Triggs. Histograms of oriented gradients for human detection. In CVPR, 2005. 2, 3, 4, 7, 8
  19. 19.Navneet Dalal, Bill Triggs, and Cordelia Schmid. Human detection using oriented histograms of flow and appearance. In ECCV, 2006. 3
  20. 20.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. ImageNet: A large-scale hierarchical image database. In CVPR, 2009. 2, 4, 5, 7
  21. 21.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT: Pre-training of deep bidirectional transformers for language understanding. In NAACL, 2019. 1, 2, 3, 5, 8
  22. 22.Carl Doersch, Abhinav Gupta, and Alexei A Efros. Unsupervised visual representation learning by context prediction. In ICCV, 2015. 8
  23. 23.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition at scale. In ICLR, 2021. 1, 3, 4, 7, 8
  24. 24.Alexey Dosovitskiy, Philipp Fischer, Jost Tobias Springenberg, Martin Riedmiller, and Thomas Brox. Discriminative unsupervised feature learning with exemplar convolutional neural networks. TPAMI, 2015. 8
  25. 25.Haoqi Fan, Tullie Murrell, Heng Wang, Kalyan Vasudev Alwala, Yanghao Li, Yilei Li, Bo Xiong, Nikhila Ravi, Meng Li, Haichuan Yang, Jitendra Malik, Ross Girshick, Matt Feiszli, Aaron Adcock, Wan-Yen Lo, and Christoph Feichtenhofer. PyTorchVideo: A deep learning library for video understanding. In ACM MM, 2021. https://pytorchvideo.org/. 2
  26. 26.Haoqi Fan, Bo Xiong, Karttikeya Mangalam, Yanghao Li, Zhicheng Yan, Jitendra Malik, and Christoph Feichtenhofer. Multiscale vision transformers. In ICCV, 2021. 2, 4, 5, 6
  27. 27.Christoph Feichtenhofer. X3D: Expanding architectures for efficient video recognition. In CVPR, 2020. 5, 6
  28. 28.Christoph Feichtenhofer, Haoqi Fan, Jitendra Malik, and Kaiming He. Slowfast networks for video recognition. In ICCV, 2019. 5, 6
  29. 29.Christoph Feichtenhofer, Haoqi Fan, Bo Xiong, Ross Girshick, and Kaiming He. A large-scale study on unsupervised spatiotemporal representation learning. In CVPR, 2021. 2, 8
  30. 30.Basura Fernando, Hakan Bilen, Efstratios Gavves, and Stephen Gould. Self-supervised video representation learning with odd-one-out networks. In CVPR, 2017. 8
  31. 31.Spyros Gidaris, Praveer Singh, and Nikos Komodakis. Unsupervised representation learning by predicting image rotations. In ICLR, 2018. 8
  32. 32.Ross Goroshin, Joan Bruna, Jonathan Tompson, David Eigen, and Yann LeCun. Unsupervised learning of spatiotemporally coherent metrics. In ICCV, 2015. 8
  33. 33.Raghav Goyal, Samira Ebrahimi Kahou, Vincent Michalski, Joanna Materzynska, Susanne Westphal, Heuna Kim, Valentin Haenel, Ingo Fruend, Peter Yianilos, Moritz Mueller-Freitag, et al. The “something something” video database for learning and evaluating visual common sense. In ICCV, 2017. 2, 6
  34. 34.Jean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec, Pierre H. Richemond, Elena Buchatskaya, Carl Doersch, Bernardo Avila Pires, Zhaohan Daniel Guo, Mohammad Gheshlaghi Azar, Bilal Piot, Koray Kavukcuoglu, Rémi Munos, and Michal Valko. Bootstrap your own latent: A new approach to self-supervised learning. In NeruIPS, 2020. 8
  35. 35.Chunhui Gu, Chen Sun, Sudheendra Vijayanarasimhan, Caroline Pantofaru, David A. Ross, George Toderici, Yeqing Li, Susanna Ricco, Rahul Sukthankar, Cordelia Schmid, and Jitendra Malik. AVA: A video dataset of spatio-temporally localized atomic visual actions. In CVPR, 2018. 2, 6
  36. 36.Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollar, and Ross Girshick. Masked autoencoders are scalable vision learners. arXiv preprint arXiv:2111.06377, 2021. 8
  37. 37.Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick. Momentum contrast for unsupervised visual representation learning. In CVPR, 2020. 2, 8
  38. 38.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In CVPR, 2016. 4
  39. 39.Roei Herzig, Elad Ben-Avraham, Karttikeya Mangalam, Amir Bar, Gal Chechik, Anna Rohrbach, Trevor Darrell, and Amir Globerson. Object-region video transformers. arXiv preprint arXiv:2110.06915, 2021. 6
  40. 40.Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. Distilling the knowledge in a neural network. In NeurIPS, 2015. 3
  41. 41.Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. Distilling the knowledge in a neural network. arXiv preprint arXiv:1503.02531, 2015. 5
  42. 42.Zihang Jiang, Qibin Hou, Li Yuan, Daquan Zhou, Yujun Shi, Xiaojie Jin, Anran Wang, and Jiashi Feng. All tokens matter: Token labeling for training better vision transformers. In NeurIPS, 2021. 4
  43. 43.Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, et al. The kinetics human action video dataset. arXiv preprint arXiv:1705.06950, 2017. 2, 4
  44. 44.Dan Kondratyuk, Liangzhe Yuan, Yandong Li, Li Zhang, Mingxing Tan, Matthew Brown, and Boqing Gong. MoviNets: Mobile video networks for efficient video recognition. In CVPR, 2021. 5
  45. 45.Christian Ledig, Lucas Theis, Ferenc Huszar, Jose Caballero, Andrew Cunningham, Alejandro Acosta, Andrew Aitken, Alykhan Tejani, Johannes Totz, Zehan Wang, et al. Photorealistic single image super-resolution using a generative adversarial network. In CVPR, 2017. 8
  46. 46.Yanghao Li, Chao-Yuan Wu, Haoqi Fan, Karttikeya Mangalam, Bo Xiong, Jitendra Malik, and Christoph Feichtenhofer. Improved multiscale vision transformers for classification and detection. arXiv preprint arXiv:2112.01526, 2021. 1, 2, 4, 5, 6
  47. 47.Ze Liu, Han Hu, Yutong Lin, Zhuliang Yao, Zhenda Xie, Yixuan Wei, Jia Ning, Yue Cao, Zheng Zhang, Li Dong, Furu Wei, and Baining Guo. Swin transformer v2: Scaling up capacity and resolution. arXiv preprint arXiv:2111.09883, 2021. 5, 6
  48. 48.Ze Liu, Jia Ning, Yue Cao, Yixuan Wei, Zheng Zhang, Stephen Lin, and Han Hu. Video swin transformer. arXiv preprint arXiv:2106.13230, 2021. 5, 6
  49. 49.David G Lowe. Object recognition from local scale-invariant features. In ICCV, 1999. 2, 3, 8
  50. 50.Ishan Misra, C Lawrence Zitnick, and Martial Hebert. Shuffle and learn: unsupervised learning using temporal order verification. In ECCV, 2016. 8
  51. 51.Mehdi Noroozi and Paolo Favaro. Unsupervised learning of visual representations by solving jigsaw puzzles. In ECCV, 2016. 8
  52. 52.Junting Pan, Siyu Chen, Mike Zheng Shou, Yu Liu, Jing Shao, and Hongsheng Li. Actor-context-actor relation network for spatio-temporal action localization. In CVPR, 2021. 6
  53. 53.Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. PyTorch: An imperative style, high-performance deep learning library. In NeurIPS, 2019. 4
  54. 54.Deepak Pathak, Ross Girshick, Piotr Dollar, Trevor Darrell, and Bharath Hariharan. Learning features by watching objects move. In CVPR, 2017. 8
  55. 55.Deepak Pathak, Philipp Krahenbuhl, Jeff Donahue, Trevor Darrell, and Alexei A Efros. Context encoders: Feature learning by inpainting. In CVPR, 2016. 3, 8
  56. 56.Mandela Patrick, Dylan Campbell, Yuki M. Asano, Ishan Misra Florian Metze, Christoph Feichtenhofer, Andrea Vedaldi, and Joao F. Henriques. Keeping your eye on the ball: Trajectory attention in video transformers. In NeurIPS, 2021. 6
  57. 57.Rui Qian, Tianjian Meng, Boqing Gong, Ming-Hsuan Yang, Huisheng Wang, Serge Belongie, and Yin Cui. Spatiotemporal contrastive video representation learning. In CVPR, 2021. 8
  58. 58.Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever. Zero-shot text-to-image generation. In ICML, 2021. 1, 2, 3, 4, 8
  59. 59.Michael S Ryoo, AJ Piergiovanni, Anurag Arnab, Mostafa Dehghani, and Anelia Angelova. Tokenlearner: What can 8 learned tokens do for images and videos? arXiv preprint arXiv:2106.11297, 2021. 5
  60. 60.Tim Salimans, Ian Goodfellow, Wojciech Zaremba, Vicki Cheung, Alec Radford, and Xi Chen. Improved techniques for training GANs. In NeurIPS, 2016. 3
  61. 61.Chen Sun, Abhinav Shrivastava, Saurabh Singh, and Abhinav Gupta. Revisiting unreasonable effectiveness of data in deep learning era. In ICCV, 2017. 2, 5
  62. 62.Hao Tan, Jie Lei, Thomas Wolf, and Mohit Bansal. VIMPAC: Video pre-training via masked token prediction and contrastive learning. arXiv preprint arXiv:2106.11250, 2021. 2, 8
  63. 63.Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Herve Jegou. Training data-efficient image transformers & distillation through attention. In ICML, 2021. 4, 5, 7
  64. 64.Koen Van De Sande, Theo Gevers, and Cees Snoek. Evaluating color descriptors for object and scene recognition. TPAMI, 2009. 7, 8
  65. 65.Aaron van den Oord, Oriol Vinyals, and Koray Kavukcuoglu. Neural discrete representation learning. In NeurIPS, 2017. 8
  66. 66.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In NeurIPS, 2017. 1
  67. 67.Pascal Vincent, Hugo Larochelle, Isabelle Lajoie, Yoshua Bengio, Pierre-Antoine Manzagol, and Leon Bottou. Stacked denoising autoencoders: Learning useful representations in a deep network with a local denoising criterion. JMLR, 2010. 8
  68. 68.Xiaolong Wang, Kaiming He, and Abhinav Gupta. Transitive invariance for self-supervised visual representation learning. In ICCV, 2017. 8
  69. 69.Chen Wei, Lingxi Xie, Xutong Ren, Yingda Xia, Chi Su, Jiaying Liu, Qi Tian, and Alan L Yuille. Iterative reorganization with weak spatial constraints: Solving arbitrary jigsaw puzzles for unsupervised representation learning. In CVPR, 2019. 8
  70. 70.Chao-Yuan Wu and Philipp Krahenbuhl. Towards long-form video understanding. In CVPR, 2021. 6
  71. 71.Zhirong Wu, Yuanjun Xiong, X Yu Stella, and Dahua Lin. Unsupervised feature learning via non-parametric instance discrimination. In CVPR, 2018. 8
  72. 72.Lu Yuan, Dongdong Chen, Yi-Ling Chen, Noel Codella, Xiyang Dai, Jianfeng Gao, Houdong Hu, Xuedong Huang, Boxin Li, Chunyuan Li, Ce Liu, Mengchen Liu, Zicheng Liu, Yumao Lu, Yu Shi, Lijuan Wang, Jianfeng Wang, Bin Xiao, Zhen Xiao, Jianwei Yang, Michael Zeng, Luowei Zhou, and Pengchuan Zhang. Florence: A new foundation model for computer vision. arXiv preprint arXiv:2111.11432, 2021. 5, 6
  73. 73.Richard Zhang, Phillip Isola, and Alexei A Efros. Colorful image colorization. In ECCV, 2016. 8
  74. 74.Nanxuan Zhao, Zhirong Wu, Rynson W. H. Lau, and Stephen Lin. What makes instance discrimination good for transfer learning? In ICLR, 2021. 3
  75. 75.Bolei Zhou, Aditya Khosla, Agata Lapedriza, Aude Oliva, and Antonio Torralba. Object detectors emerge in deep scene cnns. In ICLR, 2015. 5

Citation

MLA
Wei, C., et al. “Masked Feature Prediction for Self-Supervised Visual Pre-Training”. arXiv, 2021, http://arxiv.org/abs/2112.09133v2.
APA
Wei, C., Fan, H., Xie, S., Wu, C.-Y., Yuille, A., & Feichtenhofer, C. (2021). Masked Feature Prediction for Self-Supervised Visual Pre-Training. arXiv. http://arxiv.org/abs/2112.09133v2
Chicago
Wei, C., H. Fan, S. Xie, C.-Y. Wu, A. Yuille, and C. Feichtenhofer. 2021. “Masked Feature Prediction for Self-Supervised Visual Pre-Training”. arXiv. http://arxiv.org/abs/2112.09133v2.
Harvard
Wei, C. et al. (2021) “Masked Feature Prediction for Self-Supervised Visual Pre-Training”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2112.09133v2.
Vancouver
1. Wei C, Fan H, Xie S, Wu C-Y, Yuille A, Feichtenhofer C (2021) Masked Feature Prediction for Self-Supervised Visual Pre-Training. arXiv

BibTeX

@article{wei2021masked,
  title = {Masked Feature Prediction for Self-Supervised Visual Pre-Training},
  author = {Wei, Chen and Fan, Haoqi and Xie, Saining and Wu, Chao-Yuan and Yuille, Alan and Feichtenhofer, Christoph},
  year = {2021},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2112.09133v2},
  eprint = {2112.09133}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE