Masked Feature Prediction for Self-Supervised Visual Pre-Training
Chen WeiHaoqi FanSaining XieChao-Yuan WuAlan L. YuilleChristoph Feichtenhofer
Demonstrates that regressing simple Histograms of Oriented Gradients for masked video patches provides an efficient, highly effective self-supervised pre-training objective that sets state-of-the-art performance across major video action recognition benchmarks without relying on external tokenizers or extra supervision.
Modern computer vision models built on Transformer architectures require massive amounts of training data. Unlike natural language processing models, which excel by predicting masked words from unlabelled text, visual models have historically depended on expensive, manually labeled image datasets or complex multi-stage pipelines to avoid severe performance issues. In video recognition, this dependency is especially problematic because the gap between models trained on labeled datasets versus models trained from scratch on raw video has exceeded 5% in accuracy.
The article introduces and evaluates Masked Feature Prediction (MaskFeat), a self-supervised pre-training method designed directly for unlabeled video and image data. Its objective is to demonstrate that directly predicting simple visual feature descriptors in masked regions enables visual Transformers to achieve superior recognition performance without external labels or complex tokenizer networks.
MaskFeat operates by dividing a video or image into small space-time blocks, randomly concealing approximately 40% of them, and training a single neural network to predict feature representations of the hidden regions using only the visible surrounding content. The authors systematically evaluated five distinct prediction targets: raw pixel colors, Histograms of Oriented Gradients (HOG—a classical hand-crafted edge and texture descriptor), discrete visual tokens requiring extra pre-trained neural networks, continuous deep network features, and pseudo-class labels. These variants were tested across standard benchmark datasets, including Kinetics-400, Kinetics-600, AVA action detection, Something-Something v2, and ImageNet-1K.
The investigation produced several key findings. First, predicting standard HOG descriptors proved to be the most effective strategy, balancing high performance with computational efficiency while avoiding the need for secondary teacher networks or discrete visual vocabularies. Second, on the Kinetics-400 video benchmark, a MaskFeat-trained model achieved 86.7% top-1 accuracy using no external data, surpassing the best previous scratch-trained baseline by 5.2 percentage points and matching models trained on hundreds of millions of labeled images. Third, the pre-trained models transferred exceptionally well to fine-grained downstream tasks, reaching state-of-the-art results of 38.8 mAP on action localization in AVA and 75.0% accuracy on human-object interaction in Something-Something v2, outperforming models supervised on large image sets. Finally, experiments on single images showed that ViT models trained with MaskFeat on ImageNet-1K attained 85.7% accuracy, outperforming supervised models trained on datasets ten times larger.
These findings demonstrate that pre-training directly on raw, unlabeled video effectively replaces the costly practice of pre-training on massive supervised image collections. By using standard gradient descriptors with local contrast normalization, the model learns essential structural and motion patterns while ignoring irrelevant color and illumination variations. This significantly reduces data annotation costs, pipeline complexity, and the training overhead associated with complex multi-view contrastive setups or multi-stage tokenizers.
Organizations developing video and image understanding systems should adopt direct masked feature prediction with HOG descriptors to reduce reliance on external labeled data. Teams should prioritize spatiotemporal cube masking over single-frame masking when processing video. Further work should explore deploying this framework across longer video sequences, domain-specific industry video archives, and real-time inference constraints. Given the robust empirical results across multiple public benchmarks and diverse network scales, confidence in MaskFeat's core capability to streamline visual pre-training is high.
- Paper: Masked Autoencoders Are Scalable Vision Learners, Kaiming He et al. (2022). Read this image-based masked-autoencoder foundation first to understand the masking-and-prediction framework MaskFeat adapts for video.
- Paper: SimMIM: a Simple Framework for Masked Image Modeling, Zhenda Xie et al. (2021). SimMIM establishes direct prediction of masked visual content as a simple alternative to tokenizers, framing MaskFeat’s comparison of prediction targets.
- Paper: BEiT: BERT Pre-Training of Image Transformers, Hangbo Bao et al. (2022). BEiT shows how masked-language modeling can transfer to vision through discrete visual tokens, one of the target types MaskFeat evaluates against HOG.
- Paper: ViViT: A Video Vision Transformer, Anurag Arnab et al. (2021). ViViT introduces the video-transformer architectures whose spatiotemporal tokens provide essential context for MaskFeat’s video pre-training design.
- Paper: Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture, Mahmoud Assran et al. (2023). I-JEPA carries feature-space prediction further by learning abstract target representations, extending MaskFeat’s central question of what a masked visual model should predict.
- Paper: Efficient Self-supervised Learning with Contextualized Target Representations for Vision, Speech and Language, Alexei Baevski et al. (2023). data2vec 2.0 continues masked target prediction with contextualized teacher representations and a cross-modal efficiency focus, broadening the predictive approach beyond visual descriptors.
