Hard Patches Mining for Masked Image Modeling

Haochen WangKaiyou SongJunsong FanYuxi WangJin XieZhaoxiang Zhang

article2023CVPR101 citations

Proposes Hard Patches Mining, an adaptive masked image modeling framework that trains a model to predict patch-wise reconstruction difficulty and mask the most challenging regions, outperforming standard masked autoencoders on ImageNet-1K with half the pre-training epochs.

Listen

Visual AI models increasingly rely on self-supervised pre-training to learn general visual representations from unannotated image datasets. A prominent technique, masked image modeling, hides parts of an image and trains the computer vision model to reconstruct the missing sections. However, conventional methods rely on pre-defined or random masking rules that act only as rigid assignments for the model to solve. Because images contain substantial repetitive background information, random masking frequently hides uninformative areas rather than the core, discriminative objects necessary for robust understanding.

The article introduces and evaluates Hard Patches Mining, a self-supervised training framework designed to make the AI system act as both a teacher and a student. The main objective is to demonstrate that an AI model can autonomously identify which image regions are hardest to reconstruct and use that knowledge to generate progressively more demanding training tasks, ultimately improving representation quality across downstream vision applications.

To test this concept, the authors conducted extensive experiments using standard Vision Transformer backbones evaluated on established benchmarks, including ImageNet-1K classification, COCO object detection and instance segmentation, and ADE20k semantic segmentation. The system pairs a student network with a momentum-updated teacher network. The teacher predicts the relative reconstruction difficulty across image patches using a relative loss formulation, and an easy-to-hard scheduling strategy gradually increases the proportion of difficult patches hidden from the student during training.

The experimental findings show substantial performance and efficiency gains. First, Hard Patches Mining achieves 84.2% and 85.8% Top-1 fine-tuning accuracy on ImageNet-1K using standard base and large Vision Transformer models with 800 pre-training epochs, outperforming baseline Masked Autoencoders trained for twice as long (1600 epochs) by 0.6% and 0.7%, respectively. Second, under a shorter 200-epoch schedule, the method surpasses the baseline by 0.8% on the base model and 1.2% on the large model. Third, the benefits transfer strongly to downstream dense prediction tasks, improving object detection by 1.58 average precision points on COCO and semantic segmentation by 1.0 to 1.6 mean Intersection-over-Union points on ADE20k. Finally, ablation studies confirm that balancing difficult masks with a baseline degree of randomness is necessary to prevent eliminating all contextual clues.

These results indicate that teaching an AI model to identify salient, difficult image regions during pre-training significantly improves visual representation quality while reducing required training epochs. Engineering teams can integrate this approach as a flexible module into existing self-supervised pipelines—whether reconstructing raw pixels or distilling features—to achieve higher downstream task accuracy and lower pre-training compute budgets.

Organizations developing large-scale visual models should consider adopting learnable, difficulty-aware masking to enhance pre-training pipelines. Practitioners should implement the easy-to-hard schedule and retain a portion of random masking to prevent task collapse. Before large-scale production deployment, engineering teams should conduct pilot tests to optimize the framework for specific downstream workloads.

The findings are supported with high confidence across multiple architectures, pre-training objectives, and evaluation benchmarks. However, leaders should note key constraints: the framework requires an auxiliary prediction head that increases per-epoch training time by approximately 10%, and like other masked image modeling methods, it does not match contrastive learning baselines when evaluated via linear probing or nearest-neighbor classification.

  • Paper: Masked Autoencoders Are Scalable Vision Learners, Kaiming He et al. (2022). Introduces the asymmetric masked autoencoder framework and standard random patch masking strategy that Hard Patches Mining actively seeks to improve via adaptive difficulty estimation.
  • Paper: SimMIM: a Simple Framework for Masked Image Modeling, Zhenda Xie et al. (2021). Establishes a foundational masked image modeling framework based on direct pixel reconstruction and pre-defined masking, providing the core baseline context for HPM.
  • Paper: BEiT: BERT Pre-Training of Image Transformers, Hangbo Bao et al. (2022). Introduced the masked image modeling paradigm for Vision Transformers, formalizing patch corruption and recovery as a self-supervised pre-training objective.
  • Paper: Context Encoders: Feature Learning by Inpainting, Deepak Pathak et al. (2016). Pioneered representation learning through inpainting missing image regions, laying the historical foundation for reconstruction-driven visual pre-training.

No sufficiently relevant recommendations were found.

Cover for Hard Patches Mining for Masked Image Modeling

Abstract

Masked image modeling (MIM) has attracted much research attention due to its promising potential for learning scalable visual representations. In typical approaches, models usually focus on predicting specific contents of masked patches, and their performances are highly related to pre-defined mask strategies. Intuitively, this procedure can be considered as training a student (the model) on solving given problems (predict masked patches). However, we argue that the model should not only focus on solving given problems, but also stand in the shoes of a teacher to produce a more challenging problem by itself. To this end, we propose Hard Patches Mining (HPM), a brand-new framework for MIM pre-training. We observe that the reconstruction loss can naturally be the metric of the difficulty of the pre-training task. Therefore, we introduce an auxiliary loss predictor, predicting patch-wise losses first and deciding where to mask next. It adopts a relative relationship learning strategy to prevent overfitting to exact reconstruction loss values. Experiments under various settings demonstrate the effectiveness of HPM in constructing masked images. Furthermore, we empirically find that solely introducing the loss prediction objective leads to powerful representations, verifying the efficacy of the ability to be aware of where is hard to reconstruct.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Method
  • 3.1. Overview
  • 3.2. Image Reconstructor
  • 3.3. Hard Patches Mining with a Loss Predictor
  • 3.4. Easy-to-Hard Mask Generation
  • 4. Experiments
  • 4.1. Ablation Study
  • 4.2. Comparison with Previous Alternatives
  • 5. Conclusion
  • Acknowledgements
  • References

Knowls

  1. Knowl 1 — Hard Patches Mining Framework for Masked Image Modeling

    model/method

    Hard Patches Mining (HPM) is a masked image modeling (MIM) self-supervised pre-training framework that dynamically identifies and masks difficult-to-reconstruct image patches using a dual teacher-student architecture.

    The framework consists of a student network and a teacher network sharing an identical architecture composed of three sub-networks: an encoder fθ(⋅)f_\theta(\cdot), an image reconstructor dϕ(⋅)d_\phi(\cdot), and a patch reconstruction loss predictor dψ(⋅)d_\psi(\cdot). Parameter sets are denoted as Θs=(θs,ϕs,ψs)\Theta_s = (\theta_s, \phi_s, \psi_s) for the student and Θt=(θt,ϕt,ψt)\Theta_t = (\theta_t, \phi_t, \psi_t) for the teacher. The teacher parameters are updated via exponential moving average (EMA) momentum updates: Θt←mΘt+(1−m)Θs\Theta_t \leftarrow m\Theta_t + (1 - m)\Theta_s where m∈[0,1]m \in [0, 1] is the momentum coefficient.

    Given an input image I∈RH×W×C\mathbf{I} \in \mathbb{R}^{H \times W \times C}, it is flattened and partitioned into a sequence of NN 2D patches x∈RN×(P2C)\mathbf{x} \in \mathbb{R}^{N \times (P^2 C)}, where (H,W)(H, W) is the image resolution, CC is the channel dimension, PP is the patch size, and N=HW/P2N = HW/P^2.

    In each iteration:

    1. The unmasked patch sequence x\mathbf{x} is passed through the teacher to obtain patch-wise predicted reconstruction difficulty scores L^t=dψt(fθt(x))∈RN\hat{\mathcal{L}}^t = d_{\psi_t}(f_{\theta_t}(\mathbf{x})) \in \mathbb{R}^N.
    2. A binary mask M∈{0,1}N\mathbf{M} \in \{0, 1\}^N (where Mi=1M_i = 1 denotes visible patches and Mi=0M_i = 0 denotes masked patches) is constructed using a curriculum-based easy-to-hard strategy based on L^t\hat{\mathcal{L}}^t.
    3. The student receives only visible patches x⊙M\mathbf{x} \odot \mathbf{M} and is trained by minimizing the joint multi-task loss: L=Lrec+Lpred\mathcal{L} = \mathcal{L}_{\text{rec}} + \mathcal{L}_{\text{pred}} where Lrec=M(dϕs(fθs(x⊙M)),T(x⊙(1−M)))\mathcal{L}_{\text{rec}} = \mathcal{M}(d_{\phi_s}(f_{\theta_s}(\mathbf{x} \odot \mathbf{M})), \mathcal{T}(\mathbf{x} \odot (1 - \mathbf{M}))) measures image reconstruction error against target transformation T(⋅)\mathcal{T}(\cdot) (such as normalized pixel values or teacher representations) via similarity metric M(⋅,⋅)\mathcal{M}(\cdot, \cdot), and Lpred\mathcal{L}_{\text{pred}} supervises the loss predictor dψsd_{\psi_s}.
  2. Knowl 2 — Relative Binary Cross-Entropy Loss for Hard Patch Prediction

    equation

    In Hard Patches Mining (HPM), the auxiliary loss predictor is trained to learn the relative ranking of patch-wise reconstruction difficulty rather than exact loss magnitudes. Given the ground-truth per-patch reconstruction loss sequence Lrec∈RN\mathcal{L}_{\text{rec}} \in \mathbb{R}^N (detached from gradients) and student-predicted patch losses L^s=dψs(fθs(x⊙M))∈RN\hat{\mathcal{L}}^s = d_{\psi_s}(f_{\theta_s}(\mathbf{x} \odot \mathbf{M})) \in \mathbb{R}^N, the relative prediction loss Lpred\mathcal{L}_{\text{pred}} is defined over all pairs of masked patches as:

    Lpred=−∑i=1N∑j=1j≠iN1ij+log⁡(σ(L^is−L^js))−∑i=1N∑j=1j≠iN1ij−log⁡(1−σ(L^is−L^js))\mathcal{L}_{\text{pred}} = - \sum_{i=1}^N \sum_{\substack{j=1 \\ j \neq i}}^N \mathbb{1}^+_{ij} \log \left( \sigma(\hat{\mathcal{L}}^s_i - \hat{\mathcal{L}}^s_j) \right) - \sum_{i=1}^N \sum_{\substack{j=1 \\ j \neq i}}^N \mathbb{1}^-_{ij} \log \left( 1 - \sigma(\hat{\mathcal{L}}^s_i - \hat{\mathcal{L}}^s_j) \right)

    where σ(z)=ezez+1\sigma(z) = \frac{e^z}{e^z + 1} is the standard sigmoid function, and the pairwise relationship indicator variables 1ij+\mathbb{1}^+_{ij} and 1ij−\mathbb{1}^-_{ij} are defined as:

    1ij+={1,if Lrec(i)>Lrec(j) and Mi=Mj=00,otherwise\mathbb{1}^+_{ij} = \begin{cases} 1, & \text{if } \mathcal{L}_{\text{rec}}(i) > \mathcal{L}_{\text{rec}}(j) \text{ and } M_i = M_j = 0 \\ 0, & \text{otherwise} \end{cases}

    1ij−={1,if Lrec(i)<Lrec(j) and Mi=Mj=00,otherwise\mathbb{1}^-_{ij} = \begin{cases} 1, & \text{if } \mathcal{L}_{\text{rec}}(i) < \mathcal{L}_{\text{rec}}(j) \text{ and } M_i = M_j = 0 \\ 0, & \text{otherwise} \end{cases}

    Here Mi=Mj=0M_i = M_j = 0 ensures that relative pairwise comparisons are computed exclusively over patches that were masked during the training step. Formulating the loss as dense pairwise binary cross-entropy avoids optimization instability caused by the monotonic decrease in absolute reconstruction error scales throughout training.

  3. Knowl 3 — Easy-to-Hard Dynamic Mask Generation Strategy

    model/method

    To prevent the model from getting stuck in early training iterations when representations are dominated by high-frequency texture noise rather than semantic content, Hard Patches Mining (HPM) employs an easy-to-hard mask scheduling mechanism.

    Given the total number of patches NN, the total masking ratio γ∈(0,1)\gamma \in (0, 1) (such that γN\gamma N patches are masked in total), the current training epoch tt, and total training epochs TT, the proportion αt∈[0,1]\alpha_t \in [0, 1] of masked patches selected by difficulty is scheduled linearly: αt=α0+tT(αT−α0)\alpha_t = \alpha_0 + \frac{t}{T} (\alpha_T - \alpha_0) where α0\alpha_0 and αT\alpha_T are boundary hyperparameters specifying the initial and final difficulty proportions.

    Mask construction at epoch tt proceeds as follows:

    1. The teacher outputs predicted reconstruction losses L^t=dψt(fθt(x))∈RN\hat{\mathcal{L}}^t = d_{\psi_t}(f_{\theta_t}(\mathbf{x})) \in \mathbb{R}^N.
    2. The top αt⋅γN\alpha_t \cdot \gamma N patches with the largest predicted loss values in L^t\hat{\mathcal{L}}^t (obtained via descending argsort\text{argsort}) are designated as masked patches.
    3. The remaining (1−αt)⋅γN(1 - \alpha_t) \cdot \gamma N masked patches are sampled uniformly at random from the remaining unselected patches.

    By default, setting α0=0\alpha_0 = 0, αT=0.5\alpha_T = 0.5, and γ=0.75\gamma = 0.75 provides an optimal balance between targeted hard-region masking and necessary context retention.

  4. Knowl 4 — Training Algorithm of Hard Patches Mining

    algorithm

    The complete training step of HPM executes according to the following procedure:

    Input: Input patchified image x∈RN×(P2C)\mathbf{x} \in \mathbb{R}^{N \times (P^2 C)}, current epoch tt, total epochs TT, mask ratio γ\gamma, schedule bounds α0,αT\alpha_0, \alpha_T, student model Θs=(fθs,dϕs,dψs)\Theta_s = (f_{\theta_s}, d_{\phi_s}, d_{\psi_s}), teacher model Θt=(fθt,dϕt,dψt)\Theta_t = (f_{\theta_t}, d_{\phi_t}, d_{\psi_t}), momentum coefficient mm
    Output: Updated student parameters Θs\Theta_s and teacher parameters Θt\Theta_t
    Teacher inference:
      L^t←dψt(fθt(x))\hat{\mathcal{L}}^t \leftarrow d_{\psi_t}(f_{\theta_t}(\mathbf{x}))
    Mask generation:
      αt←α0+tT(αT−α0)\alpha_t \leftarrow \alpha_0 + \frac{t}{T}(\alpha_T - \alpha_0)
      Khard←⌊αt⋅γN⌋K_{\text{hard}} \leftarrow \lfloor \alpha_t \cdot \gamma N \rfloor
      Krand←⌊(1−αt)⋅γN⌋K_{\text{rand}} \leftarrow \lfloor (1 - \alpha_t) \cdot \gamma N \rfloor
      Ihard←topk_indices(L^t,Khard)I_{\text{hard}} \leftarrow \text{topk\_indices}(\hat{\mathcal{L}}^t, K_{\text{hard}})
      Irand←random_sample({1,…,N}∖Ihard,Krand)I_{\text{rand}} \leftarrow \text{random\_sample}(\{1, \dots, N\} \setminus I_{\text{hard}}, K_{\text{rand}})
      Imask←Ihard∪IrandI_{\text{mask}} \leftarrow I_{\text{hard}} \cup I_{\text{rand}}
      Construct binary mask M∈{0,1}N\mathbf{M} \in \{0, 1\}^N where Mi=0M_i = 0 if i∈Imaski \in I_{\text{mask}} else Mi=1M_i = 1
    Student forward pass:
      zs←fθs(x⊙M)\mathbf{z}_s \leftarrow f_{\theta_s}(\mathbf{x} \odot \mathbf{M})
      x^←dϕs(zs)\hat{\mathbf{x}} \leftarrow d_{\phi_s}(\mathbf{z}_s)
      L^s←dψs(zs)\hat{\mathcal{L}}^s \leftarrow d_{\psi_s}(\mathbf{z}_s)
    Compute losses:
      Lrec←meani∈Imask∥x^i−xi∥2\mathcal{L}_{\text{rec}} \leftarrow \text{mean}_{i \in I_{\text{mask}}} \| \hat{\mathbf{x}}_i - \mathbf{x}_i \|^2
      Lgt←(∥x^−x∥2).detach()\mathbf{L}_{\text{gt}} \leftarrow (\| \hat{\mathbf{x}} - \mathbf{x} \|^2).\text{detach}()
      Lpred←0\mathcal{L}_{\text{pred}} \leftarrow 0
      for each pair i,j∈Imaski, j \in I_{\text{mask}} with i≠ji \neq j do
        if Lgt(i)>Lgt(j)\mathbf{L}_{\text{gt}}(i) > \mathbf{L}_{\text{gt}}(j) then
          Lpred←Lpred−log⁡σ(L^is−L^js)\mathcal{L}_{\text{pred}} \leftarrow \mathcal{L}_{\text{pred}} - \log \sigma(\hat{\mathcal{L}}^s_i - \hat{\mathcal{L}}^s_j)
        else if Lgt(i)<Lgt(j)\mathbf{L}_{\text{gt}}(i) < \mathbf{L}_{\text{gt}}(j) then
          Lpred←Lpred−log⁡(1−σ(L^is−L^js))\mathcal{L}_{\text{pred}} \leftarrow \mathcal{L}_{\text{pred}} - \log (1 - \sigma(\hat{\mathcal{L}}^s_i - \hat{\mathcal{L}}^s_j))
      Lpred←Lpred/(∣Imask∣(∣Imask∣−1))\mathcal{L}_{\text{pred}} \leftarrow \mathcal{L}_{\text{pred}} / (|I_{\text{mask}}|(|I_{\text{mask}}| - 1))
      L←Lrec+Lpred\mathcal{L} \leftarrow \mathcal{L}_{\text{rec}} + \mathcal{L}_{\text{pred}}
    Optimization:
      Update Θs\Theta_s via gradient descent on L\mathcal{L}
      Θt←mΘt+(1−m)Θs\Theta_t \leftarrow m\Theta_t + (1 - m)\Theta_s
  5. Knowl 5 — ImageNet-1K Classification Performance of HPM

    data/table

    When evaluated on ImageNet-1K classification via fine-tuning at 224×224224 \times 224 resolution, HPM achieves consistent performance gains across different model scales and training durations compared to supervised baselines, contrastive learning approaches, and previous MIM methods.

    Method Effective Epochs ViT-B Top-1 (%) ViT-L Top-1 (%)
    Supervised (from scratch) - 80.9 82.6
    Contrastive Learning
    MoCo v3 600 83.2 84.1
    DINO 1600 83.6 -
    MIM with Pixel Regression
    MAE 200 82.2 83.3
    HPM (Ours) 200 83.0 84.5
    MAE 1600 83.6 85.1
    SimMIM 800 83.8 -
    HPM (Ours) 800 84.2 85.8
    MIM with Feature Distillation
    BEiT 800 83.2 85.2
    iBOT 1600 84.0 -
    BootMAE 800 84.2 85.9

    At 200 pre-training epochs with raw pixel targets, HPM improves ViT-B accuracy by +0.8%+0.8\% (83.0%83.0\% vs 82.2%82.2\%) and ViT-L accuracy by +1.2%+1.2\% (84.5%84.5\% vs 83.3%83.3\%) over standard MAE. At 800 pre-training epochs, HPM reaches 84.2%84.2\% on ViT-B and 85.8%85.8\% on ViT-L, surpassing MAE pre-trained for 1600 epochs by +0.6%+0.6\% and +0.7%+0.7\%, respectively.

  6. Knowl 6 — Ablation on Reconstruction Targets and Auxiliary Loss Predictor Components

    data/table

    The effectiveness of HPM is evaluated across multiple reconstruction targets: raw RGB pixel regression and feature distillation using EMA teacher features, DINO ViT-B features, and CLIP ViT-B features. All models use a ViT-B/16 backbone pre-trained for 200 epochs on ImageNet-1K and fine-tuned for 100 epochs.

    Target Lpred\mathcal{L}_{\text{pred}} Learn to Mask Fine-tune Top-1 (%) Linear Probing (%) kk-NN (%)
    Pixel Regression
    RGB (MAE baseline) - - 82.23 50.80 29.84
    RGB ✓ - 82.49 (+0.26) 51.26 31.98
    RGB ✓ ✓ 82.95 (+0.72) 54.92 36.09
    Feature Distillation
    EMA features - - 82.99 32.65 20.69
    EMA features ✓ - 83.13 (+0.14) 52.06 35.73
    EMA features ✓ ✓ 83.47 (+0.48) 55.25 35.94
    DINO features - - 83.46 61.31 41.53
    DINO features ✓ - 83.58 (+0.12) 63.25 43.02
    DINO features ✓ ✓ 84.13 (+0.67) 64.17 47.25
    CLIP features - - 83.20 59.80 42.51
    CLIP features ✓ - 83.31 (+0.11) 60.62 43.26
    CLIP features ✓ ✓ 83.58 (+0.38) 62.22 45.08

    Two critical conclusions emerge:

    1. Solely adding the loss prediction objective Lpred\mathcal{L}_{\text{pred}} with random masking consistently improves feature representation quality across all targets (+0.11% to +0.26% fine-tuning Top-1).
    2. Utilizing the predicted loss to drive the easy-to-hard masking strategy ("Learn to Mask") provides substantial additional improvements across all regression and feature distillation frameworks.
  7. Knowl 7 — Ablation of Mask Difficulty, Scheduling, and Selection Criteria

    data/table

    Ablation experiments analyze the interaction between mask ratio γ\gamma, schedule boundaries α0,αT\alpha_0, \alpha_T, patch difficulty selection criteria, and scheduling direction using ViT-B/16 pre-trained for 200 epochs on ImageNet-1K.

    Mask Strategy / Variant Selection / Manner γ\gamma (%) α0\alpha_0 αT\alpha_T Fine-tune Top-1 (%)
    Random baseline - 75 0 0 82.49
    Learn to mask (Default) argmax(L^t)\text{argmax}(\hat{\mathcal{L}}^t) 75 0 0.5 82.95 (+0.46)
    Learn to mask argmax(L^t)\text{argmax}(\hat{\mathcal{L}}^t) 75 0 1.0 82.67 (+0.18)
    Hard-only (No random) argmax(L^t)\text{argmax}(\hat{\mathcal{L}}^t) 75 1.0 1.0 81.40 (-1.09)
    Random baseline - 50 0 0 82.36
    Learn to mask argmax(L^t)\text{argmax}(\hat{\mathcal{L}}^t) 50 0 0.5 82.56 (+0.20)
    Hard-only (No random) argmax(L^t)\text{argmax}(\hat{\mathcal{L}}^t) 50 1.0 1.0 82.19 (-0.17)
    Random baseline - 90 0 0 82.48
    Learn to mask argmax(L^t)\text{argmax}(\hat{\mathcal{L}}^t) 90 0 0.5 82.66 (+0.18)
    Hard-only (No random) argmax(L^t)\text{argmax}(\hat{\mathcal{L}}^t) 90 1.0 1.0 80.59 (-1.89)
    Easy patches masking argmin(L^t)\text{argmin}(\hat{\mathcal{L}}^t) 75 0 0.5 82.36 (-0.13)
    Hard-to-easy schedule Hard-to-Easy 75 0.5 0 81.71 (-0.78)

    The empirical results show that:

    • Masking exclusively hard patches (α0=αT=1.0\alpha_0=\alpha_T=1.0) causes performance degradation (81.40%81.40\%) because visible patches contain only background without necessary contextual cues for reconstruction.
    • An easy-to-hard schedule (α0=0,αT=0.5\alpha_0=0, \alpha_T=0.5) with 50% hard and 50% random patch distribution at the end of training yields optimal results.
    • Masking easy patches (argmin\text{argmin}) or reversing the curriculum (hard-to-easy) underperforms random masking.
  8. Knowl 8 — Comparison of Loss Prediction Objectives: Relative BCE vs. Absolute MSE

    empirical result

    Comparing formulations for the auxiliary loss prediction objective dψsd_{\psi_s} on a ViT-B/16 backbone pre-trained for 200 epochs on ImageNet-1K demonstrates the superiority of relative ranking over absolute value regression.

    • Baseline (No loss prediction / MAE baseline): Fine-tuning Top-1 is 82.23%82.23\%, Linear probing is 51.26%51.26\%, and kk-NN accuracy is 31.98%31.98\%.
    • Absolute MSE Loss (Lpred=(dψs(fθs(x⊙M))−Lrec)2⊙(1−M)\mathcal{L}_{\text{pred}} = (d_{\psi_s}(f_{\theta_s}(\mathbf{x} \odot \mathbf{M})) - \mathcal{L}_{\text{rec}})^2 \odot (1 - \mathbf{M})): Fine-tuning Top-1 reaches 82.77%82.77\% (+0.54%+0.54\%), Linear probing reaches 51.85%51.85\%, and kk-NN accuracy reaches 34.47%34.47\%.
    • Relative BCE Loss (pairwise cross-entropy): Fine-tuning Top-1 reaches 82.95%82.95\% (+0.72%+0.72\%), Linear probing reaches 54.92%54.92\%, and kk-NN accuracy reaches 36.09%36.09\%.

    Relative BCE outperforms absolute MSE by +0.18%+0.18\% in fine-tuning, +3.07%+3.07\% in linear probing, and +1.62%+1.62\% in kk-NN, confirming that learning the relative difficulty order among patches avoids scale collapse issues caused by decaying reconstruction error magnitudes.

  9. Knowl 9 — Downstream Transfer Performance on COCO and ADE20k

    data/table

    Representations pre-trained with HPM on ImageNet-1K for 200 epochs transfer effectively to dense visual prediction benchmarks including object detection/instance segmentation on COCO (Mask R-CNN, 1×1\times 12-epoch schedule, 1024×10241024 \times 1024 resolution) and semantic segmentation on ADE20k (UperNet, 160k iterations for comparison with state-of-the-art or 80k for ablations, 512×512512 \times 512 resolution).

    Pre-train Method Target Lpred\mathcal{L}_{\text{pred}} Learn to Mask COCO APbox\text{AP}^{\text{box}} COCO APmask\text{AP}^{\text{mask}} ADE20k mIoU (80k)
    MAE Baseline RGB - - 40.45 37.01 40.49
    HPM (Loss Pred only) RGB ✓ - 40.98 (+0.53) 37.34 (+0.33) 41.45 (+0.96)
    HPM (Full) RGB ✓ ✓ 42.03 (+1.58) 38.15 (+1.14) 42.09 (+1.60)
    Distill Baseline CLIP - - 46.21 41.55 46.59
    HPM (Loss Pred only) CLIP ✓ - 46.43 (+0.22) 41.80 (+0.25) 46.97 (+0.38)
    HPM (Full) CLIP ✓ ✓ 46.57 (+0.36) 41.96 (+0.41) 47.35 (+0.76)

    When evaluated on ADE20k (160k schedule), HPM achieves 48.548.5 mIoU with ViT-B (outperforming MAE's 48.148.1 and BEiT's 47.147.1) and 54.654.6 mIoU with ViT-L (outperforming MAE's 53.653.6, BEiT's 53.353.3, and supervised pre-training's 49.949.9).

  10. Knowl 10 — Limitations and Computational Overhead of HPM

    limitation

    The Hard Patches Mining framework exhibits two primary limitations:

    1. Computational Overhead: Incorporating an auxiliary decoder dψd_\psi for loss prediction increases the computational budget, resulting in approximately ∼1.1×\sim 1.1\times training wall-clock time for a ViT-L backbone compared to standard MAE.
    2. Linear Probing / kk-NN Performance Gap: Although HPM improves linear probing (e.g., 54.92%54.92\% vs 50.80%50.80\%) and kk-NN evaluation over baseline MAE on ViT-B, its frozen representation metrics remain lower than those of contrastive self-supervised learning methods (such as DINO at 61.31%61.31\% linear probing Top-1).

Coverage note — None omitted. All substantial contributions—including the teacher-student architecture, relative loss formulation, easy-to-hard mask curriculum, pseudocode algorithm, main classification and segmentation results, and full ablation studies—are represented in the knowls.

References

  1. 1.Roman Bachmann, David Mizrahi, Andrei Atanov, and Amir Zamir. Multimae: Multi-modal multi-task masked autoencoders. In European Conference on Computer Vision (ECCV), 2022.
  2. 2.Alexei Baevski, Wei-Ning Hsu, Qiantong Xu, Arun Babu, Jiatao Gu, and Michael Auli. Data2vec: A general framework for self-supervised learning in speech, vision and language. In International Conference on Machine Learning (ICML), 2022.
  3. 3.Hangbo Bao, Li Dong, and Furu Wei. Beit: Bert pre-training of image transformers. In International Conference on Learning Representations (ICLR), 2022.
  4. 4.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. Advances in Neural Information Processing Systems (NeurIPS), 2020.
  5. 5.Mathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. Emerging properties in self-supervised vision transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2021.
  6. 6.Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. A simple framework for contrastive learning of visual representations. In International Conference on Machine Learning (ICML), 2020.
  7. 7.Xiaokang Chen, Mingyu Ding, Xiaodi Wang, Ying Xin, Shentong Mo, Yunhao Wang, Shumin Han, Ping Luo, Gang Zeng, and Jingdong Wang. Context autoencoder for self-supervised representation learning. arXiv preprint arXiv:2202.03026, 2022.
  8. 8.Xinlei Chen and Kaiming He. Exploring simple siamese representation learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2021.
  9. 9.Xinlei Chen, Saining Xie, and Kaiming He. An empirical study of training self-supervised vision transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2021.
  10. 10.MMSegmentation Contributors. MMSegmentation: Openmmlab semantic segmentation toolbox and benchmark. https://github.com/open-mmlab/mmsegmentation, 2020.
  11. 11.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. In NAACL, 2019.
  12. 12.Carl Doersch, Abhinav Gupta, and Alexei A Efros. Unsupervised visual representation learning by context prediction. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2015.
  13. 13.Xiaoyi Dong, Jianmin Bao, Ting Zhang, Dongdong Chen, Weiming Zhang, Lu Yuan, Dong Chen, Fang Wen, and Nenghai Yu. Bootstrapped masked autoencoders for vision bert pretraining. In European Conference on Computer Vision (ECCV), 2022.
  14. 14.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. In International Conference on Learning Representations (ICLR), 2021.
  15. 15.Ye Du, Yujun Shen, Haochen Wang, Jingjing Fei, Wei Li, Liwei Wu, Rui Zhao, Zehua Fu, and Qingjie Liu. Learning from future: A novel self-training framework for semantic segmentation. Advances in Neural Information Processing Systems (NeurIPS), 2022.
  16. 16.Mark Everingham, SM Eslami, Luc Van Gool, Christopher KI Williams, John Winn, and Andrew Zisserman. The pascal visual object classes challenge: A retrospective. International Journal of Computer Vision (IJCV), 2015.
  17. 17.Xinyang Geng, Hao Liu, Lisa Lee, Dale Schuurams, Sergey Levine, and Pieter Abbeel. Multimodal masked autoencoders learn transferable representations. In International Conference on Machine Learning Workshop (ICMLW), 2022.
  18. 18.Jean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec, Pierre Richemond, Elena Buchatskaya, Carl Doersch, Bernardo Avila Pires, Zhaohan Guo, Mohammad Gheshlaghi Azar, et al. Bootstrap your own latent-a new approach to self-supervised learning. Advances in Neural Information Processing Systems (NeurIPS), 2020.
  19. 19.Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross Girshick. Masked autoencoders are scalable vision learners. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022.
  20. 20.Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick. Momentum contrast for unsupervised visual representation learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020.
  21. 21.Kaiming He, Georgia Gkioxari, Piotr Dollár, and Ross Girshick. Mask r-cnn. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2017.
  22. 22.Zejiang Hou, Fei Sun, Yen-Kuang Chen, Yuan Xie, and Sun-Yuan Kung. Milan: Masked image pretraining on language assisted representation. arXiv preprint arXiv:2208.06049, 2022.
  23. 23.Ioannis Kakogeorgiou, Spyros Gidaris, Bill Psomas, Yannis Avrithis, Andrei Bursuc, Konstantinos Karantzalos, and Nikos Komodakis. What to hide from your students: Attention-guided masked image modeling. In European Conference on Computer Vision (ECCV), 2022.
  24. 24.Tero Karras, Samuli Laine, and Timo Aila. A style-based generator architecture for generative adversarial networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019.
  25. 25.Xiangwen Kong and Xiangyu Zhang. Understanding masked image modeling via learning occlusion invariant feature. arXiv preprint arXiv:2208.04164, 2022.
  26. 26.Gukyeong Kwon, Zhaowei Cai, Avinash Ravichandran, Erhan Bas, Rahul Bhotika, and Stefano Soatto. Masked vision and language modeling for multi-modal representation learning. arXiv preprint arXiv:2208.02131, 2022.
  27. 27.Buyu Li, Yu Liu, and Xiaogang Wang. Gradient harmonized single-stage detector. In Proceedings of the AAAI Conference on Artificial Intelligence (AAAI), 2019.
  28. 28.Gang Li, Heliang Zheng, Daqing Liu, Bing Su, and Changwen Zheng. Semmae: Semantic-guided masking for learning masked autoencoders. Advances in Neural Information Processing Systems (NeurIPS), 2022.
  29. 29.Xiang Li, Wenhai Wang, Lingfeng Yang, and Jian Yang. Uniform masking: Enabling mae pre-training for pyramid-based vision transformers with locality. arXiv preprint arXiv:2205.10063, 2022.
  30. 30.Yanghao Li, Saining Xie, Xinlei Chen, Piotr Dollar, Kaiming He, and Ross Girshick. Benchmarking detection transfer learning with vision transformers. arXiv preprint arXiv:2111.11429, 2021.
  31. 31.Tsung-Yi Lin, Piotr Dollár, Ross Girshick, Kaiming He, Bharath Hariharan, and Serge Belongie. Feature pyramid networks for object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2017.
  32. 32.Tsung-Yi Lin, Priya Goyal, Ross Girshick, Kaiming He, and Piotr Dollár. Focal loss for dense object detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2017.
  33. 33.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In European Conference on Computer Vision (ECCV), 2014.
  34. 34.Haotian Liu, Mu Cai, and Yong Jae Lee. Masked discrimination for self-supervised learning on point clouds. In European Conference on Computer Vision (ECCV), 2022.
  35. 35.Hao Liu, Xinghua Jiang, Xin Li, Antai Guo, Deqiang Jiang, and Bo Ren. The devil is in the frequency: Geminated gestalt autoencoder for self-supervised visual pre-training. arXiv preprint arXiv:2204.08227, 2022.
  36. 36.Jihao Liu, Xin Huang, Yu Liu, and Hongsheng Li. Mixmim: Mixed and masked image modeling for efficient visual representation learning. arXiv preprint arXiv:2205.13137, 2022.
  37. 37.Xingbin Liu, Jinghao Zhou, Tao Kong, Xianming Lin, and Rongrong Ji. Exploring target representations for masked autoencoders. arXiv preprint arXiv:2209.03917, 2022.
  38. 38.Ze Liu, Han Hu, Yutong Lin, Zhuliang Yao, Zhenda Xie, Yixuan Wei, Jia Ning, Yue Cao, Zheng Zhang, Li Dong, et al. Swin transformer v2: Scaling up capacity and resolution. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022.
  39. 39.Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2021.
  40. 40.Chen Min, Dawei Zhao, Liang Xiao, Yiming Nie, and Bin Dai. Voxel-mae: Masked autoencoders for pre-training large-scale point clouds. arXiv preprint arXiv:2206.09900, 2022.
  41. 41.Aaron van den Oord, Yazhe Li, and Oriol Vinyals. Representation learning with contrastive predictive coding. arXiv preprint arXiv:1807.03748, 2018.
  42. 42.Yatian Pang, Wenxiao Wang, Francis EH Tay, Wei Liu, Yonghong Tian, and Li Yuan. Masked autoencoders for point cloud self-supervised learning. In European Conference on Computer Vision (ECCV), 2022.
  43. 43.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In International Conference on Machine Learning (ICML), 2021.
  44. 44.Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al. Improving language understanding by generative pre-training. 2018.
  45. 45.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. Language models are unsupervised multitask learners. 2019.
  46. 46.Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever. Zero-shot text-to-image generation. In International Conference on Machine Learning (ICML), 2021.
  47. 47.Jason Tyler Rolfe. Discrete variational autoencoders. In International Conference on Learning Representations (ICLR), 2017.
  48. 48.Olaf Ronneberger, Philipp Fischer, and Thomas Brox. U-net: Convolutional networks for biomedical image segmentation. In International Conference on Medical image computing and computer-assisted intervention, 2015.
  49. 49.Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, et al. Imagenet large scale visual recognition challenge. International Journal of Computer Vision (IJCV), 2015.
  50. 50.Yuge Shi, N Siddharth, Philip Torr, and Adam R Kosiorek. Adversarial masking for self-supervised learning. In International Conference on Machine Learning (ICML), 2022.
  51. 51.Abhinav Shrivastava, Abhinav Gupta, and Ross Girshick. Training region-based object detectors with online hard example mining. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2016.
  52. 52.Zhan Tong, Yibing Song, Jue Wang, and Limin Wang. Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training. Advances in Neural Information Processing Systems (NeurIPS), 2022.
  53. 53.Wouter Van Gansbeke, Simon Vandenhende, Stamatios Georgoulis, and Luc Van Gool. Unsupervised semantic segmentation by contrasting object mask proposals. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2021.
  54. 54.Wouter Van Gansbeke, Simon Vandenhende, and Luc Van Gool. Discovering object masks with transformers for unsupervised semantic segmentation. arXiv preprint arXiv:2206.06363, 2022.
  55. 55.Rui Wang, Dongdong Chen, Zuxuan Wu, Yinpeng Chen, Xiyang Dai, Mengchen Liu, Yu-Gang Jiang, Luowei Zhou, and Lu Yuan. Bevt: Bert pretraining of video transformers. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022.
  56. 56.Wenguan Wang, Qiuxia Lai, Huazhu Fu, Jianbing Shen, Haibin Ling, and Ruigang Yang. Salient object detection in the deep learning era: An in-depth survey. IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), 2021.
  57. 57.Wenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan, Kaitao Song, Ding Liang, Tong Lu, Ping Luo, and Ling Shao. Pyramid vision transformer: A versatile backbone for dense prediction without convolutions. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2021.
  58. 58.Xiaolong Wang and Abhinav Gupta. Unsupervised learning of visual representations using videos. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2015.
  59. 59.Yuchao Wang, Jingjing Fei, Haochen Wang, Wei Li, Liwei Wu, Rui Zhao, and Yujun Shen. Balancing logit variation for long-tail semantic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023.
  60. 60.Yuchao Wang, Haochen Wang, Yujun Shen, Jingjing Fei, Wei Li, Guoqiang Jin, Liwei Wu, Rui Zhao, and Xinyi Le. Semi-supervised semantic segmentation using unreliable pseudo-labels. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022.
  61. 61.Chen Wei, Haoqi Fan, Saining Xie, Chao-Yuan Wu, Alan Yuille, and Christoph Feichtenhofer. Masked feature prediction for self-supervised visual pre-training. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022.
  62. 62.Yuxin Wu, Alexander Kirillov, Francisco Massa, Wan-Yen Lo, and Ross Girshick. Detectron2. https://github.com/facebookresearch/detectron2, 2019.
  63. 63.Zhirong Wu, Zihang Lai, Xiao Sun, and Stephen Lin. Extreme masking for learning instance and distributed visual representations. arXiv preprint arXiv:2206.04667, 2022.
  64. 64.Zhirong Wu, Yuanjun Xiong, Stella X Yu, and Dahua Lin. Unsupervised feature learning via non-parametric instance discrimination. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2018.
  65. 65.Tete Xiao, Yingcheng Liu, Bolei Zhou, Yuning Jiang, and Jian Sun. Unified perceptual parsing for scene understanding. In European Conference on Computer Vision (ECCV), 2018.
  66. 66.Jiahao Xie, Wei Li, Xiaohang Zhan, Ziwei Liu, Yew Soon Ong, and Chen Change Loy. Masked frequency modeling for self-supervised visual pre-training. arXiv preprint arXiv:2206.07706, 2022.
  67. 67.Zhenda Xie, Zheng Zhang, Yue Cao, Yutong Lin, Jianmin Bao, Zhuliang Yao, Qi Dai, and Han Hu. Simmim: A simple framework for masked image modeling. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022.
  68. 68.Weijian Xu, Yifan Xu, Tyler Chang, and Zhuowen Tu. Co-scale conv-attentional image transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2021.
  69. 69.Kun Yi, Yixiao Ge, Xiaotong Li, Shusheng Yang, Dian Li, Jianping Wu, Ying Shan, and Xiaohu Qie. Masked image modeling with denoising contrast. arXiv preprint arXiv:2205.09616, 2022.
  70. 70.Xumin Yu, Lulu Tang, Yongming Rao, Tiejun Huang, Jie Zhou, and Jiwen Lu. Point-bert: Pre-training 3d point cloud transformers with masked point modeling. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022.
  71. 71.Richard Zhang, Phillip Isola, and Alexei A Efros. Colorful image colorization. In European Conference on Computer Vision (ECCV), 2016.
  72. 72.Bolei Zhou, Hang Zhao, Xavier Puig, Sanja Fidler, Adela Barriuso, and Antonio Torralba. Scene parsing through ade20k dataset. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2017.
  73. 73.Jinghao Zhou, Chen Wei, Huiyu Wang, Wei Shen, Cihang Xie, Alan Yuille, and Tao Kong. Image bert pre-training with online tokenizer. In International Conference on Learning Representations (ICLR), 2022.

Citation

MLA
Wang, H., et al. “Hard Patches Mining for Masked Image Modeling”. arXiv, 2023, http://arxiv.org/abs/2304.05919v1.
APA
Wang, H., Song, K., Fan, J., Wang, Y., Xie, J., & Zhang, Z. (2023). Hard Patches Mining for Masked Image Modeling. arXiv. http://arxiv.org/abs/2304.05919v1
Chicago
Wang, H., K. Song, J. Fan, Y. Wang, J. Xie, and Z. Zhang. 2023. “Hard Patches Mining for Masked Image Modeling”. arXiv. http://arxiv.org/abs/2304.05919v1.
Harvard
Wang, H. et al. (2023) “Hard Patches Mining for Masked Image Modeling”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2304.05919v1.
Vancouver
1. Wang H, Song K, Fan J, Wang Y, Xie J, Zhang Z (2023) Hard Patches Mining for Masked Image Modeling. arXiv

BibTeX

@article{wang2023hard,
  title = {Hard Patches Mining for Masked Image Modeling},
  author = {Wang, Haochen and Song, Kaiyou and Fan, Junsong and Wang, Yuxi and Xie, Jin and Zhang, Zhaoxiang},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2304.05919v1},
  eprint = {2304.05919}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE