A Closer Look at Self-Supervised Lightweight Vision Transformers

Shaoru WangJin GaoZeming LiXiaoqin ZhangWeiming Hu

article2023ICML70 citations

Demonstrates that proper self-supervised pre-training enables vanilla lightweight Vision Transformers to match specialized architectures on visual benchmarks, while introducing a pre-training distillation strategy to overcome performance drops on data-limited downstream tasks.

Listen

Deploying computer vision models on mobile and edge devices requires lightweight architectures that combine low latency with high accuracy. While self-supervised learning—which trains models on unlabeled data—has significantly advanced large-scale vision transformers, its impact on lightweight vision transformers has been largely overlooked. Consequently, the prevailing industry practice has been to develop complex, hybrid architectures to compensate for the perceived weaknesses of standard, simple vision transformers in resource-constrained settings.

The article systematically evaluates how self-supervised pre-training paradigms affect lightweight vision transformers and explores whether effective training strategies can eliminate the need for complicated architectural designs.

To investigate this, the researchers benchmarked self-supervised methods, including masked image modeling (Masked Autoencoders or MAE) and contrastive learning (MoCo-v3), against supervised baselines using standard lightweight models like ViT-Tiny (containing approximately 5.7 million parameters). Evaluation encompassed core benchmarks on ImageNet classification, transfer learning across six smaller classification datasets, and dense prediction tasks such as object detection and segmentation on COCO. They further used internal representation similarity and attention-mapping metrics to analyze layer behaviors.

The investigation produced several key findings. First, a simple lightweight vision transformer pre-trained with MAE achieved up to 79.0% top-1 accuracy on ImageNet, matching or outperforming heavily engineered state-of-the-art networks while maintaining high inference speed. Second, unlike large vision models, lightweight models do not benefit from scaling up pre-training data; accuracy remained flat even when training data increased by roughly tenfold. Third, standard self-supervised pre-training transferred poorly to data-limited downstream tasks, trailing fully supervised baselines on smaller datasets (for instance, lagging by roughly 10–17 percentage points on fine-grained benchmarks like Aircraft and Pets). Fourth, structural analysis showed that while lower model layers learn strong general patterns during reconstruction tasks, higher layers fail to develop rich semantic features necessary for small-data classification.

These findings suggest that engineering teams do not necessarily need to design intricate, proprietary architectures to deploy high-performing edge vision models; standard transformer architectures with streamlined operations are sufficient when pre-trained effectively. However, teams should exercise caution when deploying standard self-supervised lightweight models in data-scarce downstream environments, as higher-layer semantic degradation can compromise transfer performance.

To overcome this limitation, the authors developed a pre-training distillation approach that transfers attention patterns from a larger pre-trained teacher model directly to the student's highest layer. This distillation strategy substantially closed the transfer gap, improving accuracy on data-scarce tasks by up to 14.6 percentage points and outperforming supervised models on object detection and instance segmentation. Organizations building lightweight vision systems should adopt masked image modeling pre-training combined with attention distillation from an appropriately sized teacher model (such as a base model rather than an oversized large model) to balance training efficiency and transferability.

The study's primary limitations lie in its focus on classification, detection, and segmentation within standard benchmark datasets, leaving other edge vision domains (such as video analysis or 3D vision) for future validation. Nevertheless, the experimental results provide high confidence that effective pre-training and distillation strategies can bridge the performance gap between simple and complex lightweight vision architectures.

Cover for A Closer Look at Self-Supervised Lightweight Vision Transformers

Abstract

Self-supervised learning on large-scale Vision Transformers (ViTs) as pre-training methods has achieved promising downstream performance. Yet, how much these pre-training paradigms promote lightweight ViTs' performance is considerably less studied. In this work, we develop and benchmark several self-supervised pre-training methods on image classification tasks and some downstream dense prediction tasks. We surprisingly find that if proper pre-training is adopted, even vanilla lightweight ViTs show comparable performance to previous SOTA networks with delicate architecture design. It breaks the recently popular conception that vanilla ViTs are not suitable for vision tasks in lightweight regimes. We also point out some defects of such pre-training, e.g., failing to benefit from large-scale pre-training data and showing inferior performance on data-insufficient downstream tasks. Furthermore, we analyze and clearly show the effect of such pre-training by analyzing the properties of the layer representation and attention maps for related models. Finally, based on the above analyses, a distillation strategy during pre-training is developed, which leads to further downstream performance improvement for MAE-based pre-training. Code is available at https://github.com/wangsr126/mae-lite.

Table of Contents

  • 1. Introduction
  • 2. Preliminaries and Experimental Setup
  • 3. How Well Does Pre-Training Work on Lightweight ViTs?
  • 3.1. Benchmarks on ImageNet Classification Tasks
  • 3.2. Benchmarks on Transfer Performance
  • 4. Revealing the Secrets of the Pre-Training
  • 4.1. Layer Representation Analyses
  • 4.2. Attention Map Analyses
  • 5. Distillation Improves Pre-Trained Models
  • 6. Related Works
  • 7. Discussions
  • References
  • A. Experimental Details
  • A.1. Evaluation Details for MAE and MoCo-v3 on ImageNet
  • A.2. Pre-Training Details of MAE
  • A.3. Pre-Training Details of MoCo-v3
  • A.4. Transfer Evaluation Details on Classification Tasks
  • A.5. Transfer Evaluation Details on Dense Prediction Tasks
  • A.6. Analysis Methods
  • B. More Analyses on the Pre-Training
  • B.1. Analyses with More Models as Reference
  • B.2. Analyses Based on Linear Probing Evaluation
  • B.3. Analyses for More Self-Supervised Pre-Training Methods
  • C. More Analyses on Distillation
  • C.1. Illustration of the Distillation Process
  • C.2. Attention Map Analyses for the Distilled Pre-trained Models
  • C.3. Applying Distillation on More Networks
  • C.4. Distilling with Larger Teachers

Knowls

  1. Knowl 1 — MAE pre-training gives the strongest default ImageNet result for ViT-Tiny

    empirical result

    The study uses a vanilla ViT-Tiny with 12 Transformer blocks, embedding dimension 192, 12 attention heads, and 5.7M parameters. For MAE pre-training, the encoder is paired with a one-block, 192-dimensional decoder and a 75% masking ratio; MAE and MoCo-v3 are each pre-trained for 400 epochs on ImageNet-1k without labels. With fine-tuning on ImageNet-1k and evaluation on its validation set, the improved from-scratch supervised baseline reaches 75.8% top-1 accuracy, MoCo-v3 reaches 76.8%, and MAE reaches 78.0%. Supervised pre-training on labeled ImageNet-21k reaches 76.9% after 30 epochs or 77.8% after 300 epochs. The reported pre-training times are 23 hours for MAE, 52 hours for MoCo-v3, and 20 or 200 hours for the two supervised runs, measured on eight V100 GPUs. For the MoCo-v3 result, global average pooling at fine-tuning is important: it yields 76.8% versus 73.7% with the class-token configuration.

  2. Knowl 2 — Attention-map distillation transfers MAE-Base representations to a lightweight student

    model/method

    The distillation method pre-trains a lightweight MAE student alongside a frozen MAE-Base teacher. At corresponding Transformer layers, it matches teacher and student attention maps with a learnable mapping across their different numbers of attention heads:

    Lattn=MSE(AT,MAS).L_{\mathrm{attn}}=\mathrm{MSE}(A^T, M A^S).

    Here, AT∈Rh×l×lA^T\in\mathbb{R}^{h\times l\times l} and AS∈Rh′×l×lA^S\in\mathbb{R}^{h'\times l\times l} are the teacher and student attention maps, respectively; hh and h′h' are their numbers of attention heads, and ll is the number of image tokens. The learnable matrix M∈Rh×h′M\in\mathbb{R}^{h\times h'} maps student head channels to teacher head channels, and MSE is mean squared error over the aligned attention maps. The teacher and student encoder receive the same visible image patches. Student parameters are updated using the joint gradients from this attention-matching loss and MAE's reconstruction loss; teacher parameters remain fixed. The reported ViT-Tiny method distills the final Transformer layer.

  3. Knowl 3 — Distillation improves classification and dense-prediction transfer

    empirical result

    For a ViT-Tiny pre-trained for 400 epochs on ImageNet-1k with MAE, attention distillation from MAE-Base produces D-MAE-Tiny. The following values compare fine-tuned MAE-Tiny with D-MAE-Tiny; classification metrics are top-1 accuracy (%) and COCO metrics are average precision (AP): Flowers, 85.8 to 95.2 (+9.4); Pets, 76.5 to 89.1 (+12.6); Aircraft, 64.6 to 79.2 (+14.6); Cars, 78.8 to 87.5 (+8.7); CIFAR100, 78.9 to 85.0 (+6.1); iNaturalist 2018, 60.6 to 63.6 (+3.0); ImageNet-1k, 78.0 to 78.4 (+0.4); COCO object detection, 39.9 to 42.3 AP (+2.4); and COCO instance segmentation, 35.4 to 37.4 AP (+2.0). The largest gains occur on smaller classification datasets and on the two COCO tasks.

  4. Knowl 4 — Which pre-trained layers matter depends on downstream data availability

    empirical result

    The study tests layer contributions by retaining a chosen number of leading pre-trained Transformer blocks, randomly initializing the remaining blocks, and fine-tuning the resulting model. On ImageNet-1k, with 100 fine-tuning epochs, retaining leading blocks from either MAE-Tiny or MoCo-v3-Tiny recovers most of the gain over training all blocks from random initialization; adding higher blocks brings comparatively small further gains. On CIFAR100, Cars, and Pets, performance improves as more pre-trained blocks are retained, and the contribution of higher blocks becomes more important as the downstream dataset gets smaller. For these smaller-data transfers, MoCo-v3-Tiny's higher layers provide more benefit than MAE-Tiny's. Representation comparisons with supervised recognition models are consistent with this pattern: MAE-Tiny aligns relatively well in lower layers but less in higher layers, whereas MoCo-v3-Tiny aligns across more of its depth. The alignment is an indicator of representation similarity, not a direct measurement of downstream accuracy.

  5. Knowl 5 — More pre-training data does not improve the tested lightweight models

    empirical result

    With the number of pre-training iterations held constant, the authors compare ImageNet-1k top-1 validation accuracy (%) after MAE or MoCo-v3 pre-training on different data scales and distributions. On all of ImageNet-1k, MoCo-v3 scores 76.8 and MAE 78.0. On 1% of ImageNet-1k, the scores are 76.2 (-0.6) and 77.9 (-0.1); on 10%, 76.5 (-0.3) and 78.0 (+0.0); on a long-tailed ImageNet-1k subset, 76.1 (-0.7) and 77.9 (-0.1); and on the larger ImageNet-21k, 76.9 (+0.1) and 78.0 (+0.0). Changes in parentheses are relative to pre-training on all of ImageNet-1k. Thus, neither method shows a meaningful gain from the larger pre-training dataset in this fixed-iteration comparison, and MAE's results vary less across the tested subsets.

  6. Knowl 6 — Self-supervised initialization transfers less reliably than supervised initialization

    empirical result

    Fine-tuning ViT-Tiny initializations on classification and COCO tasks yields the following top-1 accuracy (%) for classification and AP for COCO detection/segmentation. Dataset training/test sizes and class counts are: Flowers 2k/6k/102; Pets 4k/4k/37; Aircraft 7k/3k/100; Cars 8k/8k/196; CIFAR100 50k/10k/100; iNaturalist 2018 438k/24k/8142; and COCO 118k/50k/80. In that task order, supervised DeiT-Tiny scores 96.4, 93.1, 73.5, 85.6, 85.8, 63.6, 40.4, and 35.5. MoCo-v3-Tiny scores 94.8, 87.8, 73.7, 83.9, 83.9, 54.5, 39.7, and 35.1; MAE-Tiny scores 85.8, 76.5, 64.6, 78.8, 78.9, 60.6, 39.9, and 35.4. Self-supervised initialization is generally below supervised initialization, while the gap tends to narrow on larger downstream datasets. MoCo-v3 slightly exceeds the supervised result on Aircraft, so the comparison is not uniformly in favor of the supervised model on every task.

  7. Knowl 7 — MAE pre-training is associated with more local and concentrated attention

    empirical result

    The study characterizes attention for token jj in head hh using attention distance and entropy. Let AhA_h be the l×ll\times l attention matrix for that head, where ll is the number of image tokens; let ph,i,j=softmax(Ah)i,jp_{h,i,j}=\mathrm{softmax}(A_h)_{i,j} be the normalized weight from token jj to token ii; and let Gi,jG_{i,j} be the Euclidean distance between their spatial positions. The metrics are

    Dh,j=∑iph,i,jGi,j,Eh,j=−∑iph,i,jlog⁡ph,i,j.D_{h,j}=\sum_i p_{h,i,j}G_{i,j},\qquad E_{h,j}=-\sum_i p_{h,i,j}\log p_{h,i,j}.

    Lower attention distance indicates a more local focus, and lower entropy indicates attention concentrated on fewer tokens. Compared with a randomly initialized, supervisedly trained ViT-Tiny, an MAE-pre-trained and fine-tuned model has lower distance and entropy in middle layers, consistent with a more local and concentrated attention pattern. MoCo-v3-Tiny generally has more global and broader attention than MAE-Tiny. The authors hypothesize that broad attention may let MoCo-v3-Tiny rely on global features while overlooking local patterns, which could disadvantage fine-grained recognition; they also suggest that MAE-Tiny's local, concentrated attention in higher layers may impede transfer to data-insufficient tasks. Attention remains diverse across heads, so the reported locality does not mean all heads attend only to nearby tokens.

  8. Knowl 8 — MAE-pre-trained vanilla ViT-Tiny reaches 79.0% on ImageNet-1k

    empirical result

    Under a stronger fine-tuning protocol using relative position embeddings, MAE-Tiny fine-tuned for 300 epochs reaches 78.5% top-1 accuracy on ImageNet-1k, compared with 76.2% for the from-scratch DeiT-Tiny baseline at 300 epochs. At 1,000 fine-tuning epochs, MAE-Tiny reaches 79.0%, compared with 77.8% for the from-scratch model. The fine-tuned MAE model has 5.7M parameters and reported throughput of 4,020 images/s. Its 79.0% result exceeds reported results for several lightweight alternatives, including MobileViT-S at 78.3% (5.6M parameters), LeViT-128 at 78.6% (9.2M), and ConvNeXt V2-F at 78.5% (5.2M, with self-supervised pre-training). These comparisons support the paper's claim that pre-training can make a plain, lightweight ViT competitive with architecturally specialized models; they do not establish superiority over every listed network.

  9. Knowl 9 — Distillation works best from a suitably sized teacher and on later layers

    empirical result

    The authors compare attention distillation at corresponding teacher and student layers 3, 6, 9, or 12, using a 100-epoch ImageNet-1k fine-tuning evaluation. Distilling the final layer gives the strongest result among these choices. They also compare MAE-Base with smaller and larger teachers for a ViT-Tiny student. In the order Flowers, Pets, Aircraft, Cars, CIFAR100, iNaturalist 2018, and ImageNet-1k, no teacher yields 85.8, 76.5, 64.6, 78.8, 78.9, 60.6, and 78.0%; MAE-Small yields 89.4, 78.6, 65.2, 78.9, 79.6, 61.5, and 78.1%; MAE-Base yields 95.2, 89.1, 79.2, 87.5, 85.0, 63.6, and 78.4%; and MAE-Large yields 94.0, 87.3, 77.1, 85.2, 84.2, 63.1, and 78.3%. MAE-Base is best among the tested teachers on every listed task; the authors associate the weaker results from the smallest and largest teachers with, respectively, insufficient high-layer representations and a capacity mismatch. The method also improves a ViT-Small student: MAE-Small to D-MAE-Small changes top-1 accuracy from 91.2 to 95.8% on Flowers, 82.0 to 91.4% on Pets, 65.8 to 80.7% on Aircraft, 79.2 to 88.3% on Cars, 80.8 to 87.8% on CIFAR100, 63.2 to 66.9% on iNaturalist 2018, and 82.1 to 82.5% on ImageNet-1k.

  10. Knowl 10 — The evaluation is limited to classification and selected dense-prediction tasks

    limitation

    The study evaluates classification and some dense-prediction tasks, including object detection and instance segmentation, but does not establish how the pre-training findings or distillation method behave on other vision task families. The authors leave evaluation on additional tasks for future work.

Coverage note — Supplementary linear-probing results and the full DINO/SimMIM comparisons are omitted because they are supporting evaluations rather than load-bearing for the main fine-tuning, analysis, and distillation findings.

References

  1. 1.Abbasi Koohpayegani, S., Tejankar, A., and Pirsiavash, H. Compress: Self-supervised learning by compressing representations. Adv. Neural Inform. Process. Syst., 33:12980–12992, 2020.
  2. 2.Ali, A., Touvron, H., Caron, M., Bojanowski, P., Douze, M., Joulin, A., Laptev, I., Neverova, N., Synnaeve, G., Verbeek, J., et al. Xcit: Cross-covariance image transformers. Adv. Neural Inform. Process. Syst., 34, 2021.
  3. 3.Assran, M., Caron, M., Misra, I., Bojanowski, P., Joulin, A., Ballas, N., and Rabbat, M. Semi-supervised learning of visual features by non-parametrically predicting view assignments with support samples. In Int. Conf. Comput. Vis., pp. 8443–8452, 2021.
  4. 4.Ba, J. L., Kiros, J. R., and Hinton, G. E. Layer normalization. ArXiv, abs/1607.06450, 2016.
  5. 5.Bao, H., Dong, L., and Wei, F. Beit: Bert pre-training of image transformers. ArXiv, abs/2106.08254, 2021.
  6. 6.Bucilua, C., Caruana, R., and Niculescu-Mizil, A. Model compression. In ACM Int. Conf. on Knowledge Discovery and Data Mining, pp. 535–541, 2006.
  7. 7.Caron, M., Misra, I., Mairal, J., Goyal, P., Bojanowski, P., and Joulin, A. Unsupervised learning of visual features by contrasting cluster assignments. In Adv. Neural Inform. Process. Syst., 2020.
  8. 8.Caron, M., Touvron, H., Misra, I., Jegou, H., Mairal, J., Bojanowski, P., and Joulin, A. Emerging properties in self-supervised vision transformers. In Int. Conf. Comput. Vis., pp. 9650–9660, 2021.
  9. 9.Chen, X., Fan, H., Girshick, R., and He, K. Improved baselines with momentum contrastive learning. ArXiv, abs/2003.04297, 2020.
  10. 10.Chen, X., Xie, S., and He, K. An empirical study of training self-supervised vision transformers. In Int. Conf. Comput. Vis., pp. 9640–9649, 2021a.
  11. 11.Chen, Y., Dai, X., Chen, D., Liu, M., Dong, X., Yuan, L., and Liu, Z. Mobile-former: Bridging mobilenet and transformer. ArXiv, abs/2108.05895, 2021b.
  12. 12.Cho, J. H. and Hariharan, B. On the efficacy of knowledge distillation. In Int. Conf. Comput. Vis., pp. 4794–4802, 2019.
  13. 13.Choi, H. M., Kang, H., and Oh, D. Unsupervised representation transfer for small networks: I believe i can distill on-the-fly. In Adv. Neural Inform. Process. Syst., 2021.
  14. 14.Cortes, C., Mohri, M., and Rostamizadeh, A. Algorithms for learning kernels based on centered alignment. The Journal of Machine Learning Research, 13:795–828, 2012.
  15. 15.Cubuk, E. D., Zoph, B., Shlens, J., and Le, Q. Randaugment: Practical automated data augmentation with a reduced search space. In Adv. Neural Inform. Process. Syst., volume 33, pp. 18613–18624, 2020.
  16. 16.Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L. Imagenet: A large-scale hierarchical image database. In IEEE Conf. Comput. Vis. Pattern Recog., pp. 248–255, 2009.
  17. 17.Dosovitskiy, A., Springenberg, J. T., Riedmiller, M., and Brox, T. Discriminative unsupervised feature learning with convolutional neural networks. Adv. Neural Inform. Process. Syst., 27:766–774, 2014.
  18. 18.Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al. An image is worth 16x16 words: Transformers for image recognition at scale. In Int. Conf. Learn. Represent., 2020.
  19. 19.Fang, Z., Wang, J., Wang, L., Zhang, L., Yang, Y., and Liu, Z. Seed: Self-supervised distillation for visual representation. In Int. Conf. Learn. Represent., 2020.
  20. 20.Gidaris, S., Singh, P., and Komodakis, N. Unsupervised representation learning by predicting image rotations. In Int. Conf. Learn. Represent., 2018.
  21. 21.Goyal, P., Dollar, P., Girshick, R., Noordhuis, P., Wesolowski, L., Kyrola, A., Tulloch, A., Jia, Y., and He, K. Accurate, large minibatch sgd: Training imagenet in 1 hour. ArXiv, abs/1706.02677, 2017.
  22. 22.Graham, B., El-Nouby, A., Touvron, H., Stock, P., Joulin, A., Jegou, H., and Douze, M. Levit: A vision transformer in convnet’s clothing for faster inference. In Int. Conf. Comput. Vis., pp. 12259–12269, 2021.
  23. 23.Grill, J.-B., Strub, F., Altche, F., Tallec, C., Richemond, P., Buchatskaya, E., Doersch, C., Pires, B., Guo, Z., Azar, M., et al. Bootstrap your own latent: A new approach to self-supervised learning. In Adv. Neural Inform. Process. Syst., 2020.
  24. 24.He, K., Zhang, X., Ren, S., and Sun, J. Deep residual learning for image recognition. In IEEE Conf. Comput. Vis. Pattern Recog., pp. 770–778, 2016.
  25. 25.He, K., Fan, H., Wu, Y., Xie, S., and Girshick, R. Momentum contrast for unsupervised visual representation learning. In IEEE Conf. Comput. Vis. Pattern Recog., pp. 9729–9738, 2020.
  26. 26.He, K., Chen, X., Xie, S., Li, Y., Dollar, P., and Girshick, R. Masked autoencoders are scalable vision learners. ArXiv, abs/2111.06377, 2021.
  27. 27.Heo, B., Yun, S., Han, D., Chun, S., Choe, J., and Oh, S. J. Rethinking spatial dimensions of vision transformers. In Int. Conf. Comput. Vis., pp. 11936–11945, 2021.
  28. 28.Hinton, G., Vinyals, O., and Dean, J. Distilling the knowledge in a neural network. ArXiv, abs/1503.02531, 2015.
  29. 29.Howard, A., Sandler, M., Chu, G., Chen, L.-C., Chen, B., Tan, M., Wang, W., Zhu, Y., Pang, R., Vasudevan, V., Le, Q. V., and Adam, H. Searching for mobilenetv3. In Int. Conf. Comput. Vis., 2019.
  30. 30.Huang, G., Sun, Y., Liu, Z., Sedra, D., and Weinberger, K. Q. Deep networks with stochastic depth. In Eur. Conf. Comput. Vis., pp. 646–661, 2016.
  31. 31.Jiao, X., Yin, Y., Shang, L., Jiang, X., Chen, X., Li, L., Wang, F., and Liu, Q. Tinybert: Distilling bert for natural language understanding. In Findings of Empirical Methods in Natural Language Process., pp. 4163–4174, 2020.
  32. 32.Jin, X., Peng, B., Wu, Y., Liu, Y., Liu, J., Liang, D., Yan, J., and Hu, X. Knowledge distillation via route constrained optimization. In Int. Conf. Comput. Vis., pp. 1345–1354, 2019.
  33. 33.Krause, J., Stark, M., Deng, J., and Fei-Fei, L. 3d object representations for fine-grained categorization. In Int. Conf. Comput. Vis. Worksh., pp. 554–561, 2013.
  34. 34.Krizhevsky, A. et al. Learning multiple layers of features from tiny images. Technical Report, 2009.
  35. 35.Li, Y., Xie, S., Chen, X., Dollar, P., He, K., and Girshick, R. Benchmarking detection transfer learning with vision transformers. ArXiv, abs/2111.11429, 2021.
  36. 36.Li, Y., Mao, H., Girshick, R., and He, K. Exploring plain vision transformer backbones for object detection. ArXiv, abs/2203.16527, 2022.
  37. 37.Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollar, P., and Zitnick, C. L. Microsoft coco: Common objects in context. In Eur. Conf. Comput. Vis., pp. 740–755, 2014.
  38. 38.Liu, Z., Miao, Z., Zhan, X., Wang, J., Gong, B., and Yu, S. X. Large-scale long-tailed recognition in an open world. In IEEE Conf. Comput. Vis. Pattern Recog., 2019.
  39. 39.Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., and Guo, B. Swin transformer: Hierarchical vision transformer using shifted windows. In Int. Conf. Comput. Vis., pp. 10012–10022, 2021.
  40. 40.Liu, Z., Mao, H., Wu, C.-Y., Feichtenhofer, C., Darrell, T., and Xie, S. A convnet for the 2020s. In IEEE Conf. Comput. Vis. Pattern Recog., pp. 11966–11976, 2022.
  41. 41.Loshchilov, I. and Hutter, F. Sgdr: Stochastic gradient descent with warm restarts. ArXiv, abs/1608.03983, 2016.
  42. 42.Maji, S., Rahtu, E., Kannala, J., Blaschko, M., and Vedaldi, A. Fine-grained visual classification of aircraft. ArXiv, abs/1306.5151, 2013.
  43. 43.Mehta, S. and Rastegari, M. Mobilevit: Light-weight, general-purpose, and mobile-friendly vision transformer. In Int. Conf. Learn. Represent., 2022.
  44. 44.Mirzadeh, S. I., Farajtabar, M., Li, A., Levine, N., Matsukawa, A., and Ghasemzadeh, H. Improved knowledge distillation via teacher assistant. In AAAI Conf. on Artificial Intelligence, volume 34, pp. 5191–5198, 2020.
  45. 45.Newell, A. and Deng, J. How useful is self-supervised pretraining for visual tasks? In IEEE Conf. Comput. Vis. Pattern Recog., pp. 7345–7354, 2020.
  46. 46.Nguyen, T., Raghu, M., and Kornblith, S. Do wide and deep networks learn the same things? uncovering how neural network representations vary with width and depth. In Int. Conf. Learn. Represent., 2020.
  47. 47.Nilsback, M.-E. and Zisserman, A. Automated flower classification over a large number of classes. In Indian Conference on Computer Vision, Graphics & Image Processing, pp. 722–729, 2008.
  48. 48.Noroozi, M. and Favaro, P. Unsupervised learning of visual representations by solving jigsaw puzzles. In Eur. Conf. Comput. Vis., pp. 69–84, 2016.
  49. 49.Pan, J., Bulat, A., Tan, F., Zhu, X., Dudziak, L., Li, H., Tzimiropoulos, G., and Martinez, B. Edgevits: Competing light-weight cnns on mobile devices with vision transformers. ArXiv, abs/2205.0343, 2022.
  50. 50.Parkhi, O. M., Vedaldi, A., Zisserman, A., and Jawahar, C. V. Cats and dogs. In IEEE Conf. Comput. Vis. Pattern Recog., pp. 3498–3505, 2012.
  51. 51.Raghu, M., Unterthiner, T., Kornblith, S., Zhang, C., and Dosovitskiy, A. Do vision transformers see like convolutional neural networks? Adv. Neural Inform. Process. Syst., 34, 2021.
  52. 52.Ridnik, T., Ben-Baruch, E., Noy, A., and Zelnik-Manor, L. Imagenet-21k pretraining for the masses. ArXiv, abs/2104.10972, 2021.
  53. 53.Romero, A., Ballas, N., Kahou, S. E., Chassang, A., Gatta, C., and Bengio, Y. Fitnets: Hints for thin deep nets. ArXiv, abs/1412.6550, 2014.
  54. 54.Sandler, M., Howard, A., Zhu, M., Zhmoginov, A., and Chen, L.-C. Mobilenetv2: Inverted residuals and linear bottlenecks. In IEEE Conf. Comput. Vis. Pattern Recog., pp. 4510–4520, 2018.
  55. 55.Sanh, V., Debut, L., Chaumond, J., and Wolf, T. Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter. ArXiv, abs/1910.01108, 2019.
  56. 56.Song, L., Smola, A., Gretton, A., Bedo, J., and Borgwardt, K. Feature selection via dependence maximization. The Journal of Machine Learning Research, 13(5), 2012.
  57. 57.Steiner, A., Kolesnikov, A., Zhai, X., Wightman, R., Uszkoreit, J., and Beyer, L. How to train your vit? data, augmentation, and regularization in vision transformers. ArXiv, abs/2106.10270, 2021.
  58. 58.Sun, Z., Yu, H., Song, X., Liu, R., Yang, Y., and Zhou, D. Mobilebert: a compact task-agnostic bert for resource-limited devices. In Association for Computational Linguistics, pp. 2158–2170, 2020.
  59. 59.Tan, M. and Le, Q. Efficientnet: Rethinking model scaling for convolutional neural networks. In Int. Conf. Machine Learning., pp. 6105–6114, 2019.
  60. 60.Touvron, H., Cord, M., Douze, M., Massa, F., Sablayrolles, A., and Jegou, H. Training data-efficient image transformers & distillation through attention. In Int. Conf. Machine Learning., volume 139, pp. 10347–10357, 2021a.
  61. 61.Touvron, H., Cord, M., Sablayrolles, A., Synnaeve, G., and Jegou, H. Going deeper with image transformers. In Int. Conf. Comput. Vis., pp. 32–42, 2021b.
  62. 62.Touvron, H., Cord, M., and Jegou, H. Deit iii: Revenge of the vit. ArXiv, abs/2204.07118, 2022.
  63. 63.Van Horn, G., Mac Aodha, O., Song, Y., Cui, Y., Sun, C., Shepard, A., Adam, H., Perona, P., and Belongie, S. The inaturalist species classification and detection dataset. In IEEE Conf. Comput. Vis. Pattern Recog., 2018.
  64. 64.Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. Attention is all you need. Adv. Neural Inform. Process. Syst., 30, 2017.
  65. 65.Wang, W., Bao, H., Huang, S., Dong, L., and Wei, F. Minilmv2: Multi-head self-attention relation distillation for compressing pretrained transformers. In Findings of Int. Joint Conf. on Natural Language Process., pp. 2140–2151, 2021.
  66. 66.Wightman, R. Pytorch image models. https://github.com/rwightman/pytorch-image-models, 2019.
  67. 67.Wightman, R., Touvron, H., and Jegou, H. Resnet strikes back: An improved training procedure in timm. ArXiv, abs/2110.00476, 2021.
  68. 68.Woo, S., Debnath, S., Hu, R., Chen, X., Liu, Z., Kweon, I.-S., and Xie, S. Convnext v2: Co-designing and scaling convnets with masked autoencoders. ArXiv, abs/2301.00808, 2023.
  69. 69.Xie, Z., Zhang, Z., Cao, Y., Lin, Y., Bao, J., Yao, Z., Dai, Q., and Hu, H. Simmim: A simple framework for masked image modeling. In IEEE Conf. Comput. Vis. Pattern Recog., 2022.
  70. 70.Yun, S., Han, D., Oh, S. J., Chun, S., Choe, J., and Yoo, Y. Cutmix: Regularization strategy to train strong classifiers with localizable features. In Int. Conf. Comput. Vis., pp. 6023–6032, 2019.
  71. 71.Zhang, H., Cisse, M., Dauphin, Y. N., and Lopez-Paz, D. mixup: Beyond empirical risk minimization. In Int. Conf. Learn. Represent., 2018.
  72. 72.Zhang, R., Isola, P., and Efros, A. A. Colorful image colorization. In Eur. Conf. Comput. Vis., pp. 649–666, 2016.
  73. 73.Zhou, J., Wei, C., Wang, H., Shen, W., Xie, C., Yuille, A., and Kong, T. ibot: Image bert pre-training with online tokenizer. Int. Conf. Learn. Represent., 2022.

Citation

MLA
Wang, S., et al. “A Closer Look at Self-Supervised Lightweight Vision Transformers”. International Conference on Machine Learning, vol. 202, 2023, pp. 35624–41, https://proceedings.mlr.press/v202/wang23e.html.
APA
Wang, S., Gao, J., Li, Z., Zhang, X., & Hu, W. (2023). A Closer Look at Self-Supervised Lightweight Vision Transformers. International Conference on Machine Learning, 202, 35624–35641. https://proceedings.mlr.press/v202/wang23e.html
Chicago
Wang, S., J. Gao, Z. Li, X. Zhang, and W. Hu. 2023. “A Closer Look at Self-Supervised Lightweight Vision Transformers”. International Conference on Machine Learning 202: 35624–41. https://proceedings.mlr.press/v202/wang23e.html.
Harvard
Wang, S. et al. (2023) “A Closer Look at Self-Supervised Lightweight Vision Transformers”, International Conference on Machine Learning. PMLR, pp. 35624–35641. Available at: https://proceedings.mlr.press/v202/wang23e.html.
Vancouver
1. Wang S, Gao J, Li Z, Zhang X, Hu W (2023) A Closer Look at Self-Supervised Lightweight Vision Transformers. In: International Conference on Machine Learning. PMLR, pp 35624–35641

BibTeX

@InProceedings{pmlr-v202-wang23e,
  title = 	 {A Closer Look at Self-Supervised Lightweight Vision Transformers},
  author =       {Wang, Shaoru and Gao, Jin and Li, Zeming and Zhang, Xiaoqin and Hu, Weiming},
  booktitle = 	 {Proceedings of the 40th International Conference on Machine Learning},
  pages = 	 {35624--35641},
  year = 	 {2023},
  editor = 	 {Krause, Andreas and Brunskill, Emma and Cho, Kyunghyun and Engelhardt, Barbara and Sabato, Sivan and Scarlett, Jonathan},
  volume = 	 {202},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {23--29 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v202/wang23e/wang23e.pdf},
  url = 	 {https://proceedings.mlr.press/v202/wang23e.html},
  abstract = 	 {Self-supervised learning on large-scale Vision Transformers (ViTs) as pre-training methods has achieved promising downstream performance. Yet, how much these pre-training paradigms promote lightweight ViTs’ performance is considerably less studied. In this work, we develop and benchmark several self-supervised pre-training methods on image classification tasks and some downstream dense prediction tasks. We surprisingly find that if proper pre-training is adopted, even vanilla lightweight ViTs show comparable performance to previous SOTA networks with delicate architecture design. It breaks the recently popular conception that vanilla ViTs are not suitable for vision tasks in lightweight regimes. We also point out some defects of such pre-training, e.g., failing to benefit from large-scale pre-training data and showing inferior performance on data-insufficient downstream tasks. Furthermore, we analyze and clearly show the effect of such pre-training by analyzing the properties of the layer representation and attention maps for related models. Finally, based on the above analyses, a distillation strategy during pre-training is developed, which leads to further downstream performance improvement for MAE-based pre-training. Code is available at https://github.com/wangsr126/mae-lite.}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/