MViTv2: Improved Multiscale Vision Transformers for Classification and Detection

Yanghao LiChao-Yuan WuHaoqi FanKarttikeya MangalamBo XiongJitendra MalikChristoph Feichtenhofer

article2022CVPR997 citations

Proposes an improved Multiscale Vision Transformer architecture that integrates decomposed relative positional embeddings and residual pooling connections to establish state-of-the-art performance across image classification, object detection, and video recognition.

Listen

Modern computer vision systems increasingly rely on transformer-based architectures due to their strong performance across various visual tasks. However, applying transformers to high-resolution images, dense object detection, and complex video recognition presents severe computational and memory bottlenecks because standard attention mechanisms scale quadratically with the volume of visual data. To resolve these efficiency challenges, the article develops and evaluates an improved, unified architecture named Multiscale Vision Transformers version 2 (MViTv2) capable of serving as a general-purpose vision backbone across image classification, object detection, and video classification.

The authors approach this challenge by introducing two structural enhancements to the baseline multiscale transformer: decomposed relative positional embeddings, which enforce shift-invariance along separate spatial and temporal axes with low computational overhead, and residual pooling connections, which maintain rich information flow during feature downsampling. The architecture is instantiated across five capacity sizes (Tiny, Small, Base, Large, and Huge) and tested against standard benchmarks, including ImageNet-1K/21K for image classification, MS-COCO for object detection and instance segmentation, and Kinetics (400, 600, 700) alongside Something-Something-v2 for video recognition.

The experimental findings demonstrate state-of-the-art performance across all evaluated domains while maintaining superior computational efficiency. On ImageNet-1K, the largest MViTv2 model achieves up to 88.8% top-1 accuracy when pre-trained on ImageNet-21K, and 86.3% when trained entirely from scratch. On the COCO benchmark, MViTv2 combined with Cascade Mask R-CNN achieves an object detection score of 58.7 box Average Precision, surpassing competing architectures like Swin Transformers while utilizing fewer computational resources. On video benchmarks, MViTv2 establishes top-tier accuracy across datasets, reaching 86.1% on Kinetics-400, 87.9% on Kinetics-600, 79.4% on Kinetics-700, and 73.3% on Something-Something-v2.

These results indicate that pooling attention, supplemented by hybrid window attention for dense prediction, provides a more effective accuracy-to-compute tradeoff than standard local windowing mechanisms. For organizations deploying vision AI, adopting a unified architecture across 2D and 3D visual tasks can substantially streamline model development pipelines, reduce training and inference infrastructure costs, and mitigate memory bottlenecks on edge and server hardware. Organizations should consider transitioning to MViTv2 backbones for unified computer vision workloads and run targeted internal pilot evaluations comparing throughput, latency, and resource utilization on specific production hardware.

No sufficiently relevant recommendations were found.

Cover for MViTv2: Improved Multiscale Vision Transformers for Classification and Detection

Abstract

In this paper, we study Multiscale Vision Transformers (MViTv2) as a unified architecture for image and video classification, as well as object detection. We present an improved version of MViT that incorporates decomposed relative positional embeddings and residual pooling connections. We instantiate this architecture in five sizes and evaluate it for ImageNet classification, COCO detection and Kinetics video recognition where it outperforms prior work. We further compare MViTv2s' pooling attention to window attention mechanisms where it outperforms the latter in accuracy/compute. Without bells-and-whistles, MViTv2 has state-of-the-art performance in 3 domains: 88.8% accuracy on ImageNet classification, 58.7 AP^box on COCO object detection as well as 86.1% on Kinetics-400 video classification. Code and models are available at https://github.com/facebookresearch/mvit.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Revisiting Multiscale Vision Transformers
  • 4 Improved Multiscale Vision Transformers
  • 4.1 Improved Pooling Attention
  • 4.2 MViT for Object Detection
  • 4.3 MViT for Video Recognition
  • 4.4 MViT Architecture Variants
  • 5 Experiments: Image Recognition
  • 5.1 Image Classification on ImageNet-1K
  • 5.2 Object Detection on COCO
  • 5.3 Ablations on ImageNet and COCO
  • 6 Experiments: Video Recognition
  • 6.1 Main Results
  • 6.2 Ablations on Kinetics
  • 7 Conclusion
  • A Additional Results
  • A.1 Results: COCO Object Detection
  • A.2 Results: AVA Action Detection
  • A.3 Results: ImageNet Classification
  • A.4 Ablations: ImageNet and COCO
  • A.5 Ablations: Kinetics Action Classification
  • B Additional Implementation Details
  • B.1 Other Upgrades in MViT
  • B.2 Details: ImageNet Classification
  • B.3 Details: COCO Object Detection
  • B.4 Details: Kinetics Action Classification
  • B.5 Details: Something-Something V2 (SSv2)
  • B.6 Details: AVA Action Detection
  • C Additional Discussions
  • References

Knowls

  1. Knowl 1 — Pooling attention creates a hierarchical transformer through token reduction

    model/method

    For an input token sequence X∈RL×DX\in\mathbb{R}^{L\times D} with LL tokens and channel width DD, pooling attention projects the sequence into queries, keys, and values, then locally pools each projected sequence: Q=PQ(XWQ)Q=P_Q(XW_Q), K=PK(XWK)K=P_K(XW_K), and V=PV(XWV)V=P_V(XW_V). Here WQ,WK,WV∈RD×DW_Q,W_K,W_V\in\mathbb{R}^{D\times D} are learned projections and PQ,PK,PVP_Q,P_K,P_V are pooling operators. Attention is Z=Softmax⁡(QK⊤/D)VZ=\operatorname{Softmax}(QK^\top/\sqrt{D})V; the output has the query sequence length. Pooling queries can reduce resolution between stages, while pooling keys and values reduces the cost of global attention; the three pooling factors need not be identical. MViT uses this mechanism to form a hierarchy that goes from high-resolution, lower-width features to lower-resolution, wider features. Its usual configuration pools keys and values in pooling-attention blocks, with stride 4 in the first stage and strides that decrease adaptively in later stages.

  2. Knowl 2 — Decomposed relative position embeddings encode shift-invariant spatial relations

    model/method

    MViTv2 adds a relative-position term to pooled self-attention logits. For query token ii and key token jj, let Qi∈RdQ_i\in\mathbb{R}^{d} be the query vector, let p(i)p(i) and p(j)p(j) be their coordinates on a shared spatial or spatiotemporal grid, and let Rp(i),p(j)∈RdR_{p(i),p(j)}\in\mathbb{R}^{d} be a learned relative-position vector. The attention logit bias is Eij(rel)=Qi⊤Rp(i),p(j)E^{(\mathrm{rel})}_{ij}=Q_i^\top R_{p(i),p(j)}, giving Attn⁡(Q,K,V)=Softmax⁡((QK⊤+E(rel))/d)V\operatorname{Attn}(Q,K,V)=\operatorname{Softmax}((QK^\top+E^{(\mathrm{rel})})/\sqrt{d})V, where dd is the query/key feature dimension. Rather than learn a separate embedding for every joint offset, MViTv2 decomposes the vector as Rp(i),p(j)=Rh(i),h(j)h+Rw(i),w(j)w+Rt(i),t(j)tR_{p(i),p(j)}=R^{\mathrm{h}}_{h(i),h(j)}+R^{\mathrm{w}}_{w(i),w(j)}+R^{\mathrm{t}}_{t(i),t(j)}. The coordinates h,w,th,w,t are the vertical, horizontal, and temporal positions; the temporal term is optional for images. The decomposition reduces the number of learned embeddings from a dependence on the joint spatiotemporal grid, stated as O(TWH)O(TWH), to O(T+W+H)O(T+W+H), where T,W,HT,W,H are the grid extents. Because the bias depends on relative rather than absolute location, it supplies a shift-invariance prior. The page-3 architecture schematic depicts this relative-position contribution within the attention computation.

  3. Knowl 3 — Residual pooling adds the pooled query to the attention output

    model/method

    MViTv2 places a residual path inside each pooling-attention operation: Z=Attn⁡(Q,K,V)+QZ=\operatorname{Attn}(Q,K,V)+Q. The pooled query QQ and attention output ZZ have the same sequence length, so the addition preserves the output shape. The page-3 architecture schematic depicts the pooled-query skip path alongside the attention computation. In an MViTv2-S ablation, adding the residual path changed ImageNet-1K top-1 accuracy from 83.3% to 83.6% and COCO box AP from 48.5 to 49.3. Applying query pooling in all other layers as well as the residual path kept ImageNet accuracy at 83.6% and raised COCO box AP to 49.9; adding query pooling without the residual path yielded 83.1% and 48.5, respectively. These results support the authors’ finding that query pooling and its residual path work together.

  4. Knowl 4 — Hybrid window attention restores cross-window information with global blocks

    model/method

    Hybrid window attention (Hwin) performs local attention within windows in all blocks except the last blocks of the final three stages that feed the feature pyramid network (FPN). Those final blocks use global attention, providing cross-window information to the FPN input without making every block globally attentive. On ImageNet-1K with a ViT-B backbone, Hwin achieved 82.1% top-1 accuracy, compared with 80.0% for fixed non-overlapping windows and 80.4% for shifted-window attention; full attention achieved 82.0%. On COCO with ViT-B and Mask R-CNN, Hwin reached 46.1 box AP, compared with 45.1 for shifted windows and 46.6 for full attention. For MViTv2-S on COCO, combining pooling attention with Hwin gave 49.9 box AP, 9.4 test images/s, and 5.2 GB peak training memory, versus 50.8 AP, 4.2 images/s, and 19.5 GB for pooling alone. Thus, the combination traded a small amount of AP for higher throughput and lower memory in that comparison.

  5. Knowl 5 — Five MViTv2 sizes scale stage width, depth, and attention heads

    experimental setup

    The image-classification configurations use four stages with resolutions 562,282,142,7256^2,28^2,14^2,7^2 for a 224×224224\times224 input. Each configuration below lists stage channel widths, blocks per stage, and heads per stage, followed by image-classification compute and parameter count: Tiny uses channels [96,192,384,768], blocks [1,2,5,2], heads [1,2,4,8], 4.7 GFLOPs, and 24 million parameters; Small uses [96,192,384,768], [1,2,11,2], [1,2,4,8], 7.0 GFLOPs, and 35 million; Base uses [96,192,384,768], [2,3,16,3], [1,2,4,8], 10.2 GFLOPs, and 52 million; Large uses [144,288,576,1152], [2,6,36,4], [2,4,8,16], 39.6 GFLOPs, and 218 million; Huge uses [192,384,768,1536], [4,8,60,8], [3,6,12,24], 120.6 GFLOPs, and 667 million. The authors chose relatively few heads partly to improve runtime, noting that adding heads does not change FLOPs or parameter count.

  6. Knowl 6 — MViT stage features integrate with FPN for dense prediction

    model/method

    For object detection and instance segmentation, the four-stage MViT backbone supplies feature maps at multiple resolutions to a standard top-down FPN with lateral connections. Those pyramid features can then drive detectors such as Mask R-CNN or Cascade Mask R-CNN. The page-4 backbone illustration shows stage features entering the FPN at multiple scales before prediction heads, reflecting how the hierarchical backbone is used for dense prediction. For detection inputs of varying size, the authors initialize positional embeddings from ImageNet-pretrained weights for a 224×224224\times224 input and interpolate them to the required sizes. In their COCO setup, MViTv2 backbones were initialized from ImageNet pretraining and used Hwin by default, with window sizes [56, 28, 14, 7] across the four stages.

  7. Knowl 7 — Image-to-video initialization inflates convolutional weights and separates temporal positions

    model/method

    To initialize a video MViT from an image-pretrained model, the patchification stem and pooling operators are extended from 2D to space-time convolutions. The filters for the center frame are initialized with the corresponding 2D convolution weights, and weights for the other temporal positions are set to zero. For decomposed relative-position embeddings, spatial embeddings are copied from the image model and the temporal embedding is initialized to zero. This lets the video model begin with the pretrained spatial parameters while adding temporal processing.

  8. Knowl 8 — MViTv2 reaches high ImageNet-1K accuracy with and without ImageNet-21K pretraining

    empirical result

    The ImageNet-1K experiments trained MViTv2 variants for 300 epochs without EMA. Without external pretraining, MViTv2-B achieved 84.4% top-1 accuracy at 224×224224\times224 input with 10.2 GFLOPs and 52 million parameters. At 384×384384\times384, MViTv2-L achieved 86.0% with center-crop evaluation and 86.3% when evaluating a resized full-image view; it used 140.2 GFLOPs and 218 million parameters. With ImageNet-21K pretraining, MViTv2-L at 384×384384\times384 reached 88.2% with center-crop evaluation and 88.4% with the resized full-image view. MViTv2-H at 512×512512\times512 reached 88.3% and 88.8% under those respective evaluation protocols, using 763.5 GFLOPs and 667 million parameters. The paper reports that ImageNet-21K pretraining added 2.2 percentage points to MViTv2-L.

  9. Knowl 9 — MViTv2 improves COCO detection and instance-segmentation accuracy across detector scales

    empirical result

    The COCO experiments used 118,000 training images and 5,000 validation images, with ImageNet-pretrained backbones and a default 36-epoch fine-tuning schedule. With Mask R-CNN, MViTv2-T, -S, -B, and -L achieved, respectively, box AP/mask AP of 48.2/43.8, 49.9/45.1, 51.0/45.7, and 51.8/46.2. The ImageNet-21K-pretrained MViTv2-L reached 52.7/46.8. In the Base comparison, MViTv2-B used 392 GFLOPs and 71 million parameters, versus 496 GFLOPs and 107 million for Swin-B, while reaching 51.0 rather than 48.5 box AP and 45.7 rather than 43.4 mask AP.

    With Cascade Mask R-CNN, MViTv2-T, -S, and -B achieved box AP/mask AP of 52.2/45.0, 53.2/46.0, and 54.1/46.8; ImageNet-21K-pretrained MViTv2-B reached 54.9/47.4. A longer 50-epoch schedule with stronger large-scale jitter and ImageNet-21K pretraining produced 55.8/48.3 for MViTv2-L and 56.1/48.5 for MViTv2-H. Adding Soft-NMS and multiscale testing to the MViTv2-L Cascade Mask R-CNN system yielded 58.7 box AP and 50.5 mask AP. The gains across the backbone sizes show that the multiscale transformer is effective with both detection and instance-segmentation heads.

  10. Knowl 10 — MViTv2 achieves strong video recognition results across four benchmarks

    empirical result

    In the reported video comparisons, a clip specification T×τT\times\tau denotes TT frames sampled at temporal stride τ\tau. On Kinetics-400, MViTv2-B trained from scratch with 32×332\times3 clips reached 82.9% top-1 and 95.7% top-5 accuracy; ImageNet-21K-pretrained MViTv2-L with 40×340\times3 clips and 3122312^2 spatial input reached 86.1% and 97.0%. On Kinetics-600, the scratch-trained MViTv2-B with 32×332\times3 clips reached 85.5%/97.2%, and ImageNet-21K-pretrained MViTv2-L with 40×340\times3 clips and 3522352^2 input reached 87.9%/97.9%. On Kinetics-700, MViTv2-L with ImageNet-21K pretraining and 40×340\times3 clips reached 79.4% top-1 and 94.9% top-5. On Something-Something-v2, MViTv2-S with 16×416\times4 clips achieved 68.2% top-1; MViTv2-B achieved 70.5%, rising to 72.1% with ImageNet-21K and Kinetics-400 pretraining, while MViTv2-L reached 73.3%. The paper also reports that MViTv2 improves over its MViTv1 counterparts by 2.6 and 2.7 percentage points on Kinetics-400 for the Small and Base models, respectively, and by 1.4 points for the Base model on Kinetics-600.

    The authors’ Kinetics-400 pretraining ablation used 1 spatial view by 10 temporal views: MViTv2-S with 16×416\times4 clips scored 81.2% from scratch, 82.2% with ImageNet-1K pretraining, and 82.6% with ImageNet-21K pretraining; MViTv2-B with 32×332\times3 clips scored 82.9%, 83.3%, and 84.3%, respectively. For the larger MViTv2-L, the corresponding scores were 81.4%, 83.4%, and 84.5% with 40×340\times3 clips, and 81.8%, 84.4%, and 85.7% with 40×340\times3 clips at 3122312^2 input. These results show that ImageNet pretraining particularly benefits the larger models.

Coverage note — Detailed per-task augmentation and optimizer schedules are omitted because the paper refers to supplementary appendices for those implementation details, which are not included in the supplied pages.

References

  1. 1.Anurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun, Mario Lucić, and Cordelia Schmid. Vivit: A video vision transformer. arXiv preprint arXiv:2103.15691, 2021. 1, 2, 8
  2. 2.Josh Beal, Eric Kim, Eric Tzeng, Dong Huk Park, Andrew Zhai, and Dmitry Kislyuk. Toward transformer-based object detection. arXiv preprint arXiv:2012.09958, 2020. 1
  3. 3.Gedas Bertasius, Heng Wang, and Lorenzo Torresani. Is space-time attention all you need for video understanding? arXiv preprint arXiv:2102.05095, 2021. 2, 8
  4. 4.Navaneeth Bodla, Bharat Singh, Rama Chellappa, and Larry S Davis. Soft-nms–improving object detection with one line of code. In Proc. ICCV, 2017. 6, 12
  5. 5.Andrew Brock, Soham De, Samuel L Smith, and Karen Simonyan. High-performance large-scale image recognition without normalization. arXiv preprint arXiv:2102.06171, 2021. 5, 13
  6. 6.Zhaowei Cai and Nuno Vasconcelos. Cascade r-cnn: Delving into high quality object detection. In Proc. CVPR, 2018. 2, 5, 6, 12, 15
  7. 7.Joao Carreira, Eric Noland, Chloe Hillier, and Andrew Zisserman. A short note on the kinetics-700 human action dataset. arXiv preprint arXiv:1907.06987, 2019. 7
  8. 8.Joao Carreira and Andrew Zisserman. Quo vadis, action recognition? a new model and the kinetics dataset. In Proc. CVPR, 2017. 2, 4, 7
  9. 9.Chun-Fu Chen, Quanfu Fan, and Rameswar Panda. Crossvit: Cross-attention multi-scale vision transformer for image classification. In Proc. ICCV, 2021. 4, 13
  10. 10.Yunpeng Chen, Haoqi Fang, Bing Xu, Zhicheng Yan, Yannis Kalantidis, Marcus Rohrbach, Shuicheng Yan, and Jiashi Feng. Drop an octave: Reducing spatial redundancy in convolutional neural networks with octave convolution. arXiv preprint arXiv:1904.05049, 2019. 2
  11. 11.Xiangxiang Chu, Zhi Tian, Yuqing Wang, Bo Zhang, Haibing Ren, Xiaolin Wei, Huaxia Xia, and Chunhua Shen. Twins: Revisiting the design of spatial attention in vision transformers. In NIPS, 2021. 2, 5, 13
  12. 12.Ekin D Cubuk, Barret Zoph, Jonathon Shlens, and Quoc V Le. Randaugment: Practical automated data augmentation with a reduced search space. In Proc. CVPR, 2020. 14, 15
  13. 13.Zihang Dai, Hanxiao Liu, Quoc V Le, and Mingxing Tan. Coatnet: Marrying convolution and attention for all data sizes. arXiv preprint arXiv:2106.04803, 2021. 5, 13
  14. 14.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In Proc. CVPR, pages 248–255. Ieee, 2009. 2, 4
  15. 15.Piotr Dollar, Mannat Singh, and Ross Girshick. Fast and accurate model scaling. In Proc. CVPR, 2021. 2, 5, 13
  16. 16.Xiaoyi Dong, Jianmin Bao, Dongdong Chen, Weiming Zhang, Nenghai Yu, Lu Yuan, Dong Chen, and Baining Guo. Cswin transformer: A general vision transformer backbone with cross-shaped windows. arXiv preprint arXiv:2107.00652, 2021. 5, 13
  17. 17.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929, 2020. 1, 2, 5
  18. 18.Alaaeldin El-Nouby, Hugo Touvron, Mathilde Caron, Piotr Bojanowski, Matthijs Douze, Armand Joulin, Ivan Laptev, Natalia Neverova, Gabriel Synnaeve, Jakob Verbeek, et al. Xcit: Cross-covariance image transformers. arXiv preprint arXiv:2106.09681, 2021. 5, 13
  19. 19.Haoqi Fan, Yanghao Li, Bo Xiong, Wan-Yen Lo, and Christoph Feichtenhofer. PySlowFast. https://github.com/facebookresearch/slowfast, 2020. 2, 7, 15
  20. 20.Haoqi Fan, Tullie Murrell, Heng Wang, Kalyan Vasudev Alwala, Yanghao Li, Yilei Li, Bo Xiong, Nikhila Ravi, Meng Li, Haichuan Yang, Jitendra Malik, Ross Girshick, Matt Feiszli, Aaron Adcock, Wan-Yen Lo, and Christoph Feichtenhofer. PyTorchVideo: A deep learning library for video understanding. In Proceedings of the 29th ACM International Conference on Multimedia, 2021. https://pytorchvideo.org/. 2
  21. 21.Haoqi Fan, Bo Xiong, Karttikeya Mangalam, Yanghao Li, Zhicheng Yan, Jitendra Malik, and Christoph Feichtenhofer. Multiscale vision transformers. In Proc. ICCV, 2021. 1, 2, 3, 4, 5, 6, 7, 8, 12, 13, 14, 15, 16
  22. 22.Christoph Feichtenhofer. X3D: Expanding architectures for efficient video recognition. In Proc. CVPR, pages 203–213, 2020. 2, 8, 12
  23. 23.Christoph Feichtenhofer, Haoqi Fan, Jitendra Malik, and Kaiming He. SlowFast networks for video recognition. In Proc. ICCV, 2019. 2, 7, 8, 12, 15
  24. 24.Christoph Feichtenhofer, Axel Pinz, and Richard Wildes. Spatiotemporal residual networks for video action recognition. In NIPS, 2016. 4
  25. 25.Christoph Feichtenhofer, Axel Pinz, and Andrew Zisserman. Convolutional two-stream network fusion for video action recognition. In Proc. CVPR, 2016. 2
  26. 26.Golnaz Ghiasi, Yin Cui, Aravind Srinivas, Rui Qian, Tsung-Yi Lin, Ekin D Cubuk, Quoc V Le, and Barret Zoph. Simple copy-paste is a strong data augmentation method for instance segmentation. In Proc. CVPR, 2021. 6, 12, 15
  27. 27.Golnaz Ghiasi, Tsung-Yi Lin, and Quoc V Le. Nas-fpn: Learning scalable feature pyramid architecture for object detection. In Proc. CVPR, 2019. 12
  28. 28.Rohit Girdhar, Joao Carreira, Carl Doersch, and Andrew Zisserman. Video action transformer network. In Proc. CVPR, 2019. 2
  29. 29.Ross Girshick. Fast R-CNN. In Proc. ICCV, 2015. 2, 15
  30. 30.Priya Goyal, Piotr Dollar, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He. Accurate, large minibatch SGD: training ImageNet in 1 hour. arXiv:1706.02677, 2017. 15
  31. 31.Raghav Goyal, Samira Ebrahimi Kahou, Vincent Michalski, Joanna Materzynska, Susanne Westphal, Heuna Kim, Valentin Haenel, Ingo Fruend, Peter Yianilos, Moritz Mueller-Freitag, et al. The “Something Something” video database for learning and evaluating visual common sense. In ICCV, 2017. 7, 15
  32. 32.Chunhui Gu, Chen Sun, David A. Ross, Carl Vondrick, Caroline Pantofaru, Yeqing Li, Sudheendra Vijayanarasimhan, George Toderici, Susanna Ricco, Rahul Sukthankar, Cordelia Schmid, and Jitendra Malik. AVA: A video dataset of spatiotemporally localized atomic visual actions. In Proc. CVPR, 2018. 12, 15
  33. 33.Kai Han, An Xiao, Enhua Wu, Jianyuan Guo, Chunjing Xu, and Yunhe Wang. Transformer in transformer. In NIPS, 2021. 5, 13
  34. 34.Zhang Hang, Chongruo Wu, Zhongyue Zhang, Yi Zhu, Zhi Zhang, Haibin Lin, and Yue Sun. Resnest: Split-attention networks. 2020. 2
  35. 35.Boris Hanin and David Rolnick. How to start training: The effect of initialization and architecture. arXiv preprint arXiv:1803.01719, 2018. 14
  36. 36.Kaiming He, Georgia Gkioxari, Piotr Dollar, and Ross Girshick. Mask R-CNN. In Proc. ICCV, 2017. 1, 3, 5, 6, 14, 15
  37. 37.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Delving deep into rectifiers: Surpassing human-level performance on imagenet classification. In Proc. CVPR, 2015. 1
  38. 38.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proc. CVPR, 2016. 2, 6
  39. 39.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Identity mappings in deep residual networks. In Proc. ECCV, 2016. 2
  40. 40.Elad Hoffer, Tal Ben-Nun, Itay Hubara, Niv Giladi, Torsten Hoefler, and Daniel Soudry. Augment your batch: Improving generalization through instance repetition. In Proc. CVPR, pages 8129–8138, 2020. 15
  41. 41.Gao Huang, Yu Sun, Zhuang Liu, Daniel Sedra, and Kilian Q Weinberger. Deep networks with stochastic depth. In Proc. ECCV, 2016. 14
  42. 42.Boyuan Jiang, MengMeng Wang, Weihao Gan, Wei Wu, and Junjie Yan. Stm: Spatiotemporal and motion encoding for action recognition. In Proc. CVPR, pages 2000–2009, 2019. 2
  43. 43.Zihang Jiang, Qibin Hou, Li Yuan, Daquan Zhou, Xiaojie Jin, Anran Wang, and Jiashi Feng. Token labeling: Training a 85.5% top-1 accuracy vision transformer with 56m parameters on imagenet. arXiv preprint arXiv:2104.10858, 2021. 5
  44. 44.Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, et al. The kinetics human action video dataset. arXiv:1705.06950, 2017. 2, 7
  45. 45.Dan Kondratyuk, Liangzhe Yuan, Yandong Li, Li Zhang, Mingxing Tan, Matthew Brown, and Boqing Gong. MoViNets: Mobile video networks for efficient video recognition. In Proc. CVPR, 2021. 8
  46. 46.Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. Imagenet classification with deep convolutional neural networks. In NIPS, 2012. 2
  47. 47.Yann LeCun, Bernhard Boser, John Denker, Donnie Henderson, Richard Howard, Wayne Hubbard, and Lawrence Jackel. Handwritten digit recognition with a back-propagation network. In NIPS, 1989. 3
  48. 48.Yann LeCun, Bernhard Boser, John S Denker, Donnie Henderson, Richard E Howard, Wayne Hubbard, and Lawrence D Jackel. Backpropagation applied to handwritten zip code recognition. Neural computation, 1(4):541–551, 1989. 2
  49. 49.Yan Li, Bin Ji, Xintian Shi, Jianguo Zhang, Bin Kang, and Limin Wang. Tea: Temporal excitation and aggregation for action recognition. In Proc. CVPR, pages 909–918, 2020. 8
  50. 50.Yanghao Li, Saining Xie, Xinlei Chen, Piotr Dollar, Kaiming He, and Ross Girshick. Benchmarking detection transfer learning with vision transformers. arXiv preprint arXiv:2111.11429, 2021. 7
  51. 51.Zhenyang Li, Kirill Gavrilyuk, Efstratios Gavves, Mihir Jain, and Cees GM Snoek. VideoLSTM convolves, attends and flows for action recognition. Computer Vision and Image Understanding, 166:41–50, 2018. 2
  52. 52.Ji Lin, Chuang Gan, and Song Han. Temporal shift module for efficient video understanding. In Proc. ICCV, 2019. 15
  53. 53.Tsung-Yi Lin, Piotr Dollar, Ross Girshick, Kaiming He, Bharath Hariharan, and Serge Belongie. Feature pyramid networks for object detection. In Proc. CVPR, 2017. 1, 2, 3
  54. 54.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollar, and C Lawrence Zitnick. Microsoft COCO: Common objects in context. In Proc. ECCV, 2014. 4, 5
  55. 55.Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. arXiv preprint arXiv:2103.14030, 2021. 1, 2, 3, 4, 5, 6, 7, 12, 13, 15, 16
  56. 56.Ze Liu, Jia Ning, Yue Cao, Yixuan Wei, Zheng Zhang, Stephen Lin, and Han Hu. Video swin transformer. arXiv preprint arXiv:2106.13230, 2021. 2, 8
  57. 57.Ilya Loshchilov and Frank Hutter. SGDR: Stochastic gradient descent with warm restarts. arXiv:1608.03983, 2016. 15
  58. 58.Ilya Loshchilov and Frank Hutter. Fixing weight decay regularization in adam. 2018. 14, 15
  59. 59.Daniel Neimark, Omri Bar, Maya Zohar, and Dotan Asselmann. Video transformer network. arXiv preprint arXiv:2102.00719, 2021. 2, 8
  60. 60.Junting Pan, Siyu Chen, Mike Zheng Shou, Yu Liu, Jing Shao, and Hongsheng Li. Actor-context-actor relation network for spatio-temporal action localization. In Proc. CVPR, 2021. 12
  61. 61.Zhaofan Qiu, Ting Yao, and Tao Mei. Learning spatio-temporal representation with pseudo-3d residual networks. In Proc. ICCV, 2017. 2
  62. 62.Ilija Radosavovic, Raj Prateek Kosaraju, Ross Girshick, Kaiming He, and Piotr Dollar. Designing network design spaces. In Proc. CVPR, June 2020. 2, 5, 13
  63. 63.Joseph Redmon, Santosh Divvala, Ross Girshick, and Ali Farhadi. You only look once: Unified, real-time object detection. In Proc. CVPR, 2016. 2
  64. 64.Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. Faster R-CNN: Towards real-time object detection with region proposal networks. In NIPS, 2015. 15
  65. 65.Peter Shaw, Jakob Uszkoreit, and Ashish Vaswani. Self-attention with relative position representations. arXiv preprint arXiv:1803.02155, 2018. 3
  66. 66.Karen Simonyan and Andrew Zisserman. Two-stream convolutional networks for action recognition in videos. In NIPS, 2014. 2
  67. 67.Karen Simonyan and Andrew Zisserman. Very deep convolutional networks for large-scale image recognition. In Proc. ICLR, 2015. 1, 2
  68. 68.Robin Strudel, Ricardo Garcia, Ivan Laptev, and Cordelia Schmid. Segmenter: Transformer for semantic segmentation. arXiv preprint arXiv:2105.05633, 2021. 1
  69. 69.Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich. Going deeper with convolutions. In Proc. CVPR, 2015. 2, 15
  70. 70.Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jonathon Shlens, and Zbigniew Wojna. Rethinking the inception architecture for computer vision. arXiv:1512.00567, 2015. 14
  71. 71.Mingxing Tan and Quoc V Le. Efficientnet: Rethinking model scaling for convolutional neural networks. arXiv preprint arXiv:1905.11946, 2019. 2, 5, 13
  72. 72.Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Herve Jégou. Training data-efficient image transformers & distillation through attention. arXiv preprint arXiv:2012.12877, 2020. 4, 5, 13, 14
  73. 73.Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Herve Jégou. DeiT: Data-efficient image transformers. arXiv preprint arXiv:2012.12877, 2020. 1, 2, 14, 16
  74. 74.Hugo Touvron, Matthieu Cord, Alexandre Sablayrolles, Gabriel Synnaeve, and Herve Jégou. Going deeper with image transformers. arXiv preprint arXiv:2103.17239, 2021. 5, 13
  75. 75.Du Tran, Heng Wang, Lorenzo Torresani, and Matt Feiszli. Video classification with channel-separated convolutional networks. In Proc. ICCV, 2019. 2
  76. 76.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. arXiv preprint arXiv:1706.03762, 2017. 1, 2
  77. 77.Wenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan, Kaitao Song, Ding Liang, Tong Lu, Ping Luo, and Ling Shao. Pvtv2: Improved baselines with pyramid vision transformer. arXiv preprint arXiv:2106.13797, 2021. 5, 13
  78. 78.Wenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan, Kaitao Song, Ding Liang, Tong Lu, Ping Luo, and Ling Shao. Pyramid vision transformer: A versatile backbone for dense prediction without convolutions. In IEEE ICCV, 2021. 1, 2, 6, 13
  79. 79.Chao-Yuan Wu, Christoph Feichtenhofer, Haoqi Fan, Kaiming He, Philipp Krahenbühl, and Ross Girshick. Long-term feature banks for detailed video understanding. In Proc. CVPR, 2019. 2
  80. 80.Chao-Yuan Wu and Philipp Krahenbuhl. Towards long-form video understanding. In Proc. CVPR, 2021. 12
  81. 81.Haiping Wu, Bin Xiao, Noel Codella, Mengchen Liu, Xiyang Dai, Lu Yuan, and Lei Zhang. Cvt: Introducing convolutions to vision transformers. arXiv preprint arXiv:2103.15808, 2021. 4, 5, 13
  82. 82.Yuxin Wu, Alexander Kirillov, Francisco Massa, Wan-Yen Lo, and Ross Girshick. Detectron2. https://github.com/facebookresearch/detectron2, 2019. 5, 6, 15
  83. 83.Saining Xie, Ross Girshick, Piotr Dollar, Zhuowen Tu, and Kaiming He. Aggregated residual transformations for deep neural networks. In Proc. CVPR, 2017. 6
  84. 84.Saining Xie, Chen Sun, Jonathan Huang, Zhuowen Tu, and Kevin Murphy. Rethinking spatiotemporal feature learning for video understanding. arXiv:1712.04851, 2017. 2
  85. 85.Li Yuan, Yunpeng Chen, Tao Wang, Weihao Yu, Yujun Shi, Zi-Hang Jiang, Francis E.H. Tay, Jiashi Feng, and Shuicheng Yan. Tokens-to-token vit: Training vision transformers from scratch on imagenet. In Proc. ICCV, 2021. 13
  86. 86.Li Yuan, Qibin Hou, Zihang Jiang, Jiashi Feng, and Shuicheng Yan. Volo: Vision outlooker for visual recognition, 2021. 5
  87. 87.Sangdoo Yun, Dongyoon Han, Seong Joon Oh, Sanghyuk Chun, Junsuk Choe, and Youngjoon Yoo. Cutmix: Regularization strategy to train strong classifiers with localizable features. In Proc. ICCV, 2019. 14, 15
  88. 88.Hongyi Zhang, Moustapha Cisse, Yann N Dauphin, and David Lopez-Paz. Mixup: Beyond empirical risk minimization. In Proc. ICLR, 2018. 14, 15
  89. 89.Pengchuan Zhang, Xiyang Dai, Jianwei Yang, Bin Xiao, Lu Yuan, Lei Zhang, and Jianfeng Gao. Multi-scale vision long-former: A new vision transformer for high-resolution image encoding. In Proc. ICCV, 2021. 2, 5, 6, 13
  90. 90.Sixiao Zheng, Jiachen Lu, Hengshuang Zhao, Xiatian Zhu, Zekun Luo, Yabiao Wang, Yanwei Fu, Jianfeng Feng, Tao Xiang, Philip HS Torr, et al. Rethinking semantic segmentation from a sequence-to-sequence perspective with transformers. In Proc. CVPR, 2021. 1
  91. 91.Zhun Zhong, Liang Zheng, Guoliang Kang, Shaozi Li, and Yi Yang. Random erasing data augmentation. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, pages 13001–13008, 2020. 14, 15
  92. 92.Bolei Zhou, Alex Andonian, Aude Oliva, and Antonio Torralba. Temporal relational reasoning in videos. In ECCV, 2018. 2
  93. 93.Xingyi Zhou, Dequan Wang, and Philipp Krahenbühl. Objects as points. arXiv preprint arXiv:1904.07850, 2019. 2

Citation

MLA
Li, Y., et al. “MViTv2: Improved Multiscale Vision Transformers for Classification and Detection”. arXiv, 2021, http://arxiv.org/abs/2112.01526v2.
APA
Li, Y., Wu, C.-Y., Fan, H., Mangalam, K., Xiong, B., Malik, J., & Feichtenhofer, C. (2021). MViTv2: Improved Multiscale Vision Transformers for Classification and Detection. arXiv. http://arxiv.org/abs/2112.01526v2
Chicago
Li, Y., C.-Y. Wu, H. Fan, et al. 2021. “MViTv2: Improved Multiscale Vision Transformers for Classification and Detection”. arXiv. http://arxiv.org/abs/2112.01526v2.
Harvard
Li, Y. et al. (2021) “MViTv2: Improved Multiscale Vision Transformers for Classification and Detection”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2112.01526v2.
Vancouver
1. Li Y, Wu C-Y, Fan H, Mangalam K, Xiong B, Malik J, Feichtenhofer C (2021) MViTv2: Improved Multiscale Vision Transformers for Classification and Detection. arXiv

BibTeX

@article{li2021mvitv2,
  title = {MViTv2: Improved Multiscale Vision Transformers for Classification and Detection},
  author = {Li, Yanghao and Wu, Chao-Yuan and Fan, Haoqi and Mangalam, Karttikeya and Xiong, Bo and Malik, Jitendra and Feichtenhofer, Christoph},
  year = {2021},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2112.01526v2},
  eprint = {2112.01526}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE