MViTv2: Improved Multiscale Vision Transformers for Classification and Detection
Yanghao LiChao-Yuan WuHaoqi FanKarttikeya MangalamBo XiongJitendra MalikChristoph Feichtenhofer
Proposes an improved Multiscale Vision Transformer architecture that integrates decomposed relative positional embeddings and residual pooling connections to establish state-of-the-art performance across image classification, object detection, and video recognition.
Modern computer vision systems increasingly rely on transformer-based architectures due to their strong performance across various visual tasks. However, applying transformers to high-resolution images, dense object detection, and complex video recognition presents severe computational and memory bottlenecks because standard attention mechanisms scale quadratically with the volume of visual data. To resolve these efficiency challenges, the article develops and evaluates an improved, unified architecture named Multiscale Vision Transformers version 2 (MViTv2) capable of serving as a general-purpose vision backbone across image classification, object detection, and video classification.
The authors approach this challenge by introducing two structural enhancements to the baseline multiscale transformer: decomposed relative positional embeddings, which enforce shift-invariance along separate spatial and temporal axes with low computational overhead, and residual pooling connections, which maintain rich information flow during feature downsampling. The architecture is instantiated across five capacity sizes (Tiny, Small, Base, Large, and Huge) and tested against standard benchmarks, including ImageNet-1K/21K for image classification, MS-COCO for object detection and instance segmentation, and Kinetics (400, 600, 700) alongside Something-Something-v2 for video recognition.
The experimental findings demonstrate state-of-the-art performance across all evaluated domains while maintaining superior computational efficiency. On ImageNet-1K, the largest MViTv2 model achieves up to 88.8% top-1 accuracy when pre-trained on ImageNet-21K, and 86.3% when trained entirely from scratch. On the COCO benchmark, MViTv2 combined with Cascade Mask R-CNN achieves an object detection score of 58.7 box Average Precision, surpassing competing architectures like Swin Transformers while utilizing fewer computational resources. On video benchmarks, MViTv2 establishes top-tier accuracy across datasets, reaching 86.1% on Kinetics-400, 87.9% on Kinetics-600, 79.4% on Kinetics-700, and 73.3% on Something-Something-v2.
These results indicate that pooling attention, supplemented by hybrid window attention for dense prediction, provides a more effective accuracy-to-compute tradeoff than standard local windowing mechanisms. For organizations deploying vision AI, adopting a unified architecture across 2D and 3D visual tasks can substantially streamline model development pipelines, reduce training and inference infrastructure costs, and mitigate memory bottlenecks on edge and server hardware. Organizations should consider transitioning to MViTv2 backbones for unified computer vision workloads and run targeted internal pilot evaluations comparing throughput, latency, and resource utilization on specific production hardware.
- Paper: Multiscale Vision Transformers, Haoqi Fan et al. (2021). Read the original MViT first to see the multiscale pooling-attention backbone that MViTv2 directly improves with positional embeddings and residual pooling.
- Paper: An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale, Alexey Dosovitskiy et al. (2021). ViT establishes the patch-token transformer baseline whose limitations in handling multiscale visual features motivate MViT and its MViTv2 refinement.
No sufficiently relevant recommendations were found.
