Built independently by an author, for readers. Read the story and support ChapterPal

keyword

multiscale vision transformers

Multiscale Vision Transformers are deep learning neural network architectures designed for computer vision tasks, such as image classification, object detection, and video recognition, that process visual data across multiple hierarchical scales. Unlike standard vision transformers that maintain a uniform token resolution and channel dimension throughout the entire network, multiscale vision transformers progressively reduce the spatial or spatiotemporal resolution of tokens while expanding the channel capacity across successive stages. This design creates a hierarchical feature pyramid where earlier layers model fine-grained, low-level details at high spatial resolution, while deeper layers capture complex, high-level semantic concepts at coarser resolutions, optimizing both representation capacity and computational efficiency.

2 items

MViTv2: Improved Multiscale Vision Transformers for Classification and Detection

MViTv2: Improved Multiscale Vision Transformers for Classification and Detection

Yanghao Li, Chao-Yuan Wu, Haoqi Fan, Karttikeya Mangalam, Bo Xiong, Jitendra Malik, Christoph Feichtenhofer

OrganizationsMetaUniversity of California Berkeley

Why you should read this

Proposes an improved Multiscale Vision Transformer architecture that integrates decomposed relative positional embeddings and residual pooling connections to establish state-of-the-art performance across image classification, object detection, and video recognition.

In this paper, we study Multiscale Vision Transformers (MViTv2) as a unified architecture for image and video classification, as well as object detection. We present an improved version of MViT that incorporates decomposed relative positional embeddings and residual pooling connections. We instantiate this architecture in five sizes and evaluate it for ImageNet classification, COCO detection and Kinetics video recognition where it outperforms prior work. We further compare MViTv2s' pooling attention to window attention mechanisms where it outperforms the latter in accuracy/compute. Without bells-and-whistles, MViTv2 has state-of-the-art performance in 3 domains: 88.8% accuracy on ImageNet classification, 58.7 AP^box on COCO object detection as well as 86.1% on Kinetics-400 video classification. Code and models are available at https://github.com/facebookresearch/mvit.

Added

2026-10-05

Multiscale Vision Transformers

Multiscale Vision Transformers

Haoqi Fan, Bo Xiong, Karttikeya Mangalam, Yanghao Li, Zhicheng Yan, Jitendra Malik, Christoph Feichtenhofer

OrganizationsMetaUniversity of California Berkeley

Why you should read this

Develops Multiscale Vision Transformers, a hierarchical architecture that incorporates multiscale feature pyramids into visual attention to achieve superior video and image recognition performance with up to ten times less computation and without requiring massive external pre-training.

We present Multiscale Vision Transformers (MViT) for video and image recognition, by connecting the seminal idea of multiscale feature hierarchies with transformer models. Multiscale Transformers have several channel-resolution scale stages. Starting from the input resolution and a small channel dimension, the stages hierarchically expand the channel capacity while reducing the spatial resolution. This creates a multiscale pyramid of features with early layers operating at high spatial resolution to model simple low-level visual information, and deeper layers at spatially coarse, but complex, high-dimensional features. We evaluate this fundamental architectural prior for modeling the dense nature of visual signals for a variety of video recognition tasks where it outperforms concurrent vision transformers that rely on large scale external pre-training and are 5-10x more costly in computation and parameters. We further remove the temporal dimension and apply our model for image classification where it outperforms prior work on vision transformers. Code is available at: this https URL

Added

2026-09-24