CrossViT: Cross-Attention Multi-Scale Vision Transformer for Image Classification
Chun-Fu ChenQuanfu FanRameswar Panda
Introduces CrossViT, a dual-branch vision transformer that extracts multi-scale features from different patch sizes and fuses them using a linear-time cross-attention mechanism to achieve superior ImageNet classification accuracy over standard Vision Transformers with low computational overhead.
Visual recognition systems increasingly rely on vision transformers rather than conventional convolutional networks. However, standard vision transformers typically process images using single-sized patches, missing the multi-scale contextual information that made earlier vision architectures successful. Directly combining multiple patch granularities often creates prohibitive computational and memory bottlenecks due to the quadratic complexity of standard attention mechanisms.
To address this limitation, the article introduces CrossViT, a dual-branch vision transformer architecture designed to learn multi-scale visual representations. The objective was to demonstrate that combining coarse-grained and fine-grained image patch streams via an efficient cross-attention mechanism achieves higher classification accuracy without substantial increases in computational overhead.
Researchers evaluated the approach primarily on the ImageNet benchmark, testing several model configurations varying in depth, embedding dimensions, and patch embedding methods. The architecture uses a wider, deeper primary branch for coarse patches and a lighter, narrower complementary branch for fine patches. Rather than calculating full pairwise attention across all tokens, the system uses the summary classification token of each branch as an agent to exchange information with the other branch, reducing attention complexity from quadratic to linear time. The models were also tested across five downstream transfer learning tasks, including medical and fine-grained image datasets.
Key findings confirm that this dual-branch cross-attention strategy provides consistent accuracy gains across various model sizes. On the standard ImageNet benchmark, CrossViT models outperformed baseline models like DeiT by up to 1.2 to 2.0 percentage points with modest parameter increases. When incorporating convolutional patch embeddings, performance reached 82.8% to 84.1% top-1 accuracy at higher resolutions. Notably, the CrossViT-18 configuration achieved 82.8% accuracy while cutting floating-point operations and parameter counts nearly in half compared to base baseline models. Across downstream transfer tasks, the architecture retained competitive generalization without overfitting.
These results indicate that multi-scale representation is highly practical for vision transformers when information exchange is restricted to summary tokens. Organizations deploying vision models can achieve superior accuracy and throughput compared to traditional convolutional networks and earlier transformer models, lowering operational inference costs for demanding image recognition tasks.
Decision-makers should consider adopting dual-branch transformer architectures for production computer vision pipelines where fine detail and broad context are both essential. For maximal accuracy, practitioners should pair the architecture with convolutional embedding tokenizers. While the current study validates performance in general image classification, future work should extend and pilot these multi-scale mechanisms in dense prediction tasks, such as object detection, semantic segmentation, and video understanding.
- Paper: An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale, Alexey Dosovitskiy et al. (2021). Read the original Vision Transformer first to understand the patch-token and class-token architecture that CrossViT adapts into two interacting branches.
- Paper: Training data-efficient image transformers & distillation through attention, Hugo Touvron et al. (2021). DeiT provides the data-efficient ViT training baseline that CrossViT compares against, clarifying the classification framework its multi-scale design seeks to improve.
- Paper: CrossFormer: A Versatile Vision Transformer Hinging on Cross-scale Attention, Wenxiao Wang et al. (2022). CrossFormer carries multi-scale transformer design further with explicit cross-scale attention, extending CrossViT’s central effort to connect representations at different patch granularities.
- Paper: Multi-Scale High-Resolution Vision Transformer for Semantic Segmentation, Jiaqi Gu et al. (2022). HRViT extends multi-scale transformer interaction to high-resolution dense prediction, taking the kind of cross-scale representation learning explored by CrossViT into segmentation.
