CrossFormer: A Versatile Vision Transformer Hinging on Cross-scale Attention
Wenxiao WangLu YaoLong ChenBinbin LinDeng CaiXiaofei HeWei Liu
Proposes CrossFormer, a vision transformer that establishes cross-scale feature interactions using multi-scale patch embeddings and long-short distance attention to achieve superior performance across image classification, object detection, and segmentation benchmarks.
Modern computer vision systems rely heavily on attention mechanisms to capture relationships across visual scenes. However, existing vision architectures struggle to link visual features of differing scales—such as fine-grained details alongside broad background contexts—because they partition images into uniform, single-scale patches and compress representations to manage computational load. This limitation significantly hinders performance on complex, real-world visual perception tasks such as detecting objects and segmenting scenes.
The article develops and evaluates CrossFormer, a versatile vision architecture designed to establish cross-scale visual interactions while efficiently processing arbitrary input sizes. The design combines a cross-scale embedding layer that extracts multiple patch sizes concurrently, an alternating long- and short-distance attention mechanism that retains fine details without excessive computational expense, and a trainable dynamic position bias to accommodate variable input dimensions.
The authors conducted comprehensive empirical evaluations on standard computer vision benchmarks across four core tasks: image classification on ImageNet (1.28 million training images), object detection and instance segmentation on COCO 2017 (118,000 training images), and semantic segmentation on ADE20K (20,000 training images). CrossFormer variants ranging from tiny to large were compared against leading architectures under standardized training protocols.
The evaluation yielded several key findings:
- CrossFormer consistently outperformed competing architectures across all four evaluation benchmarks while maintaining comparable parameter counts and computational demands.
- The architecture demonstrated substantial gains in dense prediction tasks, achieving up to 1.7-point improvements in object detection precision on COCO and over 4-point gains in intersection-over-union for semantic segmentation on ADE20K compared to peer models.
- In image classification on ImageNet, CrossFormer variants achieved top-1 accuracies ranging from 81.5% to 84.0%, exceeding baseline models like PVT and Swin by at least 1.2% in small configurations.
- Ablation analyses confirmed that combining multi-scale patch extraction with long-short distance attention provided a 1.0% accuracy improvement over single-scale baselines, confirming the direct value of cross-scale interactions.
These findings demonstrate that explicitly enabling cross-scale interactions resolves a fundamental efficiency and accuracy trade-off in visual recognition models. For engineering and technology leaders, adopting architectures with dynamic position handling and cross-scale attention offers improved accuracy across multi-task visual pipelines—including dense pixel-level labeling—without requiring proportional increases in computational or memory budgets.
Organizations developing vision applications should evaluate CrossFormer backbones for downstream perception systems, particularly in dense detection and segmentation workloads where gains are most pronounced. System designers can select larger grouping parameters during deployment to reduce GPU memory footprint with negligible loss in detection accuracy.
The presented evaluations focus primarily on standard academic benchmark datasets and established model frameworks. Practical deployment in live production environments may require further testing to evaluate inference latency under specific edge-hardware constraints and performance on domain-specific data distributions.
- Paper: An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale, Alexey Dosovitskiy et al. (2021). Introduces the original Vision Transformer (ViT) architecture, establishing the foundational patch-embedding and self-attention paradigms that CrossFormer modifies.
- Paper: Swin Transformer: Hierarchical Vision Transformer using Shifted Windows, Ze Liu et al. (2021). Presents hierarchical vision transformers with local window self-attention and relative position bias, addressing computational limits that CrossFormer further improves via cross-scale attention.
- Paper: CrossViT: Cross-Attention Multi-Scale Vision Transformer for Image Classification, Chun-Fu Chen et al. (2021). Explores multi-scale patch tokenization and cross-attention between distinct branches, motivating CrossFormer's single-stream Cross-scale Embedding Layer.
- Paper: Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without Convolutions, Wenhai Wang et al. (2021). Develops a pyramid vision transformer backbone with spatial reduction attention for dense downstream tasks, laying groundwork for hierarchical feature extraction.
- Paper: Twins: Revisiting the Design of Spatial Attention in Vision Transformers, Xiangxiang Chu et al. (2021). Introduces spatially separable self-attention combining local grouped attention with global sub-sampled attention, directly preceding CrossFormer's Long Short Distance Attention scheme.
- Paper: Transformer in Transformer, Kai Han et al. (2021). Investigates nested, fine-grained sub-patch attention to capture intra-patch visual details, directly addressing multi-scale feature representation.
- Paper: CSWin Transformer: A General Vision Transformer Backbone with Cross-Shaped Windows, Xiaoyi Dong et al. (2021). Introduces cross-shaped window self-attention and locally-enhanced position encodings to broaden receptive fields efficiently across variable image resolutions.
- Paper: Multiscale Vision Transformers, Haoqi Fan et al. (2021). Proposes multiscale vision transformers with pooling attention to model feature hierarchies, providing conceptual context for multi-scale attention mechanisms.
- Paper: TransNeXt: Robust Foveal Visual Perception for Vision Transformers, Dai Shi (2024). Develops aggregated attention combining fine-grained local focus and coarse-grained global perception, building on cross-scale and multi-distance attention strategies.
- Paper: Swin Transformer V2: Scaling Up Capacity and Resolution, Ze Liu et al. (2022). Scales hierarchical vision transformers to high resolutions and large parameter regimes using continuous position bias that generalizes dynamic position encodings.
- Paper: Cross Aggregation Transformer for Image Restoration, Zheng Chen et al. (2022). Extends efficient multi-axis and cross-window attention mechanisms to dense pixel-level image restoration tasks.
- Paper: SegNeXt: Rethinking Convolutional Attention Design for Semantic Segmentation, Meng-Hao Guo et al. (2022). Explores multi-scale convolutional attention as an alternative lightweight backbone paradigm for dense visual tasks like semantic segmentation.
- Paper: VMamba: Visual State Space Model, Yue Liu et al. (2024). Investigates 2D selective state-space scanning to achieve global and local receptive fields with linear complexity beyond transformer attention mechanisms.
