CrossFormer: A Versatile Vision Transformer Hinging on Cross-scale Attention

Wenxiao WangLu YaoLong ChenBinbin LinDeng CaiXiaofei HeWei Liu

article2022ICLR390 citations

Proposes CrossFormer, a vision transformer that establishes cross-scale feature interactions using multi-scale patch embeddings and long-short distance attention to achieve superior performance across image classification, object detection, and segmentation benchmarks.

Listen

Modern computer vision systems rely heavily on attention mechanisms to capture relationships across visual scenes. However, existing vision architectures struggle to link visual features of differing scales—such as fine-grained details alongside broad background contexts—because they partition images into uniform, single-scale patches and compress representations to manage computational load. This limitation significantly hinders performance on complex, real-world visual perception tasks such as detecting objects and segmenting scenes.

The article develops and evaluates CrossFormer, a versatile vision architecture designed to establish cross-scale visual interactions while efficiently processing arbitrary input sizes. The design combines a cross-scale embedding layer that extracts multiple patch sizes concurrently, an alternating long- and short-distance attention mechanism that retains fine details without excessive computational expense, and a trainable dynamic position bias to accommodate variable input dimensions.

The authors conducted comprehensive empirical evaluations on standard computer vision benchmarks across four core tasks: image classification on ImageNet (1.28 million training images), object detection and instance segmentation on COCO 2017 (118,000 training images), and semantic segmentation on ADE20K (20,000 training images). CrossFormer variants ranging from tiny to large were compared against leading architectures under standardized training protocols.

The evaluation yielded several key findings:

  1. CrossFormer consistently outperformed competing architectures across all four evaluation benchmarks while maintaining comparable parameter counts and computational demands.
  2. The architecture demonstrated substantial gains in dense prediction tasks, achieving up to 1.7-point improvements in object detection precision on COCO and over 4-point gains in intersection-over-union for semantic segmentation on ADE20K compared to peer models.
  3. In image classification on ImageNet, CrossFormer variants achieved top-1 accuracies ranging from 81.5% to 84.0%, exceeding baseline models like PVT and Swin by at least 1.2% in small configurations.
  4. Ablation analyses confirmed that combining multi-scale patch extraction with long-short distance attention provided a 1.0% accuracy improvement over single-scale baselines, confirming the direct value of cross-scale interactions.

These findings demonstrate that explicitly enabling cross-scale interactions resolves a fundamental efficiency and accuracy trade-off in visual recognition models. For engineering and technology leaders, adopting architectures with dynamic position handling and cross-scale attention offers improved accuracy across multi-task visual pipelines—including dense pixel-level labeling—without requiring proportional increases in computational or memory budgets.

Organizations developing vision applications should evaluate CrossFormer backbones for downstream perception systems, particularly in dense detection and segmentation workloads where gains are most pronounced. System designers can select larger grouping parameters during deployment to reduce GPU memory footprint with negligible loss in detection accuracy.

The presented evaluations focus primarily on standard academic benchmark datasets and established model frameworks. Practical deployment in live production environments may require further testing to evaluate inference latency under specific edge-hardware constraints and performance on domain-specific data distributions.

Cover for CrossFormer: A Versatile Vision Transformer Hinging on Cross-scale Attention

Abstract

Transformers have made great progress in dealing with computer vision tasks. However, existing vision transformers do not yet possess the ability of building the interactions among features of different scales, which is perceptually important to visual inputs. The reasons are two-fold: (1) Input embeddings of each layer are equal-scale, so no cross-scale feature can be extracted; (2) to lower the computational cost, some vision transformers merge adjacent embeddings inside the self-attention module, thus sacrificing small-scale (fine-grained) features of the embeddings and also disabling the cross-scale interactions. To this end, we propose Cross-scale Embedding Layer (CEL) and Long Short Distance Attention (LSDA). On the one hand, CEL blends each embedding with multiple patches of different scales, providing the self-attention module itself with cross-scale features. On the other hand, LSDA splits the self-attention module into a short-distance one and a long-distance counterpart, which not only reduces the computational burden but also keeps both small-scale and large-scale features in the embeddings. Through the above two designs, we achieve cross-scale attention. Besides, we put forward a dynamic position bias for vision transformers to make the popular relative position bias apply to variable-sized images. Hinging on the cross-scale attention module, we construct a versatile vision architecture, dubbed CrossFormer, which accommodates variable-sized inputs. Extensive experiments show that CrossFormer outperforms the other vision transformers on image classification, object detection, instance segmentation, and semantic segmentation tasks. The code has been released: this https URL.

Table of Contents

  • 1 Introduction
  • 2 Background
  • 3 CrossFormer
  • 3.1 Cross-scale Embedding Layer (CEL)
  • 3.2 CrossFormer Block
  • 3.2.1 Long Short Distance Attention (LSDA)
  • 3.2.2 Dynamic Position Bias (DPB)
  • 3.3 Variants of CrossFormer
  • 4 Experiments
  • 4.1 Image Classification
  • 4.2 Object Detection and Instance Segmentation
  • 4.3 Semantic Segmentation
  • 4.4 Ablation Studies
  • 5 Conclusions
  • References
  • A CrossFormer
  • A.1 Pseudo code of LSDA
  • A.2 Dynamic Position Bias (DPB)
  • A.3 Variants of CrossFormer for Detection and Segmentation
  • B Experiments
  • B.1 Object Detection
  • B.2 Semantic Segmentation

Knowls

  1. Knowl 1 — Cross-Scale Embedding Layer

    model/method

    The Cross-scale Embedding Layer (CEL) constructs visual embeddings containing multi-scale representations by sampling an input feature map (or image) using multiple convolution kernels of different spatial dimensions simultaneously with a shared stride.

    In Stage 1 of the architecture, CEL applies four convolution kernels of sizes 4×44 \times 4, 8×88 \times 8, 16×1616 \times 16, and 32×3232 \times 32 with a constant stride of 4×44 \times 4. The four receptive fields share the same spatial center, and their projected outputs are concatenated along the channel dimension to form a single token embedding. In Stages 2, 3, and 4, CEL employs two convolution kernels of sizes 2×22 \times 2 and 4×44 \times 4 with a stride of 2×22 \times 2, reducing the spatial resolution by a factor of 4 while doubling channel capacity.

    To constrain the computational complexity of convolution, which scales with K2D2K^2 D^2 for kernel size KK and channel dimension DD, CEL allocates lower channel dimensions to larger kernels and higher channel dimensions to smaller kernels. For a target total embedding dimension DtD_t in Stage 1, the dimensions are allocated as:

    • Kernel 4×44 \times 4: Dt2\frac{D_t}{2} dimensions
    • Kernel 8×88 \times 8: Dt4\frac{D_t}{4} dimensions
    • Kernel 16×1616 \times 16: Dt8\frac{D_t}{8} dimensions
    • Kernel 32×3232 \times 32: Dt8\frac{D_t}{8} dimensions

    For example, when Dt=128D_t = 128, the kernel channel dimensions are 64, 32, 16, and 16, respectively.

  2. Knowl 2 — Long Short Distance Attention

    model/method

    Long Short Distance Attention (LSDA) factorizes self-attention over a 2D feature map of size S×SS \times S into two complementary attention modules: Short Distance Attention (SDA) and Long Distance Attention (LDA). SDA and LDA alternate across successive transformer blocks without merging adjacent embeddings.

    1. Short Distance Attention (SDA): The feature map is partitioned into non-overlapping local grids of size G×GG \times G adjacent embeddings. Self-attention is computed exclusively among tokens within the same local window, modeling short-range dependencies.

    2. Long Distance Attention (LDA): For an input feature map of size S×SS \times S, embeddings are sampled at a fixed grid interval II, grouping together embeddings that share the same sampling phase offset. The resulting group size is G×GG \times G, where G=S/IG = S / I. Self-attention is computed exclusively among tokens within each sampled grid, modeling global dependencies.

    For an S×SS \times S spatial grid containing N=S2N = S^2 tokens, standard multi-head self-attention requires O(S4)O(S^4) computation and memory. In contrast, LSDA operates on (S/G)2(S / G)^2 independent groups of G2G^2 tokens each, requiring:

    O((SG)2⋅(G2)2)=O(S2G2)\mathcal{O}\left(\left(\frac{S}{G}\right)^2 \cdot (G^2)^2\right) = \mathcal{O}(S^2 G^2)

    Because G≪SG \ll S in high-resolution stages, computational cost is substantially reduced while preserving fine-grained token representations.

  3. Knowl 3 — Dynamic Position Bias

    model/method

    Dynamic Position Bias (DPB) computes relative position representations dynamically via a multilayer perceptron (MLP), enabling relative position biases to generalize to variable image and group sizes.

    In self-attention with relative position bias, the attention map within a group of G2G^2 embeddings is computed as:

    Attention=Softmax(QKTd+B)V\text{Attention} = \text{Softmax}\left(\frac{QK^T}{\sqrt{d}} + B\right)V

    where Q,K,V∈RG2×DQ, K, V \in \mathbb{R}^{G^2 \times D}, d\sqrt{d} is the head dimension scaling factor, and B∈RG2×G2B \in \mathbb{R}^{G^2 \times G^2} is the position bias matrix. For two embeddings located at 2D grid coordinates (xi,yi)(x_i, y_i) and (xj,yj)(x_j, y_j) with coordinate offset (Δxij,Δyij)=(xi−xj,yi−yj)(\Delta x_{ij}, \Delta y_{ij}) = (x_i - x_j, y_i - y_j), DPB calculates the bias scalar entry Bi,jB_{i,j} as:

    Bi,j=DPB(Δxij,Δyij)B_{i,j} = \text{DPB}(\Delta x_{ij}, \Delta y_{ij})

    The DPB architecture consists of three fully-connected layers with Layer Normalization (LN) and ReLU activations:

    Linear(2,D/4)→LN→ReLU→Linear(D/4,D/4)→LN→ReLU→Linear(D/4,1)\text{Linear}(2, D/4) \to \text{LN} \to \text{ReLU} \to \text{Linear}(D/4, D/4) \to \text{LN} \to \text{ReLU} \to \text{Linear}(D/4, 1)

    where DD is the embedding dimension.

    For an embedding group of size G×GG \times G, the coordinate offsets satisfy 1−G≤Δxij,Δyij≤G−11 - G \le \Delta x_{ij}, \Delta y_{ij} \le G - 1. An intermediate lookup matrix B^∈R(2G−1)×(2G−1)\hat{B} \in \mathbb{R}^{(2G-1) \times (2G-1)} is precomputed in O(G2)\mathcal{O}(G^2) time as B^u,v=DPB(1−G+u,1−G+v)\hat{B}_{u, v} = \text{DPB}(1 - G + u, 1 - G + v) for 0≤u,v<2G−10 \le u, v < 2G - 1. The attention bias matrix BB is then populated by indexing Bi,j=B^Δxij+G−1,Δyij+G−1B_{i,j} = \hat{B}_{\Delta x_{ij} + G - 1, \Delta y_{ij} + G - 1}.

  4. Knowl 4 — Tensor Grouping and Reshaping Procedure for LSDA

    algorithm

    Long Short Distance Attention (LSDA) partitions an input tensor of shape (H,W,D)(H, W, D) into independent attention groups using tensor reshape and permutation operations without token merging.

    Input: Tensor xx of shape (H,W,D)(H, W, D), group size GG, attention type type∈{"SDA","LDA"}\text{type} \in \{\text{"SDA"}, \text{"LDA"}\}
    Output: Tensor xoutx_{\text{out}} of shape (H,W,D)(H, W, D)
    if type is "SDA" then
        x←reshape(x,(H//G,G,W//G,G,D))x \leftarrow \text{reshape}(x, (H // G, G, W // G, G, D))
        x←permute(x,(0,2,1,3,4))x \leftarrow \text{permute}(x, (0, 2, 1, 3, 4))
    elif type is "LDA" then
        x←reshape(x,(G,H//G,G,W//G,D))x \leftarrow \text{reshape}(x, (G, H // G, G, W // G, D))
        x←permute(x,(1,3,0,2,4))x \leftarrow \text{permute}(x, (1, 3, 0, 2, 4))
    x←reshape(x,(H⋅W//G2,G2,D))x \leftarrow \text{reshape}(x, (H \cdot W // G^2, G^2, D))
    x←Attention(x)x \leftarrow \text{Attention}(x)
    x←reshape(x,(H//G,W//G,G,G,D))x \leftarrow \text{reshape}(x, (H // G, W // G, G, G, D))
    if type is "SDA" then
        x←permute(x,(0,2,1,3,4))x \leftarrow \text{permute}(x, (0, 2, 1, 3, 4))
    elif type is "LDA" then
        x←permute(x,(2,0,3,1,4))x \leftarrow \text{permute}(x, (2, 0, 3, 1, 4))
    xout←reshape(x,(H,W,D))x_{\text{out}} \leftarrow \text{reshape}(x, (H, W, D))
    return xoutx_{\text{out}}
  5. Knowl 5 — CrossFormer Hierarchical Architecture

    model/method

    CrossFormer is organized into a four-stage hierarchical pyramid. Given an input image of size H0×W0×3H_0 \times W_0 \times 3, each stage i∈{1,2,3,4}i \in \{1, 2, 3, 4\} comprises a Cross-scale Embedding Layer (CEL) followed by nin_i CrossFormer blocks:

    • Stage 1: CEL (stride 4) yields feature maps of size H04×W04×D1\frac{H_0}{4} \times \frac{W_0}{4} \times D_1.
    • Stage 2: CEL (stride 2) yields feature maps of size H08×W08×D2\frac{H_0}{8} \times \frac{W_0}{8} \times D_2.
    • Stage 3: CEL (stride 2) yields feature maps of size H016×W016×D3\frac{H_0}{16} \times \frac{W_0}{16} \times D_3.
    • Stage 4: CEL (stride 2) yields feature maps of size H032×W032×D4\frac{H_0}{32} \times \frac{W_0}{32} \times D_4.

    Within each CrossFormer block ll, the computation uses Layer Normalization (LN), residual connections, Long Short Distance Attention (LSDA) with Dynamic Position Bias (DPB), and a Multilayer Perceptron (MLP):

    Xl′=Attention(LN(Xl−1))+Xl−1X'_l = \text{Attention}(\text{LN}(X_{l-1})) + X_{l-1}

    Xl=MLP(LN(Xl′))+Xl′X_l = \text{MLP}(\text{LN}(X'_l)) + X'_l

    where Attention\text{Attention} alternates between Short Distance Attention (SDA) and Long Distance Attention (LDA) in consecutive blocks.

  6. Knowl 6 — CrossFormer Architectural Variants and Layer Configurations

    data/table

    CrossFormer is configured into four standard model capacities: CrossFormer-Tiny (-T), Small (-S), Base (-B), and Large (-L). The table below lists layer configurations for an input resolution of 224×224224 \times 224, where DiD_i denotes embedding dimension, HiH_i denotes the number of self-attention heads, GiG_i is group size, and IiI_i is sampling interval in Stage ii.

    Stage CrossFormer-T CrossFormer-S CrossFormer-B CrossFormer-L
    Stage 1 D1=64,H1=2D_1=64, H_1=2 D1=96,H1=3D_1=96, H_1=3 D1=96,H1=3D_1=96, H_1=3 D1=128,H1=4D_1=128, H_1=4
    (56×5656 \times 56) G1=7,I1=8G_1=7, I_1=8 (×1\times 1) G1=7,I1=8G_1=7, I_1=8 (×2\times 2) G1=7,I1=8G_1=7, I_1=8 (×2\times 2) G1=7,I1=8G_1=7, I_1=8 (×2\times 2)
    Stage 2 D2=128,H2=4D_2=128, H_2=4 D2=192,H2=6D_2=192, H_2=6 D2=192,H2=6D_2=192, H_2=6 D2=256,H2=8D_2=256, H_2=8
    (28×2828 \times 28) G2=7,I2=4G_2=7, I_2=4 (×1\times 1) G2=7,I2=4G_2=7, I_2=4 (×2\times 2) G2=7,I2=4G_2=7, I_2=4 (×2\times 2) G2=7,I2=4G_2=7, I_2=4 (×2\times 2)
    Stage 3 D3=256,H3=8D_3=256, H_3=8 D3=384,H3=12D_3=384, H_3=12 D3=384,H3=12D_3=384, H_3=12 D3=512,H3=16D_3=512, H_3=16
    (14×1414 \times 14) G3=7,I3=2G_3=7, I_3=2 (×8\times 8) G3=7,I3=2G_3=7, I_3=2 (×6\times 6) G3=7,I3=2G_3=7, I_3=2 (×18\times 18) G3=7,I3=2G_3=7, I_3=2 (×18\times 18)
    Stage 4 D4=512,H4=16D_4=512, H_4=16 D4=768,H4=24D_4=768, H_4=24 D4=768,H4=24D_4=768, H_4=24 D4=1024,H4=32D_4=1024, H_4=32
    (7×77 \times 7) G4=7,I4=1G_4=7, I_4=1 (×6\times 6) G4=7,I4=1G_4=7, I_4=1 (×2\times 2) G4=7,I4=1G_4=7, I_4=1 (×2\times 2) G4=7,I4=1G_4=7, I_4=1 (×2\times 2)

    For downstream tasks with large images, adapted variants (denoted with ‡\ddagger) configure G1=G2=14G_1 = G_2 = 14, I1=16I_1 = 16, and I2=8I_2 = 8 for the first two stages while maintaining identical weight dimensions.

  7. Knowl 7 — ImageNet-1K Image Classification Performance

    data/table

    CrossFormer variants were evaluated on ImageNet-1K (1.28M training images, 50K validation images) with 224×224224 \times 224 input resolution trained for 300 epochs using AdamW.

    Architecture Parameters FLOPs Top-1 Accuracy (%)
    PVT-S 24.5M 3.8G 79.8
    RegionViT-T 13.8M 2.4G 80.4
    Twins-SVT-S 24.0M 2.8G 81.3
    CrossFormer-T 27.8M 2.9G 81.5
    DeiT-S 22.1M 4.6G 79.8
    Swin-T 29.0M 4.5G 81.3
    ViL-S 24.6M 4.9G 81.8
    RegionViT-S 30.6M 5.3G 82.5
    CrossFormer-S 30.7M 4.9G 82.5
    PVT-L 61.4M 9.8G 81.7
    Swin-S 50.0M 8.7G 83.0
    Twins-SVT-B 56.0M 8.3G 83.1
    NesT-S 38.0M 10.4G 83.3
    CrossFormer-B 52.0M 9.2G 83.4
    DeiT-B 86.0M 17.5G 81.8
    RegionViT-B 72.0M 13.3G 83.3
    Twins-SVT-L 99.2M 14.8G 83.3
    Swin-B 88.0M 15.4G 83.3
    NesT-B 68.0M 17.9G 83.8
    CrossFormer-L 92.0M 16.1G 84.0
  8. Knowl 8 — COCO Object Detection and Instance Segmentation Performance

    data/table

    CrossFormer backbones were evaluated on COCO 2017 validation set using RetinaNet for object detection (1×1\times schedule, 12 epochs) and Mask R-CNN for instance segmentation (1×1\times and 3×3\times multi-scale schedules).

    Detector Backbone #Params FLOPs APb\text{AP}^b AP50b\text{AP}^b_{50} AP75b\text{AP}^b_{75} APm\text{AP}^m AP50m\text{AP}^m_{50} AP75m\text{AP}^m_{75}
    RetinaNet (1×1\times) Swin-T 38.5M 245.0G 41.5 62.1 44.2 – – –
    PVT-M 53.9M – 41.9 63.1 44.3 – – –
    TransCNN-B 36.5M – 43.4 64.2 46.5 – – –
    CrossFormer-S 40.8M 282.0G 44.4 65.8 47.4 – – –
    CrossFormer-S‡^{\ddagger} 40.8M 272.1G 44.2 65.7 47.2 – – –
    Swin-B 98.4M 477.0G 44.7 65.9 49.2 – – –
    Twins-SVT-L 110.9M 455.0G 44.8 66.1 48.1 – – –
    CrossFormer-B 62.1M 389.0G 46.2 67.8 49.5 – – –
    CrossFormer-B‡^{\ddagger} 62.1M 379.1G 46.1 67.7 49.0 – – –
    Mask R-CNN (1×1\times) Swin-T 47.8M 264.0G 42.2 64.6 46.2 39.1 61.6 42.0
    Twins-PCPVT-S 44.3M 245.0G 42.9 65.8 47.1 40.0 62.7 42.9
    CrossFormer-S 50.2M 301.0G 45.4 68.0 49.7 41.4 64.8 44.6
    CrossFormer-S‡^{\ddagger} 50.2M 291.1G 45.0 67.9 49.1 41.2 64.6 44.3
    Swin-S 69.1M 354.0G 44.8 66.6 48.9 40.9 63.4 44.2
    Twins-SVT-L 119.7M 474.0G 45.2 67.5 49.4 41.2 64.5 44.5
    CrossFormer-B 71.5M 407.9G 47.2 69.9 51.8 42.7 66.6 46.2
    CrossFormer-B‡^{\ddagger} 71.5M 398.1G 47.1 69.9 52.0 42.7 66.5 46.1
    Mask R-CNN (3×3\times) Swin-T 47.8M 264.0G 46.0 68.2 50.2 41.6 65.1 44.8
    Shuffle-T 48.0M 268.0G 46.8 68.9 51.5 42.3 66.0 45.6
    CrossFormer-S‡^{\ddagger} 50.2M 291.1G 48.7 70.7 53.7 43.9 67.9 47.3
    Swin-S 69.1M 354.0G 48.5 70.2 53.5 43.3 67.3 46.6
    Shuffle-S 69.0M 359.0G 48.4 70.1 53.5 43.3 67.3 46.7
    CrossFormer-B‡^{\ddagger} 71.5M 398.1G 49.8 71.6 54.9 44.5 68.8 47.9

    CrossFormer variants with ‡\ddagger configure (G1=14,I1=16,G2=14,I2=8)(G_1=14, I_1=16, G_2=14, I_2=8) in the first two stages to reduce memory usage during high-resolution dense training.

  9. Knowl 9 — ADE20K Semantic Segmentation Performance

    data/table

    CrossFormer backbones were evaluated on the ADE20K semantic segmentation validation set (150 semantic categories) using Semantic FPN (80K iterations) and UPerNet (160K iterations).

    Semantic FPN (80K iters) UPerNet (160K iters)
    Backbone #Params FLOPs mIoU Backbone #Params FLOPs mIoU MS mIoU
    PVT-M 48.0M 219.0G 41.6 Swin-T 60.0M 945.0G 44.5 45.8
    Twins-SVT-B 60.4M 261.0G 45.0 Shuffle-T 60.0M 949.0G 46.6 47.6
    Swin-S 53.2M 274.0G 45.2 CrossFormer-S 62.3M 979.5G 47.6 48.4
    CrossFormer-S 34.3M 220.7G 46.0 CrossFormer-S‡^{\ddagger} 62.3M 968.5G 47.4 48.2
    CrossFormer-S‡^{\ddagger} 34.3M 209.8G 46.4
    PVT-L 65.1M 283.0G 42.1 Swin-S 81.0M 1038.0G 47.6 49.5
    CAT-B 55.0M 276.0G 43.6 Shuffle-S 81.0M 1044.0G 48.4 49.6
    CrossFormer-B 55.6M 331.0G 47.7 CrossFormer-B 83.6M 1089.7G 49.7 50.6
    CrossFormer-B‡^{\ddagger} 55.6M 320.1G 48.0 CrossFormer-B‡^{\ddagger} 83.6M 1078.8G 49.2 50.1
    Twins-SVT-L 103.7M 397.0G 45.8 Swin-B 121.0M 1088.0G 48.1 49.7
    CrossFormer-L 95.4M 497.0G 48.7 Shuffle-B 121.0M 1096.0G 49.0 –
    CrossFormer-L‡^{\ddagger} 95.4M 482.7G 49.1 CrossFormer-L 125.5M 1257.8G 50.4 51.4
    CrossFormer-L‡^{\ddagger} 125.5M 1243.5G 50.5 51.4

    MS mIoU indicates multi-scale testing evaluation.

  10. Knowl 10 — Ablation Analysis of CrossFormer Architectural Components

    empirical result

    Ablation experiments conducted on the CrossFormer-S baseline (82.5% ImageNet Top-1 accuracy) evaluate individual architectural designs:

    1. Cross-scale Embedding Layer vs Single-Scale Embeddings:

      • Single-scale 4×44 \times 4 kernel across all stages achieves 81.5% accuracy.
      • Single-scale 8×88 \times 8 kernel in Stage 1 achieves 81.9% accuracy (+0.4% from larger receptive field).
      • Multi-scale kernels in Stage 1 (4×4,8×8,16×16,32×324 \times 4, 8 \times 8, 16 \times 16, 32 \times 32) and later stages (2×2,4×42 \times 2, 4 \times 4) achieve 82.5% accuracy (+1.0% total gain over single-scale).
      • Varying kernel configurations across stages consistently yields between 82.3% and 82.5% accuracy.
    2. Long Short Distance Attention vs Alternative Self-Attentions:

      • PVT-like spatial reduction attention + CEL: 81.3% accuracy.
      • Swin-like local window attention + CEL: 81.9% accuracy.
      • LSDA without CEL: 81.5% accuracy.
      • LSDA + CEL: 82.5% accuracy (an improvement of ≥0.6%\ge 0.6\% over substitute attention mechanisms).
    3. Dynamic Position Bias vs Other Position Representations:

      • Absolute Position Embedding (APE): 82.1% accuracy, 686 imgs/sec throughput.
      • Relative Position Bias (RPB): 82.5% accuracy, 684 imgs/sec throughput (restricted to fixed input sizes).
      • Dynamic Position Bias (DPB): 82.5% accuracy, 672 imgs/sec throughput (accommodates arbitrary input sizes).
      • DPB with residual connections: 82.4% accuracy, 672 imgs/sec throughput.

Coverage note — None was omitted; all contributed mechanisms (CEL, LSDA, DPB), architectural specifications, algorithm implementations, and empirical benchmark results on ImageNet, COCO, ADE20K, and ablations are included.

References

  1. 1.Lei Jimmy Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton. Layer normalization. CoRR, abs/1607.06450, 2016.
  2. 2.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. Language models are few-shot learners. In Neural Information Processing Systems, NeurIPS, 2020.
  3. 3.Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. End-to-end object detection with transformers. In Computer Vision - ECCV 2020 - 16th European Conference, Glasgow, UK, August 23-28, 2020, Proceedings, Part I, volume 12346, pp. 213–229, 2020.
  4. 4.Chun-Fu Chen, Quanfu Fan, and Rameswar Panda. Crossvit: Cross-attention multi-scale vision transformer for image classification. CoRR, abs/2103.14899, 2021a.
  5. 5.Chun-Fu Chen, Rameswar Panda, and Quanfu Fan. Regionvit: Regional-to-local attention for vision transformers. CoRR, abs/2106.02689, 2021b.
  6. 6.Kai Chen, Jiaqi Wang, Jiangmiao Pang, Yuhang Cao, Yu Xiong, Xiaoxiao Li, Shuyang Sun, Wansen Feng, Ziwei Liu, Jiarui Xu, Zheng Zhang, Dazhi Cheng, Chenchen Zhu, Tianheng Cheng, Qijie Zhao, Buyu Li, Xin Lu, Rui Zhu, Yue Wu, Jifeng Dai, Jingdong Wang, Jianping Shi, Wanli Ouyang, Chen Change Loy, and Dahua Lin. MMDetection: Open mmlab detection toolbox and benchmark. arXiv preprint arXiv:1906.07155, 2019.
  7. 7.Xiangxiang Chu, Zhi Tian, Yuqing Wang, Bo Zhang, Haibing Ren, Xiaolin Wei, Huaxia Xia, and Chunhua Shen. Twins: Revisiting spatial attention design in vision transformers. CoRR, abs/2104.13840, 2021.
  8. 8.MMSegmentation Contributors. MMSegmentation: Openmmlab semantic segmentation toolbox and benchmark, 2020.
  9. 9.Ekin Dogus Cubuk, Barret Zoph, Jon Shlens, and Quoc Le. Randaugment: Practical automated data augmentation with a reduced search space. In Neural Information Processing Systems, NeurIPS, 2020.
  10. 10.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT: pre-training of deep bidirectional transformers for language understanding. In Conference of the North American Chapter of the Association for Computational Linguistics, NAACL, pp. 4171–4186, 2019.
  11. 11.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition at scale. In International Conference on Learning Representations, ICLR, 2021.
  12. 12.Kai Han, An Xiao, Enhua Wu, Jianyuan Guo, Chunjing Xu, and Yunhe Wang. Transformer in transformer. CoRR, abs/2103.00112, 2021.
  13. 13.Kaiming He, Georgia Gkioxari, Piotr Doll'ar, and Ross B. Girshick. Mask R-CNN. In ' International Conference on Computer Vision, ICCV, pp. 2980–2988, 2017.
  14. 14.Gao Huang, Yu Sun, Zhuang Liu, Daniel Sedra, and Kilian Q. Weinberger. Deep networks with stochastic depth. In Bastian Leibe, Jiri Matas, Nicu Sebe, and Max Welling (eds.), European Conference on Computer Vision, ECCV, volume 9908, pp. 646–661, 2016.
  15. 15.Zilong Huang, Youcheng Ben, Guozhong Luo, Pei Cheng, Gang Yu, and Bin Fu. Shuffle transformer: Rethinking spatial shuffle for vision transformer. CoRR, abs/2106.03650, 2021.
  16. 16.Diederik P. Kingma and Jimmy Ba. Adam: A method for stochastic optimization. In International Conference on Learning Representations, ICLR, 2015.
  17. 17.Alexander Kirillov, Ross B. Girshick, Kaiming He, and Piotr Doll'ar. Panoptic feature pyramid ' networks. In Conference on Computer Vision and Pattern Recognition, CVPR, pp. 6399–6408, 2019.
  18. 18.Hezheng Lin, Xing Cheng, Xiangyu Wu, Fan Yang, Dong Shen, Zhongyuan Wang, Qing Song, and Wei Yuan. CAT: cross attention in vision transformer. CoRR, abs/2106.05786, 2021.
  19. 19.Tsung-Yi Lin, Michael Maire, Serge J. Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll'ar, and C. Lawrence Zitnick. Microsoft COCO: common objects in context. In ' European Conference on Computer Vision, ECCV, volume 8693, pp. 740–755, 2014.
  20. 20.Tsung-Yi Lin, Priya Goyal, Ross B. Girshick, Kaiming He, and Piotr Doll'ar. Focal loss for dense ' object detection. Transactions on Pattern Analysis and Machine Intelligence, PAMI, 42(2):318– 327, 2020.
  21. 21.Yun Liu, Guolei Sun, Yu Qiu, Le Zhang, Ajad Chhatkuli, and Luc Van Gool. Transformer in convolutional neural networks. CoRR, abs/2106.03180, 2021a.
  22. 22.Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. CoRR, abs/2103.14030, 2021b.
  23. 23.Vinod Nair and Geoffrey E. Hinton. Rectified linear units improve restricted boltzmann machines. In Johannes Fürnkranz and Thorsten Joachims (eds.), ¨ International Conference on Machine Learning, ICML, pp. 807–814, 2010.
  24. 24.Zizheng Pan, Bohan Zhuang, Jing Liu, Haoyu He, and Jianfei Cai. Scalable visual transformers with hierarchical pooling. CoRR, abs/2103.10619, 2021.
  25. 25.Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael S. Bernstein, Alexander C. Berg, and Fei-Fei Li. Imagenet large scale visual recognition challenge. International Journal of Computer Vision, IJCV, 2015.
  26. 26.Peter Shaw, Jakob Uszkoreit, and Ashish Vaswani. Self-attention with relative position representations. In Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL, pp. 464–468, 2018.
  27. 27.Aravind Srinivas, Tsung-Yi Lin, Niki Parmar, Jonathon Shlens, Pieter Abbeel, and Ashish Vaswani. Bottleneck transformers for visual recognition. In Conference on Computer Vision and Pattern Recognition, CVPR, 2021.
  28. 28.Peize Sun, Rufeng Zhang, Yi Jiang, Tao Kong, Chenfeng Xu, Wei Zhan, Masayoshi Tomizuka, Lei Li, Zehuan Yuan, Changhu Wang, and Ping Luo. Sparse R-CNN: end-to-end object detection with learnable proposals. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2021, virtual, June 19-25, 2021, pp. 14454–14463, 2021.
  29. 29.Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Herve J ' egou. Training data-efficient image transformers & distillation through attention. In ' International Conference on Machine Learning, ICML, volume 139, pp. 10347–10357, 2021.
  30. 30.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Neural Information Processing Systems, NeurIPS, pp. 5998–6008, 2017.
  31. 31.Wenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan, Kaitao Song, Ding Liang, Tong Lu, Ping Luo, and Ling Shao. Pyramid vision transformer: A versatile backbone for dense prediction without convolutions. CoRR, abs/2102.12122, 2021.
  32. 32.Haiping Wu, Bin Xiao, Noel Codella, Mengchen Liu, Xiyang Dai, Lu Yuan, and Lei Zhang. Cvt: Introducing convolutions to vision transformers. CoRR, abs/2103.15808, 2021.
  33. 33.Tete Xiao, Yingcheng Liu, Bolei Zhou, Yuning Jiang, and Jian Sun. Unified perceptual parsing for scene understanding. In European Conference on Computer Vision, ECCV, volume 11209, pp. 432–448, 2018.
  34. 34.Yufei Xu, Qiming Zhang, Jing Zhang, and Dacheng Tao. Vitae: Vision transformer advanced by exploring intrinsic inductive bias. CoRR, abs/2106.03348, 2021.
  35. 35.Li Yuan, Yunpeng Chen, Tao Wang, Weihao Yu, Yujun Shi, Francis E. H. Tay, Jiashi Feng, and Shuicheng Yan. Tokens-to-token vit: Training vision transformers from scratch on imagenet. CoRR, abs/2101.11986, 2021.
  36. 36.Sangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh, Youngjoon Yoo, and Junsuk Choe. Cutmix: Regularization strategy to train strong classifiers with localizable features. In International Conference on Computer Vision, ICCV, pp. 6022–6031, 2019.
  37. 37.Hongyi Zhang, Moustapha Cissé, Yann N. Dauphin, and David Lopez-Paz. mixup: Beyond empiri- ' cal risk minimization. In International Conference on Learning Representations, ICLR, 2018.
  38. 38.Pengchuan Zhang, Xiyang Dai, Jianwei Yang, Bin Xiao, Lu Yuan, Lei Zhang, and Jianfeng Gao. Multi-scale vision longformer: A new vision transformer for high-resolution image encoding. CoRR, abs/2103.15358, 2021a.
  39. 39.Qing-Long Zhang and Yubin Yang. Rest: An efficient transformer for visual recognition. CoRR, abs/2105.13677, 2021.
  40. 40.Zizhao Zhang, Han Zhang, Long Zhao, Ting Chen, and Tomas Pfister. Aggregating nested transformers. CoRR, abs/2105.12723, 2021b.
  41. 41.Zhun Zhong, Liang Zheng, Guoliang Kang, Shaozi Li, and Yi Yang. Random erasing data augmentation. In Association for the Advancement of Artificial Intelligence, AAAI, pp. 13001–13008, 2020.
  42. 42.Bolei Zhou, Hang Zhao, Xavier Puig, Sanja Fidler, Adela Barriuso, and Antonio Torralba. Scene parsing through ADE20K dataset. In Conference on Computer Vision and Pattern Recognition, CVPR, pp. 5122–5130, 2017.

Citation

MLA
Wang, W., et al. “CrossFormer: A Versatile Vision Transformer Hinging on Cross-scale Attention”. arXiv, 2021, http://arxiv.org/abs/2108.00154v2.
APA
Wang, W., Yao, L., Chen, L., Lin, B., Cai, D., He, X., & Liu, W. (2021). CrossFormer: A Versatile Vision Transformer Hinging on Cross-scale Attention. arXiv. http://arxiv.org/abs/2108.00154v2
Chicago
Wang, W., L. Yao, L. Chen, et al. 2021. “CrossFormer: A Versatile Vision Transformer Hinging on Cross-scale Attention”. arXiv. http://arxiv.org/abs/2108.00154v2.
Harvard
Wang, W. et al. (2021) “CrossFormer: A Versatile Vision Transformer Hinging on Cross-scale Attention”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2108.00154v2.
Vancouver
1. Wang W, Yao L, Chen L, Lin B, Cai D, He X, Liu W (2021) CrossFormer: A Versatile Vision Transformer Hinging on Cross-scale Attention. arXiv

BibTeX

@article{wang2021crossformer,
  title = {CrossFormer: A Versatile Vision Transformer Hinging on Cross-scale Attention},
  author = {Wang, Wenxiao and Yao, Lu and Chen, Long and Lin, Binbin and Cai, Deng and He, Xiaofei and Liu, Wei},
  year = {2021},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2108.00154v2},
  eprint = {2108.00154}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors