Twins: Revisiting the Design of Spatial Attention in Vision Transformers

Xiangxiang ChuZhi TianYuqing WangBo ZhangHaibing RenXiaolin WeiHuaxia XiaChunhua Shen

article2021NeurIPS1,379 citations

Proposes two efficient vision transformer architectures that simplify spatial attention using standard matrix operations to deliver competitive performance across classification, detection, and segmentation tasks.

Listen

Vision transformers are emerging as powerful alternatives to traditional convolutional neural networks for computer vision, offering greater flexibility and natural multi-modal integration. However, their standard self-attention mechanisms suffer from quadratic computational complexity relative to image resolution. This makes applying transformers to high-resolution, dense visual tasks—such as object detection and semantic segmentation—computationally expensive. While existing approaches like shifted local windows mitigate this cost, they introduce deployment complexities due to memory-unfriendly cyclic shifts and uneven window partitions that hinder optimization on production runtimes.

The article demonstrates that simplified spatial attention mechanisms and proper positional encodings can outperform or match leading transformer models while improving efficiency and deployment ease. To achieve this, the article evaluates two novel vision transformer backbone architectures: Twins-PCPVT, which enhances pyramid transformers with dynamic conditional position encodings, and Twins-SVT, which introduces spatially separable self-attention by interleaving locally-grouped self-attention with global sub-sampled attention.

The evaluation was conducted across standard computer vision benchmarks using rigorous, controlled comparisons. Models were assessed on the ImageNet-1K dataset for image classification, the ADE20K dataset for semantic scene segmentation, and the COCO 2017 dataset for object detection and instance segmentation across multiple detector frameworks.

The findings establish that the proposed architectures achieve superior accuracy and efficiency compared to prior models. On ImageNet-1K classification, the small Twins-SVT variant achieves 81.7% top-1 accuracy, outperforming Swin Transformer while requiring approximately 35% fewer floating-point operations. On ADE20K semantic segmentation, Twins-PCPVT and Twins-SVT surpass earlier pyramid vision transformers and Swin baselines, reaching a state-of-the-art 50.2% mean intersection-over-union. For object detection and segmentation on COCO, both architectures consistently yield 1.5% to 6.7% improvements in average precision across single-scale and multi-scale training schedules. Furthermore, ablation experiments confirm that regular strided convolutions serve as the most effective sub-sampling mechanism for global attention.

These results demonstrate that complex window-shifting mechanisms are unnecessary to maintain wide receptive fields in vision transformers. Because the proposed spatially separable self-attention relies exclusively on standard matrix multiplications, it removes engineering bottlenecks associated with hardware-unfriendly operations. In production deployment tests, converting the architecture to an optimized inference framework yielded a 1.7-fold boost in processing throughput, directly lowering serving costs and latency.

For practical implementation, engineering and product teams should consider adopting spatially separable transformer backbones for high-resolution visual processing systems to capture both accuracy and hardware-efficiency gains. When deploying to edge or server environments, teams should prioritize models based on standard tensor operations to maximize runtime acceleration. Future development should explore automated stage-by-stage optimization of sub-window sizes and validate these backbones across additional domains such as video processing and 3D vision.

Confidence in these findings is high, as the empirical gains are demonstrated across multiple established benchmarks and detector frameworks under standardized training protocols. A minor limitation is that sub-window dimensions were set uniformly across model stages rather than tuned per resolution level, suggesting that custom-tuned configurations might unlock further performance gains.

arXiv: 2104.13840
Cover for Twins: Revisiting the Design of Spatial Attention in Vision Transformers

Abstract

Very recently, a variety of vision transformer architectures for dense prediction tasks have been proposed and they show that the design of spatial attention is critical to their success in these tasks. In this work, we revisit the design of the spatial attention and demonstrate that a carefully-devised yet simple spatial attention mechanism performs favourably against the state-of-the-art schemes. As a result, we propose two vision transformer architectures, namely, Twins-PCPVT and Twins-SVT. Our proposed architectures are highly-efficient and easy to implement, only involving matrix multiplications that are highly optimized in modern deep learning frameworks. More importantly, the proposed architectures achieve excellent performance on a wide range of visual tasks, including image level classification as well as dense detection and segmentation. The simplicity and strong performance suggest that our proposed architectures may serve as stronger backbones for many vision tasks. Our code is released at this https URL .

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Our Method: Twins
  • 3.1 Twins-PCPVT
  • 3.2 Twins-SVT
  • 4 Experiments
  • 4.1 Classification on ImageNet-1K
  • 4.2 Semantic Segmentation on ADE20K
  • 4.3 Object Detection and Segmentation on COCO
  • 4.4 Ablation Studies
  • 5 Conclusion
  • References
  • A Experiment
  • B Algorithm
  • C Architecture Setting

Knowls

  1. Knowl 1 — Spatially Separable Self-Attention Mechanism

    model/method

    Spatially Separable Self-Attention (SSSA) is an attention mechanism for vision transformers designed to capture both local and global spatial dependencies with linear computational scaling relative to image size, avoiding the quadratic complexity O(H2W2d)\mathcal{O}(H^2 W^2 d) of standard full self-attention.

    SSSA decomposes spatial modeling into two alternating operations:

    1. Locally-Grouped Self-Attention (LSA): A 2D feature map of spatial dimensions H×WH \times W is partitioned into m×nm \times n non-overlapping sub-windows, where each window has spatial size k1×k2k_1 \times k_2 with k1=H/mk_1 = H/m and k2=W/nk_2 = W/n. Self-attention is executed strictly within each local sub-window, capturing fine-grained short-range context with computational cost O(k1k2HWd)\mathcal{O}(k_1 k_2 H W d).

    2. Global Sub-Sampled Attention (GSA): To establish communication across disconnected sub-windows, a summary representative is extracted from each of the m×nm \times n sub-windows using a sub-sampling function (specifically, a regular strided 2D convolution). These m×nm \times n summary tokens serve as the key and value representations, while every original feature token across the entire feature map serves as a query. This enables global context modeling with computational cost O(mnHWd)=O(H2W2dk1k2)\mathcal{O}(m n H W d) = \mathcal{O}\left(\frac{H^2 W^2 d}{k_1 k_2}\right).

    SSSA alternates LSA and GSA sequentially, mirroring the separation of depthwise and pointwise operations in depthwise separable convolutions while relying exclusively on hardware-optimized matrix multiplications.

  2. Knowl 2 — Mathematical Formulation and Complexity Minimization of Spatially Separable Self-Attention

    equation

    Given an input feature tensor zl−1∈RH×W×C\mathbf{z}^{l-1} \in \mathbb{R}^{H \times W \times C} partitioned into m×nm \times n rectangular sub-windows {zijl−1}i=1,j=1m,n\{\mathbf{z}^{l-1}_{ij}\}_{i=1, j=1}^{m, n} of spatial size k1×k2k_1 \times k_2 (where k1=H/mk_1 = H/m and k2=W/nk_2 = W/n), a Spatially Separable Self-Attention (SSSA) block is formulated as:

    z^ijl=LSA(LayerNorm(zijl−1))+zijl−1,∀i∈{1,…,m},  j∈{1,…,n}\hat{\mathbf{z}}^l_{ij} = \text{LSA}(\text{LayerNorm}(\mathbf{z}^{l-1}_{ij})) + \mathbf{z}^{l-1}_{ij}, \quad \forall i \in \{1, \dots, m\}, \; j \in \{1, \dots, n\} zijl=FFN(LayerNorm(z^ijl))+z^ijl,∀i∈{1,…,m},  j∈{1,…,n}\mathbf{z}^l_{ij} = \text{FFN}(\text{LayerNorm}(\hat{\mathbf{z}}^l_{ij})) + \hat{\mathbf{z}}^l_{ij}, \quad \forall i \in \{1, \dots, m\}, \; j \in \{1, \dots, n\} z^l+1=GSA(LayerNorm(zl))+zl\hat{\mathbf{z}}^{l+1} = \text{GSA}(\text{LayerNorm}(\mathbf{z}^l)) + \mathbf{z}^l zl+1=FFN(LayerNorm(z^l+1))+z^l+1\mathbf{z}^{l+1} = \text{FFN}(\text{LayerNorm}(\hat{\mathbf{z}}^{l+1})) + \hat{\mathbf{z}}^{l+1}

    where LSA(⋅)\text{LSA}(\cdot) is the locally-grouped multi-head self-attention operating within each sub-window, GSA(⋅)\text{GSA}(\cdot) is the global sub-sampled multi-head self-attention interacting with representative keys generated from each sub-window z^ij∈Rk1×k2×C\hat{\mathbf{z}}_{ij} \in \mathbb{R}^{k_1 \times k_2 \times C}, LayerNorm(⋅)\text{LayerNorm}(\cdot) denotes layer normalization, and FFN(⋅)\text{FFN}(\cdot) denotes a multi-layer perceptron feed-forward network.

    For feature dimension dd, the combined computational complexity of the LSA and GSA pair is:

    O(H2W2dk1k2+k1k2HWd)\mathcal{O}\left(\frac{H^2 W^2 d}{k_1 k_2} + k_1 k_2 H W d\right)

    By the arithmetic-mean geometric-mean (AM-GM) inequality:

    H2W2dk1k2+k1k2HWd≥2HWdHW\frac{H^2 W^2 d}{k_1 k_2} + k_1 k_2 H W d \ge 2 H W d \sqrt{H W}

    The global computational complexity minimum is achieved when k1⋅k2=HWk_1 \cdot k_2 = \sqrt{H W}. For square sub-windows (k1=k2k_1 = k_2), the optimal window dimension is k1=k2=(HW)1/4k_1 = k_2 = (H W)^{1/4}. In practical network configurations, k1=k2=7k_1 = k_2 = 7 is set uniformly across early stages, while GSA summarization window sizes are controlled to 4, 2, and 1 in subsequent stages to maintain an adequate number of generated keys.

  3. Knowl 3 — Twins-PCPVT Architecture Design

    model/method

    Twins-PCPVT is a pyramid vision transformer backbone that combines the multi-stage global sub-sampled attention design of the Pyramid Vision Transformer (PVT) with Conditional Positional Encodings (CPE) generated by a Positional Encoding Generator (PEG).

    Key architectural components include:

    1. Positional Encoding Generator (PEG): Instead of fixed or learnable absolute positional encodings, a PEG is inserted immediately after the first transformer encoder block of each stage. The PEG consists of a 2D depthwise convolution (kernel size 3×33 \times 3) without batch normalization. Because the generated positional encodings are conditioned dynamically on local input features, the network handles variable input resolutions seamlessly and preserves 2D translation invariance.

    2. Classifier Head: For image classification, the class token (exttt[CLS] exttt{[CLS]}) is eliminated, and global average pooling (GAP) is applied over the spatial tokens at the output of the final stage prior to the linear classification layer.

    3. Multi-Scale Feature Pyramids: For dense downstream tasks (e.g., object detection and semantic segmentation), the architecture outputs feature representations at 4 different resolutions with downsampling strides of 4×,8×,16×,4\times, 8\times, 16\times, and 32×32\times relative to the input image.

  4. Knowl 4 — Architecture Configurations for Twins-PCPVT and Twins-SVT

    data/table

    Twins-PCPVT and Twins-SVT are instantiated in Small (S), Base (B), and Large (L) variants structured over 4 progressive downsampling stages with patch reduction factors PiP_i, channel dimensions CiC_i, attention reduction ratios RiR_i, number of attention heads NiN_i, and MLP expansion ratios EiE_i.

    Stage Output Size Twins-PCPVT-S Twins-PCPVT-B Twins-PCPVT-L
    Stage 1 H4×W4\frac{H}{4} \times \frac{W}{4} P1=4,C1=64P_1=4, C_1=64 P1=4,C1=64P_1=4, C_1=64 P1=4,C1=64P_1=4, C_1=64
    [R1=8,N1=1,E1=8]×3[R_1=8, N_1=1, E_1=8] \times 3 [R1=8,N1=1,E1=8]×3[R_1=8, N_1=1, E_1=8] \times 3 [R1=8,N1=1,E1=8]×3[R_1=8, N_1=1, E_1=8] \times 3
    Stage 2 H8×W8\frac{H}{8} \times \frac{W}{8} P2=2,C2=128P_2=2, C_2=128 P2=2,C2=128P_2=2, C_2=128 P2=2,C2=128P_2=2, C_2=128
    [R2=4,N2=2,E2=8]×3[R_2=4, N_2=2, E_2=8] \times 3 [R2=4,N2=2,E2=8]×3[R_2=4, N_2=2, E_2=8] \times 3 [R2=4,N2=2,E2=8]×8[R_2=4, N_2=2, E_2=8] \times 8
    Stage 3 H16×W16\frac{H}{16} \times \frac{W}{16} P3=2,C3=320P_3=2, C_3=320 P3=2,C3=320P_3=2, C_3=320 P3=2,C3=320P_3=2, C_3=320
    [R3=2,N3=5,E3=4]×6[R_3=2, N_3=5, E_3=4] \times 6 [R3=2,N3=5,E3=4]×18[R_3=2, N_3=5, E_3=4] \times 18 [R3=2,N3=5,E3=4]×27[R_3=2, N_3=5, E_3=4] \times 27
    Stage 4 H32×W32\frac{H}{32} \times \frac{W}{32} P4=2,C4=512P_4=2, C_4=512 P4=2,C4=512P_4=2, C_4=512 P4=2,C4=512P_4=2, C_4=512
    [R4=1,N4=8,E4=4]×3[R_4=1, N_4=8, E_4=4] \times 3 [R4=1,N4=8,E4=4]×3[R_4=1, N_4=8, E_4=4] \times 3 [R4=1,N4=8,E4=4]×3[R_4=1, N_4=8, E_4=4] \times 3
    Stage Output Size Twins-SVT-S Twins-SVT-B Twins-SVT-L
    Stage 1 H4×W4\frac{H}{4} \times \frac{W}{4} P1=4,C1=64P_1=4, C_1=64 P1=4,C1=96P_1=4, C_1=96 P1=4,C1=128P_1=4, C_1=128
    [LSA,GSA]×1[\text{LSA}, \text{GSA}] \times 1 [LSA,GSA]×1[\text{LSA}, \text{GSA}] \times 1 [LSA,GSA]×1[\text{LSA}, \text{GSA}] \times 1
    Stage 2 H8×W8\frac{H}{8} \times \frac{W}{8} P2=2,C2=128P_2=2, C_2=128 P2=2,C2=192P_2=2, C_2=192 P2=2,C2=256P_2=2, C_2=256
    [LSA,GSA]×1[\text{LSA}, \text{GSA}] \times 1 [LSA,GSA]×1[\text{LSA}, \text{GSA}] \times 1 [LSA,GSA]×1[\text{LSA}, \text{GSA}] \times 1
    Stage 3 H16×W16\frac{H}{16} \times \frac{W}{16} P3=2,C3=256P_3=2, C_3=256 P3=2,C3=384P_3=2, C_3=384 P3=2,C3=512P_3=2, C_3=512
    [LSA,GSA]×5[\text{LSA}, \text{GSA}] \times 5 [LSA,GSA]×9[\text{LSA}, \text{GSA}] \times 9 [LSA,GSA]×9[\text{LSA}, \text{GSA}] \times 9
    Stage 4 H32×W32\frac{H}{32} \times \frac{W}{32} P4=2,C4=512P_4=2, C_4=512 P4=2,C4=768P_4=2, C_4=768 P4=2,C4=1024P_4=2, C_4=1024
    [GSA]×4[\text{GSA}] \times 4 [GSA]×2[\text{GSA}] \times 2 [GSA]×2[\text{GSA}] \times 2
  5. Knowl 5 — PyTorch Implementation of Locally-Grouped Self-Attention

    algorithm

    The following PyTorch module implements Locally-Grouped Self-Attention (LSA). It reshapes a 2D spatial feature sequence into non-overlapping sub-windows of size k1×k2k_1 \times k_2, executes multi-head attention independently within each local window group, and projects the concatenated results back to the original token layout.

    import torch
    import torch.nn as nn
    
    class GroupAttention(nn.Module):
        def __init__(self, dim, num_heads=8, qkv_bias=False, qk_scale=None,
                     attn_drop=0., proj_drop=0., k1=7, k2=7):
            super(GroupAttention, self).__init__()
            self.dim = dim
            self.num_heads = num_heads
            head_dim = dim // num_heads
            self.scale = qk_scale or head_dim ** -0.5
            self.qkv = nn.Linear(dim, dim * 3, bias=qkv_bias)
            self.attn_drop = nn.Dropout(attn_drop)
            self.proj = nn.Linear(dim, dim)
            self.proj_drop = nn.Dropout(proj_drop)
            self.k1 = k1
            self.k2 = k2
    
        def forward(self, x, H, W):
            B, N, C = x.shape
            h_group, w_group = H // self.k1, W // self.k2
            total_groups = h_group * w_group
            x = x.reshape(B, h_group, self.k1, w_group, self.k2, C).transpose(2, 3)
            qkv = self.qkv(x).reshape(B, total_groups, -1, 3, self.num_heads, C // self.num_heads).permute(3, 0, 1, 4, 2, 5)
            q, k, v = qkv[0], qkv[1], qkv[2]
            attn = (q @ k.transpose(-2, -1)) * self.scale
            attn = attn.softmax(dim=-1)
            attn = self.attn_drop(attn)
            attn = (attn @ v).transpose(2, 3).reshape(B, h_group, w_group, self.k1, self.k2, C)
            x = attn.transpose(2, 3).reshape(B, N, C)
            x = self.proj(x)
            x = self.proj_drop(x)
            return x
    
  6. Knowl 6 — ImageNet-1K Classification Performance

    data/table

    Top-1 accuracy on ImageNet-1K classification evaluated at 224×224224 \times 224 resolution across ConvNets, standard vision transformers, and Twins architectures. Models are trained for 300 epochs with the AdamW optimizer (batch size 1024, initial learning rate 0.001 with cosine decay, linear warm-up for 5 epochs, stochastic depth rates 0.2, 0.3, and 0.5 for small, base, and large models). Throughput is measured on a single NVIDIA V100 GPU with batch size 192.

    Method Param (M) FLOPs (G) Throughput (img/s) Top-1 (%)
    RegNetY-4G 21 4.0 1157 80.0
    RegNetY-8G 39 8.0 592 81.7
    RegNetY-16G 84 16.0 335 82.9
    DeiT-Small/16 22.1 4.6 437 79.9
    CrossViT-S 26.7 5.6 - 81.0
    T2T-ViT-14 22 5.2 - 81.5
    TNT-S 23.8 5.2 - 81.3
    CoaT-Lite Small 20 4.0 - 81.9
    PVT-Small 24.5 3.8 820 79.8
    CPVT-Small-GAP 23 4.6 817 81.5
    Twins-PCPVT-S (ours) 24.1 3.8 815 81.2
    Swin-T 29 4.5 766 81.3
    Twins-SVT-S (ours) 24 2.9 1059 81.7
    PVT-Medium 44.2 6.7 526 81.2
    Twins-PCPVT-B (ours) 43.8 6.7 525 82.7
    Swin-S 50 8.7 444 83.0
    Twins-SVT-B (ours) 56 8.6 469 83.2
    DeiT-Base/16 86.6 17.6 292 81.8
    PVT-Large 61.4 9.8 367 81.7
    Twins-PCPVT-L (ours) 60.9 9.8 367 83.1
    Swin-B 88 15.4 275 83.3
    Twins-SVT-L (ours) 99.2 15.1 288 83.7

    Twins-PCPVT-S improves upon PVT-Small by +1.4% Top-1 with identical FLOPs. Twins-SVT-S achieves 81.7% Top-1, outperforming Swin-T (81.3%) while using 35% fewer FLOPs (2.9G vs. 4.5G) and delivering 38% higher throughput (1059 vs. 766 images/s).

  7. Knowl 7 — ADE20K Semantic Segmentation Performance

    data/table

    Semantic segmentation evaluations on the ADE20K validation set comparing backbones pretrained on ImageNet-1K under Semantic FPN (80k iterations) and UperNet (160k iterations, single-scale and multi-scale testing) frameworks at 512×512512 \times 512 resolution.

    Backbone Semantic FPN 80k UperNet 160k
    FLOPs (G) Param (M) mIoU (%) FLOPs (G) Param (M) mIoU / MS mIoU (%)
    ResNet50 45 28.5 36.7 - - -
    PVT-Small 40 28.2 39.8 - - -
    Twins-PCPVT-S 40 28.4 44.3 234 54.6 46.2 / 47.5
    Swin-T 46 31.9 41.5 237 59.9 44.5 / 45.8
    Twins-SVT-S 37 28.3 43.2 228 54.4 46.2 / 47.1
    ResNet101 66 47.5 38.8 258 86.0 - / 44.9
    PVT-Medium 55 48.0 41.6 - - -
    Twins-PCPVT-B 55 48.1 44.9 250 74.3 47.1 / 48.4
    Swin-S 70 53.2 45.2 261 81.3 47.6 / 49.5
    Twins-SVT-B 67 60.4 45.3 261 88.5 47.7 / 48.9
    PVT-Large 71 65.1 42.1 - - -
    Twins-PCPVT-L 71 65.3 46.4 269 91.5 48.6 / 49.8
    Swin-B 107 91.2 46.0 299 121.0 48.1 / 49.7
    Twins-SVT-L 102 103.7 46.7 297 133.0 48.8 / 50.2

    Under the Semantic FPN setting, Twins-PCPVT-S improves by +4.5% mIoU over PVT-Small, and Twins-SVT-S surpasses Swin-T by +1.7% mIoU. Under UperNet with multi-scale inference, Twins-SVT-L establishes 50.2% mIoU, outperforming Swin-B (49.7% mIoU).

  8. Knowl 8 — COCO Object Detection and Instance Segmentation Performance

    data/table

    Object detection (RetinaNet) and instance segmentation (Mask R-CNN) results on the COCO val2017 dataset evaluated at 800×600800 \times 600 resolution under 1×1\times (12 epochs) and 3×+MS3\times + \text{MS} (36 epochs with multi-scale training) schedules.

    Backbone Complexity RetinaNet 1×1\times Mask R-CNN 1×1\times
    FLOPs (G) Param (M) AP\text{AP} AP50\text{AP}_{50} AP75\text{AP}_{75} APb\text{AP}^b AP50b\text{AP}^b_{50} APm\text{AP}^m
    ResNet50 111 37.7 36.3 55.3 38.6 38.0 58.6 34.4
    PVT-Small 118 34.2 40.4 61.3 43.0 40.4 62.9 37.8
    Twins-PCPVT-S 118 34.4 43.0 64.1 46.0 42.9 65.8 40.0
    Swin-T 118 38.5 41.5 62.1 44.2 42.2 64.6 39.1
    Twins-SVT-S 104 34.3 43.0 64.2 46.3 43.4 66.0 40.3
    PVT-Medium 151 53.9 41.9 63.1 44.3 42.0 64.4 39.0
    Twins-PCPVT-B 151 54.1 44.3 65.6 47.3 44.6 66.7 40.9
    Swin-S 162 59.8 44.5 65.7 47.5 44.8 66.6 40.9
    Twins-SVT-B 163 67.0 45.3 66.7 48.1 45.2 67.6 41.5
    PVT-Large - 71.1 42.6 63.7 45.4 42.9 - 39.5
    Twins-PCPVT-L - 71.2 45.1 66.4 48.4 45.4 - 41.5
    Swin-B - 98.4 44.7 65.9 47.8 45.5 - 41.3
    Twins-SVT-L - 110.9 45.7 67.1 49.2 45.9 - 41.6

    Under RetinaNet 1×1\times, Twins-PCPVT-S outperforms PVT-Small by +2.6% AP\text{AP}, while Twins-SVT-S exceeds Swin-T by +1.5% AP\text{AP} with 12% fewer FLOPs (104G vs. 118G). Under Mask R-CNN 1×1\times, Twins-SVT-S achieves 43.4% box APb\text{AP}^b and 40.3% mask APm\text{AP}^m, surpassing Swin-T (42.2% APb\text{AP}^b, 39.1% APm\text{AP}^m).

  9. Knowl 9 — Ablation of Spatial Attention Block Arrangements and Sub-Sampling Functions

    empirical result

    Ablation experiments on ImageNet-1K classification evaluating combinations of Locally-Grouped Self-Attention (LSA, denoted LL) and Global Sub-Sampled Attention (GSA, denoted GG) blocks, as well as different sub-sampling functions for GSA in the small model variant:

    1. Arrangement of LL and GG blocks across stages:

      • Pure local attention (L,L,L)(L, L, L): 76.9% Top-1 (8.8M params, 2.2G FLOPs). The receptive field is restricted within local windows, leading to severe performance degradation.
      • Local-Local-Global (L,LLG,LLG,G)(L, LLG, LLG, G): 81.5% Top-1 (23.5M params, 2.8G FLOPs).
      • Interleaved Local-Global (L,LG,LG,G)(L, LG, LG, G): 81.7% Top-1 (24.1M params, 2.8G FLOPs). This configuration provides the best balance and forms the default design for Twins-SVT.
      • Global attention only in the final stage (L,L,L,G)(L, L, L, G): 80.5% Top-1 (22.2M params, 2.9G FLOPs).
      • Pure global attention (PVT-Small: G,G,G,GG, G, G, G): 79.8% Top-1 (24.5M params, 3.8G FLOPs).
    2. Sub-sampling function type for GSA:

      • Standard regular strided 2D convolution: 81.7% Top-1.
      • 2D depthwise separable strided convolution: 81.2% Top-1.
      • Average pooling: 81.2% Top-1.

    Regular strided 2D convolution is selected as the default sub-sampling operation in GSA.

  10. Knowl 10 — Invariance of Shifted Window Attention to Conditional Positional Encodings

    empirical result

    Replacing relative positional encodings in Swin Transformer with Conditional Positional Encodings (CPE generated by PEG) does not produce performance improvements across classification, detection, and segmentation:

    1. ImageNet-1K Classification: Swin-T with standard relative positional encoding achieves 81.3% Top-1 accuracy; replacing it with PEG yields 81.2% Top-1 accuracy (28M params, 4.4G FLOPs).

    2. COCO Object Detection with RetinaNet (1×1\times schedule): Standard Swin-T achieves 41.5% AP\text{AP} (62.1%  AP50,44.2%  AP7562.1\% \; \text{AP}_{50}, 44.2\% \; \text{AP}_{75}), whereas Swin-T + CPVT achieves 41.3% AP\text{AP} (62.4%  AP50,44.1%  AP7562.4\% \; \text{AP}_{50}, 44.1\% \; \text{AP}_{75}).

    3. COCO Instance Segmentation with Mask R-CNN (1×1\times schedule): Standard Swin-T achieves 42.2% AP\text{AP} (64.6%  AP50,46.2%  AP7564.6\% \; \text{AP}_{50}, 46.2\% \; \text{AP}_{75}), whereas Swin-T + CPVT achieves 42.0% AP\text{AP} (64.5%  AP50,45.9%  AP7564.5\% \; \text{AP}_{50}, 45.9\% \; \text{AP}_{75}).

    These empirical results indicate that CPE does not synergize with shifted window mechanisms, confirming that the performance gains in Twins-SVT stem from the spatially separable self-attention (SSSA) design rather than positional encodings alone.

Coverage note — All primary contributed models (Twins-PCPVT and Twins-SVT), attention mechanisms (SSSA, LSA, GSA), mathematical complexity analyses, PyTorch algorithm definitions, architecture settings, and empirical results across ImageNet-1K, ADE20K, and COCO are fully covered. No substantial contributed material was omitted.

References

  1. 1.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition at scale. In Proc. Int. Conf. Learn. Representations, 2021.
  2. 2.Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Hervé Jégou. Training data-efficient image transformers & distillation through attention. arXiv preprint arXiv:2012.12877, 2020.
  3. 3.Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. End-to-end object detection with transformers. In Proc. Eur. Conf. Comp. Vis., pages 213–229. Springer, 2020.
  4. 4.Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. arXiv preprint arXiv:2103.14030, 2021.
  5. 5.Liang-Chieh Chen, George Papandreou, Iasonas Kokkinos, Kevin Murphy, and Alan L Yuille. Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs. IEEE Trans. Pattern Anal. Mach. Intell., 40(4):834–848, 2017.
  6. 6.Chao Peng, Xiangyu Zhang, Gang Yu, Guiming Luo, and Jian Sun. Large kernel matters– improve semantic segmentation by global convolutional network. In Proc. IEEE Conf. Comp. Vis. Patt. Recogn., pages 4353–4361, 2017.
  7. 7.Wei Liu, Andrew Rabinovich, and Alexander C Berg. Parsenet: Looking wider to see better. arXiv preprint arXiv:1506.04579, 2015.
  8. 8.Wenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan, Kaitao Song, Ding Liang, Tong Lu, Ping Luo, and Ling Shao. Pyramid vision transformer: A versatile backbone for dense prediction without convolutions. arXiv preprint arXiv:2102.12122, 2021.
  9. 9.Xiangxiang Chu, Zhi Tian, Bo Zhang, Xinlong Wang, Xiaolin Wei, Huaxia Xia, and Chunhua Shen. Conditional positional encodings for vision transformers. arXiv preprint arXiv:2102.10882, 2021.
  10. 10.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proc. IEEE Conf. Comp. Vis. Patt. Recogn., pages 770–778, 2016.
  11. 11.Mingxing Tan and Quoc Le. Efficientnet: Rethinking model scaling for convolutional neural networks. In Proc. Int. Conf. Mach. Learn., pages 6105–6114. PMLR, 2019.
  12. 12.François Chollet. Xception: Deep learning with depthwise separable convolutions. In Proc. IEEE Conf. Comp. Vis. Patt. Recogn., pages 1251–1258, 2017.
  13. 13.Saining Xie, Ross Girshick, Piotr Dollár, Zhuowen Tu, and Kaiming He. Aggregated residual transformations for deep neural networks. In Proc. IEEE Conf. Comp. Vis. Patt. Recogn., pages 1492–1500, 2017.
  14. 14.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Proc. Advances in Neural Inf. Process. Syst., pages 6000–6010, 2017.
  15. 15.Kai Han, An Xiao, Enhua Wu, Jianyuan Guo, Chunjing Xu, and Yunhe Wang. Transformer in transformer. arXiv preprint arXiv:2103.00112, 2021.
  16. 16.Prajit Ramachandran, Niki Parmar, Ashish Vaswani, Irwan Bello, Anselm Levskaya, and Jon Shlens. Stand-alone self-attention in vision models. In Proc. Advances in Neural Inf. Process. Syst., volume 32, pages 68–80, 2019.
  17. 17.Weijian Xu, Yifan Xu, Tyler Chang, and Zhuowen Tu. Co-scale conv-attentional image transformers, 2021.
  18. 18.Yuqing Wang, Zhaoliang Xu, Xinlong Wang, Chunhua Shen, Baoshan Cheng, Hao Shen, and Huaxia Xia. End-to-end video instance segmentation with transformers. In Proc. IEEE Conf. Comp. Vis. Patt. Recogn., 2021.
  19. 19.Hanting Chen, Yunhe Wang, Tianyu Guo, Chang Xu, Yiping Deng, Zhenhua Liu, Siwei Ma, Chunjing Xu, Chao Xu, and Wen Gao. Pre-trained image processing transformer. arXiv preprint arXiv:2012.00364, 2020.
  20. 20.Fuzhi Yang, Huan Yang, Jianlong Fu, Hongtao Lu, and Baining Guo. Learning texture trans- former network for image super-resolution. In Proc. IEEE Conf. Comp. Vis. Patt. Recogn., pages 5791–5800, 2020.
  21. 21.Yanhong Zeng, Jianlong Fu, and Hongyang Chao. Learning joint spatial-temporal transforma- tions for video inpainting. In Proc. Eur. Conf. Comp. Vis., pages 528–543. Springer, 2020.
  22. 22.Zhigang Dai, Bolun Cai, Yugeng Lin, and Junying Chen. UP-DETR: Unsupervised pre-training for object detection with transformers. arXiv preprint arXiv:2011.09094, 2020.
  23. 23.Aravind Srinivas, Tsung-Yi Lin, Niki Parmar, Jonathon Shlens, Pieter Abbeel, and Ashish Vaswani. Bottleneck transformers for visual recognition. arXiv: Comp. Res. Repository, 2021.
  24. 24.Sixiao Zheng, Jiachen Lu, Hengshuang Zhao, Xiatian Zhu, Zekun Luo, Yabiao Wang, Yanwei Fu, Jianfeng Feng, Tao Xiang, Philip HS Torr, et al. Rethinking semantic segmentation from a sequence-to-sequence perspective with transformers. arXiv preprint arXiv:2012.15840, 2020.
  25. 25.Niki Parmar, Ashish Vaswani, Jakob Uszkoreit, Lukasz Kaiser, Noam Shazeer, Alexander Ku, and Dustin Tran. Image transformer. In Proc. Int. Conf. Mach. Learn., volume 80, pages 4055–4064, 2018.
  26. 26.Hugo Touvron, Matthieu Cord, Alexandre Sablayrolles, Gabriel Synnaeve, and Hervé Jégou. Going deeper with image transformers. arXiv preprint arXiv:2103.17239, 2021.
  27. 27.Li Yuan, Yunpeng Chen, Tao Wang, Weihao Yu, Yujun Shi, Francis Tay, Jiashi Feng, and Shuicheng Yan. Tokens-to-token ViT: Training vision transformers from scratch on imagenet. arXiv preprint arXiv:2101.11986, 2021.
  28. 28.Zihang Jiang, Qibin Hou, Li Yuan, Daquan Zhou, Xiaojie Jin, Anran Wang, and Jiashi Feng. Token labeling: Training an 85.4% top-1 accuracy vision transformer with 56M parameters on ImageNet. arXiv: Comp. Res. Repository, 2021.
  29. 29.Aravind Srinivas, Tsung-Yi Lin, Niki Parmar, Jonathon Shlens, Pieter Abbeel, and Ashish Vaswani. Bottleneck transformers for visual recognition. arXiv preprint arXiv:2101.11605, 2021.
  30. 30.Chun-Fu Chen, Quanfu Fan, and Rameswar Panda. CrossViT: Cross-attention multi-scale vision transformer for image classification. arXiv preprint arXiv:2103.14899, 2021.
  31. 31.Haiping Wu, Bin Xiao, Noel Codella, Mengchen Liu, Xiyang Dai, Lu Yuan, and Lei Zhang. CvT: Introducing convolutions to vision transformers. arXiv: Comp. Res. Repository, 2021.
  32. 32.Xizhou Zhu, Weijie Su, Lewei Lu, Bin Li, Xiaogang Wang, and Jifeng Dai. Deformable detr: Deformable transformers for end-to-end object detection. In Proc. Int. Conf. Learn. Representations, 2021.
  33. 33.Tsung-Yi Lin, Piotr Dollár, Ross Girshick, Kaiming He, Bharath Hariharan, and Serge Belongie. Feature pyramid networks for object detection. In Proc. IEEE Conf. Comp. Vis. Patt. Recogn., pages 2117–2125, 2017.
  34. 34.Yuhui Yuan, Lang Huang, Jianyuan Guo, Chao Zhang, Xilin Chen, and Jingdong Wang. OCNet: Object context network for scene parsing. arXiv: Comp. Res. Repository, 2021.
  35. 35.Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton. ImageNet classification with deep convolutional neural networks. In Proc. Advances in Neural Inf. Process. Syst., volume 25, pages 1097–1105, 2012.
  36. 36.Laurent Sifre and Stéphane Mallat. Rigid-motion scattering for texture classification. arXiv preprint arXiv:1403.1687, 2014.
  37. 37.Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. In Proc. Int. Conf. Learn. Representations, 2019.
  38. 38.Gao Huang, Yu Sun, Zhuang Liu, Daniel Sedra, and Kilian Weinberger. Deep networks with stochastic depth. In Proc. Eur. Conf. Comp. Vis., pages 646–661. Springer, 2016.
  39. 39.Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, et al. Imagenet large scale visual recognition challenge. International journal of computer vision, 115(3):211–252, 2015.
  40. 40.Ilija Radosavovic, Raj Prateek Kosaraju, Ross Girshick, Kaiming He, and Piotr Dollár. Design- ing network design spaces. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10428–10436, 2020.
  41. 41.Changlin Li, Tao Tang, Guangrun Wang, Jiefeng Peng, Bing Wang, Xiaodan Liang, and Xiaojun Chang. Bossnas: Exploring hybrid cnn-transformers with block-wisely self-supervised neural architecture search. arXiv preprint arXiv:2103.12424, 2021.
  42. 42.Bolei Zhou, Hang Zhao, Xavier Puig, Sanja Fidler, Adela Barriuso, and Antonio Torralba. Scene parsing through ADE20k dataset. In Proc. IEEE Conf. Comp. Vis. Patt. Recogn., pages 633–641, 2017.
  43. 43.Alexander Kirillov, Ross Girshick, Kaiming He, and Piotr Dollár. Panoptic feature pyramid networks. In Proc. IEEE Conf. Comp. Vis. Patt. Recogn., pages 6399–6408, 2019.
  44. 44.Tete Xiao, Yingcheng Liu, Bolei Zhou, Yuning Jiang, and Jian Sun. Unified perceptual parsing for scene understanding. In Proceedings of the European Conference on Computer Vision (ECCV), pages 418–434, 2018.
  45. 45.Sixiao Zheng, Jiachen Lu, Hengshuang Zhao, Xiatian Zhu, Zekun Luo, Yabiao Wang, Yanwei Fu, Jianfeng Feng, Tao Xiang, Philip H.S. Torr, and Li Zhang. Rethinking semantic segmentation from a sequence-to-sequence perspective with transformers. In CVPR, 2021.
  46. 46.Tsung-Yi Lin, Priya Goyal, Ross Girshick, Kaiming He, and Piotr Dollár. Focal loss for dense object detection. In Proc. IEEE Int. Conf. Comp. Vis., pages 2980–2988, 2017.
  47. 47.Kaiming He, Georgia Gkioxari, Piotr Dollár, and Ross Girshick. Mask R-CNN. In Proc. IEEE Int. Conf. Comp. Vis., pages 2961–2969, 2017.
  48. 48.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In Proc. Eur. Conf. Comp. Vis., pages 740–755, 2014.
  49. 49.Kai Chen, Jiaqi Wang, Jiangmiao Pang, Yuhang Cao, Yu Xiong, Xiaoxiao Li, Shuyang Sun, Wansen Feng, Ziwei Liu, Jiarui Xu, Zheng Zhang, Dazhi Cheng, Chenchen Zhu, Tianheng Cheng, Qijie Zhao, Buyu Li, Xin Lu, Rui Zhu, Yue Wu, Jifeng Dai, Jingdong Wang, Jianping Shi, Wanli Ouyang, Chen Change Loy, and Dahua Lin. MMDetection: Open MMLab detection toolbox and benchmark. arXiv preprint arXiv:1906.07155, 2019.

Citation

MLA
Chu, X., et al. “Twins: Revisiting the Design of Spatial Attention in Vision Transformers”. Advances in Neural Information Processing Systems, vol. 34, 2021, pp. 9355–66, https://proceedings.neurips.cc/paper_files/paper/2021/file/4e0928de075538c593fbdabb0c5ef2c3-Paper.pdf.
APA
Chu, X., Tian, Z., Wang, Y., Zhang, B., Ren, H., Wei, X., Xia, H., & Shen, C. (2021). Twins: Revisiting the Design of Spatial Attention in Vision Transformers. Advances in Neural Information Processing Systems, 34, 9355–9366. https://proceedings.neurips.cc/paper_files/paper/2021/file/4e0928de075538c593fbdabb0c5ef2c3-Paper.pdf
Chicago
Chu, X., Z. Tian, Y. Wang, et al. 2021. “Twins: Revisiting the Design of Spatial Attention in Vision Transformers”. Advances in Neural Information Processing Systems 34: 9355–66. https://proceedings.neurips.cc/paper_files/paper/2021/file/4e0928de075538c593fbdabb0c5ef2c3-Paper.pdf.
Harvard
Chu, X. et al. (2021) “Twins: Revisiting the Design of Spatial Attention in Vision Transformers”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 9355–9366. Available at: https://proceedings.neurips.cc/paper_files/paper/2021/file/4e0928de075538c593fbdabb0c5ef2c3-Paper.pdf.
Vancouver
1. Chu X, Tian Z, Wang Y, Zhang B, Ren H, Wei X, Xia H, Shen C (2021) Twins: Revisiting the Design of Spatial Attention in Vision Transformers. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 9355–9366

BibTeX

@inproceedings{chu2021twins,
  title = {Twins: Revisiting the Design of Spatial Attention in Vision Transformers},
  author = {Chu, Xiangxiang and Tian, Zhi and Wang, Yuqing and Zhang, Bo and Ren, Haibing and Wei, Xiaolin and Xia, Huaxia and Shen, Chunhua},
  year = {2021},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {34},
  pages = {9355-9366},
  url = {https://proceedings.neurips.cc/paper_files/paper/2021/file/4e0928de075538c593fbdabb0c5ef2c3-Paper.pdf}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors