Castling-ViT: Compressing Self-Attention via Switching Towards Linear-Angular Attention at Vision Transformer Inference

Haoran YouYunyang XiongXiaoliang DaiBichen WuPeizhao ZhangHaoqi FanPeter VajdaYingyan Celine Lin

article2023CVPR54 citations

Proposes a training framework that pairs linear-angular attention with a decaying auxiliary softmax branch, enabling vision transformers to switch to efficient linear-complexity inference without sacrificing accuracy across classification and detection tasks.

Listen

Vision transformers deliver leading accuracy across visual recognition tasks, but their core self-attention mechanisms suffer from computational complexity that grows quadratically with image resolution. While previous efficient designs adopted local windows or linear approximations to reduce computation, they often sacrificed the model's ability to capture global context or fine local details. This runtime inefficiency creates a major deployment bottleneck, particularly for high-resolution vision applications on resource-constrained platforms.

The article evaluates Castling-ViT, a novel framework designed to close the accuracy gap between efficient linear attention and standard quadratic self-attention without introducing additional inference overhead. Castling-ViT trains vision transformers using both an efficient linear-angular attention mechanism and an auxiliary masked quadratic attention branch, then eliminates the auxiliary branch entirely during inference—a switch analogous to the castling move in chess.

To construct this approach, the authors decomposed spectral angular similarity kernels into linear terms and higher-order residual terms. The linear terms are retained for lightweight computation, while the non-linear residuals are approximated using a depthwise convolution alongside an auxiliary masked softmax attention module. During training, a thresholding regularization causes the auxiliary attention masks to naturally decay to zero. The evaluation assessed the framework across standard benchmarks: ImageNet for image classification, COCO for object detection, and ADE20K for semantic segmentation, integrating the mechanism into popular baseline architectures.

The experimental findings show substantial improvements in efficiency and accuracy. In image classification benchmarks, Castling-ViT delivered up to a 1.8% top-1 accuracy improvement under comparable computational budgets or up to a 40% reduction in multiply-accumulate operations while maintaining baseline accuracy. In object detection tasks on the COCO dataset, Castling-ViT improved average precision by up to 1.2 to 6.0 points compared to alternative convolutional and transformer baselines under similar operation counts. In semantic segmentation on ADE20K, integrating Castling-ViT into standard segmentation backbones achieved 15% to 19% reductions in computational operations while matching or exceeding baseline segmentation quality. Furthermore, ablation experiments confirmed that the proposed linear-angular kernel outperformed five other standard linear attention kernels by up to 4.6% in detection accuracy.

These results demonstrate that vision models do not need to choose between global context modeling and operational efficiency. Deploying linear-angular attention enables lower runtime latency, reduced hardware resource consumption, and lower operational costs for high-resolution visual processing. The findings challenge the conventional assumption that linear attention must underperform traditional quadratic attention, proving that auxiliary training mechanisms can effectively transfer complex feature learning into streamlined runtime architectures.

Organizations seeking to deploy vision transformers in resource-limited or low-latency environments should consider adopting linear-angular attention modules as drop-in replacements for standard self-attention. For development pipelines, engineering teams should evaluate the castling training recipe when training or fine-tuning transformer backbones for detection and segmentation tasks.

Confidence in these findings is supported by consistent empirical improvements across multiple architectures and standard vision benchmarks. However, a key limitation highlighted by the authors is that directly swapping transformer blocks into lightweight convolutional backbones can still introduce overhead compared to purely convolutional designs, as overall network structures may not yet be optimal. Further architectural exploration and on-device hardware profiling are recommended before deploying these models in strictly resource-bounded production systems.

arXiv: 2211.10526
Cover for Castling-ViT: Compressing Self-Attention via Switching Towards Linear-Angular Attention at Vision Transformer Inference

Abstract

Vision Transformers (ViTs) have shown impressive performance but still require a high computation cost as compared to convolutional neural networks (CNNs), one reason is that ViTs’ attention measures global similarities and thus has a quadratic complexity with the number of input tokens. Existing efficient ViTs adopt local attention or linear attention, which sacrifice ViTs’ capabilities of capturing either global or local context. In this work, we ask an important research question: Can ViTs learn both global and local context while being more efficient during inference? To this end, we propose a framework called Castling-ViT, which trains ViTs using both linear-angular attention and masked softmax-based quadratic attention, but then switches to having only linear-angular attention during inference. Our Castling-ViT leverages angular kernels to measure the similarities between queries and keys via spectral angles. And we further simplify it with two techniques: (1) a novel linear-angular attention mechanism: we decompose the angular kernels into linear terms and high-order residuals, and only keep the linear terms; and (2) we adopt two parameterized modules to approximate high-order residuals: a depthwise convolution and an auxiliary masked softmax attention to help learn global and local information, where the masks for softmax attention are regularized to gradually become zeros and thus incur no overhead during inference. Extensive experiments validate the effectiveness of our Castling-ViT, e.g., achieving up to a 1.8% higher accuracy or 40% MACs reduction on classification and 1.2 higher mAP on detection under comparable FLOPs, as compared to ViTs with vanilla softmax-based attentions. Project page is available at here.

Table of Contents

  • 1. Introduction
  • 2. Related Works
  • 3. The Proposed Methods
  • 3.1. Preliminary of Self-Attention
  • 3.2. The Proposed Castling-ViT Framework
  • 3.2.1 Linear-Angular Attention
  • 3.2.2 Switch Towards Linear-Angular Attention
  • 4. Experiments
  • 4.1. Experiment Settings
  • 4.2. Castling-ViT over SOTA Baselines
  • 4.3. Linear-Angular Attention over SOTA Baselines
  • 4.4. Ablation Studies of Castling-ViT
  • 4.5. Discussion on the Auxiliary Branch
  • 5. Conclusion
  • Acknowledgements
  • References

Knowls

  1. Knowl 1 — Castling-ViT training–inference switch

    model/method

    Castling-ViT is a Vision Transformer framework that combines an efficient linear-angular attention branch with an auxiliary quadratic masked-softmax attention branch during training, then removes the auxiliary branch for inference. The linear-angular branch remains active at inference and provides the low-complexity attention computation, while the auxiliary branch supplies information associated with the omitted high-order terms of the angular similarity function during early training. The design assumes that the remaining network can gradually learn these high-order components later in training, allowing the costly branch to disappear without degrading the final model accuracy.

  2. Knowl 2 — Angular kernel for spectral similarity

    equation

    For two nonzero token feature vectors xi,xj∈Rd\mathbf{x}_i,\mathbf{x}_j\in\mathbb{R}^{d}, Castling-ViT defines their spectral angle and angular similarity as

    θ(xi,xj)=arccos⁡(⟨xi,xj⟩∥xi∥2∥xj∥2),Sim⁡(xi,xj)=1−θ(xi,xj)π.\theta(\mathbf{x}_i,\mathbf{x}_j)=\arccos\left(\frac{\langle\mathbf{x}_i,\mathbf{x}_j\rangle}{\|\mathbf{x}_i\|_2\|\mathbf{x}_j\|_2}\right), \qquad \operatorname{Sim}(\mathbf{x}_i,\mathbf{x}_j)=1-\frac{\theta(\mathbf{x}_i,\mathbf{x}_j)}{\pi}.

    Here ⟨⋅,⋅⟩\langle\cdot,\cdot\rangle is the inner product, ∥⋅∥2\|\cdot\|_2 is the Euclidean norm, and θ∈[0,π]\theta\in[0,\pi], so Sim⁡∈[0,1]\operatorname{Sim}\in[0,1]. Aligned vectors have similarity near 11, whereas oppositely directed vectors have similarity near 00. The induced feature map ϕ\phi places every input on the unit sphere, because ∥ϕ(x)∥22=Sim⁡(x,x)=1\|\phi(\mathbf{x})\|_2^2=\operatorname{Sim}(\mathbf{x},\mathbf{x})=1, and satisfies

    ∥ϕ(xi)−ϕ(xj)∥22=2(1−Sim⁡(xi,xj))=2πθ(xi,xj).\|\phi(\mathbf{x}_i)-\phi(\mathbf{x}_j)\|_2^2 =2\bigl(1-\operatorname{Sim}(\mathbf{x}_i,\mathbf{x}_j)\bigr) =\frac{2}{\pi}\theta(\mathbf{x}_i,\mathbf{x}_j).

    Thus, angular separation in the original space becomes squared Euclidean distance in the induced feature space.

  3. Knowl 3 — Angular-kernel expansion and linear-angular terms

    equation

    Let q,k∈Rd\mathbf{q},\mathbf{k}\in\mathbb{R}^{d} be nonzero query and key vectors and define their normalized inner product by

    s(q,k)=qTk∥q∥2∥k∥2∈[−1,1].s(\mathbf{q},\mathbf{k})=\frac{\mathbf{q}^{\mathsf T}\mathbf{k}}{\|\mathbf{q}\|_2\|\mathbf{k}\|_2}\in[-1,1].

    The angular similarity has the expansion

    Sim⁡(q,k)=12+1πs(q,k)+1π∑r=1∞(2r)!22r(r!)2(2r+1)s(q,k)2r+1.\operatorname{Sim}(\mathbf{q},\mathbf{k}) =\frac{1}{2}+\frac{1}{\pi}s(\mathbf{q},\mathbf{k}) +\frac{1}{\pi}\sum_{r=1}^{\infty} \frac{(2r)!}{2^{2r}(r!)^2(2r+1)}s(\mathbf{q},\mathbf{k})^{2r+1}.

    The first two terms are the linear-angular terms and can be evaluated by changing the matrix order from (QKT)V(\mathbf{Q}\mathbf{K}^{\mathsf T})\mathbf{V} to Q(KTV)\mathbf{Q}(\mathbf{K}^{\mathsf T}\mathbf{V}). The infinite sum contains the nonlinear high-order residuals that would otherwise require quadratic computation in the number of tokens. Castling-ViT retains the linear-angular terms explicitly and approximates the residuals with learnable modules.

  4. Knowl 4 — Linear-angular attention with residual approximators

    model/method

    For NN tokens, let normalized query and key matrices be Q,K∈RN×d\mathbf{Q},\mathbf{K}\in\mathbb{R}^{N\times d} and let the value matrix be V∈RN×dv\mathbf{V}\in\mathbb{R}^{N\times d_v}. Castling-ViT computes its attention output approximately as

    H≈12V+1πQ(KTV)+MDWV+MSparseAttn.\mathbf{H} \approx \frac{1}{2}\mathbf{V} +\frac{1}{\pi}\mathbf{Q}(\mathbf{K}^{\mathsf T}\mathbf{V}) +\mathbf{M}_{\mathrm{DW}}\mathbf{V} +\mathbf{M}_{\mathrm{SparseAttn}}.

    The first term is the constant part of the angular expansion, and the second is the linear-angular term. MDW\mathbf{M}_{\mathrm{DW}} denotes the learnable depthwise-convolution operator applied to value tokens, while MSparseAttn\mathbf{M}_{\mathrm{SparseAttn}} denotes the output of the auxiliary masked-softmax branch. The depthwise convolution is intended to approximate nearby-token residual structure, and the sparse attention branch supplements nonlocal high-order structure. With feature dimensions treated as fixed, the matrix product Q(KTV)\mathbf{Q}(\mathbf{K}^{\mathsf T}\mathbf{V}) and the depthwise convolution scale linearly with NN; the paper reports that the depthwise-convolution MACs are negligible, below 1% of total MACs in the cited implementations. Query/key normalization is also applied to stabilize similarity computation.

  5. Knowl 5 — Thresholded masked-softmax training branch

    algorithm

    The auxiliary branch is constructed from the ordinary softmax attention matrix and a threshold ϵ>0\epsilon>0. For query and key matrices Q,K∈RN×d\mathbf{Q},\mathbf{K}\in\mathbb{R}^{N\times d}, the masked attention output is

    MSparseAttn(Q,K)=Mask⁡ϵ ⁣(Softmax⁡(QKT)),\mathbf{M}_{\mathrm{SparseAttn}}(\mathbf{Q},\mathbf{K}) =\operatorname{Mask}_{\epsilon}\!\left(\operatorname{Softmax}(\mathbf{Q}\mathbf{K}^{\mathsf T})\right),

    where softmax is applied row-wise and, elementwise,

    Mask⁡ϵ(x)={x,x>ϵ,0,x≤ϵ.\operatorname{Mask}_{\epsilon}(x)= \begin{cases} x,&x>\epsilon,\\ 0,&x\leq\epsilon. \end{cases}

    The resulting sparse attention output is added to the linear-angular output only during training. The threshold keeps high softmax scores, which are interpreted as strong local or salient interactions, while suppressing the remaining scores. Image-classification experiments use a fixed threshold of ϵ=0.02\epsilon=0.02 as one example; fixed and dynamic threshold schedules were both reported to give similar behavior for a given task.

  6. Knowl 6 — Castling-LeViT attention architecture

    model/method

    For the Castling-LeViT instantiation, the authors use post-query pooling rather than pre-query pooling in downsampling layers and retain residual query connections. Post-query pooling means that token or feature pooling is performed after the linear projections that produce the attention inputs, whereas pre-query pooling performs pooling during those projections. Their ImageNet attention-design ablations found that token pooling and feature pooling can each reduce classification accuracy, and that post-query pooling with residual connections performs better than pre-query pooling for the downsampling design used in their models. This design is intended to avoid the feature bottlenecks associated with aggressively shrinking feature dimensions and the token bottlenecks associated with aggressively reducing token counts.

  7. Knowl 7 — Evaluation protocol across three vision tasks

    experimental setup

    Castling-ViT is evaluated on ImageNet classification, COCO object detection, and ADE20K semantic segmentation. ImageNet contains approximately 1.2 million training images and 50,000 validation images; COCO uses 118,000 training and 5,000 validation images; ADE20K uses 20,000, 2,000, and 3,000 images for training, validation, and testing. Classification experiments apply Castling-ViT to LeViT, MViTv2, and DeiT models. Detection experiments use efficient detector backbones derived from ESNet or LCNet with transformer blocks in their final stages, and segmentation experiments use Mask2Former with a ViT-Base or Castling-ViT-Base backbone.

    For classification, the models are trained for 1,000 epochs with SGD, momentum 0.90.9, weight decay 2×10−52\times10^{-5}, 64 V100 GPUs, an 11-epoch warm-up, and a learning-rate decay factor of 0.98750.9875 per epoch; distillation uses a teacher with 85.5% accuracy. Detection uses SGD with momentum 0.90.9, weight decay 4×10−54\times10^{-5}, eight V100 GPUs, and batch size 80 per card, following the PicoDet recipe. Segmentation follows the Mask2Former training recipe, with MAE ImageNet pretraining used in the specified pretrained condition. Accuracy is measured with top-1/top-5 accuracy, AP metrics, or mIoU/mAcc/pAcc, while efficiency is measured with parameters and inference MACs or FLOPs.

  8. Knowl 8 — ImageNet accuracy–efficiency improvements

    empirical result

    On ImageNet, Castling-ViT improves the accuracy–MAC tradeoff across models ranging from approximately 0.4G to 17G MACs. Representative reported results are:

    • Castling-LeViT-128: 10.5M parameters, 0.49G MACs, 79.6% top-1 and 94.6% top-5 accuracy; Castling-LeViT-192: 12.7M parameters, 0.82G MACs, 81.3% top-1 and 95.5% top-5.
    • Castling-LeViT-256: 22.0M parameters, 1.40G MACs, 82.6% top-1 and 96.1% top-5; the corresponding 82.6% LeViT result uses 2.35G MACs, giving approximately 40% fewer MACs at comparable accuracy.
    • Castling-LeViT-384: 45.8M parameters, 2.90G MACs, 83.7% top-1 and 96.7% top-5.
    • Castling-MViTv2-T: 24.1M parameters, 4.50G MACs, 84.1% top-1 and 96.8% top-5, compared with 82.3% top-1 for MViTv2-T at 24.0M parameters and 4.68G MACs.
    • Castling-MViTv2-S: 34.7M parameters, 6.95G MACs, 84.6% top-1 and 97.0% top-5, compared with 83.6% top-1 for MViTv2-S at the same reported parameter and MAC counts.
    • Castling-MViTv2-B: 51.9M parameters, 9.82G MACs, 85.0% top-1 and 97.2% top-5, compared with 84.4% top-1 for MViTv2-B at 51.2M parameters and 10.07G MACs.

    Across the four reported MAC regimes—below 1G, 1–3G, 3–10G, and above 10G—the authors report top-1 improvements over comparable baselines ranging from 0.5–6.6%, 1.0–8.1%, 1.0–4.1%, and 0.6–2.6%, respectively.

  9. Knowl 9 — Transfer to detection and segmentation

    empirical result

    Castling-ViT transfers to dense prediction tasks while retaining its inference-efficiency advantage. On COCO detection, representative results include Castling-ViT-S-320 with 3.25M parameters, 0.62G MACs, 28.1 mAP, 42.3 AP50, and 29.2 AP75; Castling-ViT-M-416 with 6.01M parameters, 2.00G MACs, 34.0 mAP, 49.5 AP50, and 35.9 AP75; and Castling-ViT-L-416 with 9.32M parameters, 3.03G MACs, 35.0 mAP, 50.7 AP50, and 37.2 AP75. The LCNet-based Castling-ViT-L-416 variant reaches 37.3 mAP, 53.4 AP50, and 39.6 AP75 at 13.10M parameters and 5.31G MACs. Compared with YOLOv5, YOLOX, MobileDet, and FBNetV5 at comparable or lower MACs, the reported mAP gains are approximately 6.0, 2.2–2.3, 4.0–5.9, and 3.1–4.0 points, respectively.

    On ADE20K with Mask2Former, ViT-Base and Castling-ViT-Base both use 118M parameters. Without MAE pretraining, Castling-ViT-Base reduces total MACs from 229G to 195G and backbone MACs from 182G to 147G, while improving mIoU/mAcc/pAcc from 34.54/46.36/75.84 to 34.67/46.47/76.20. With MAE pretraining, it retains the same MAC reduction and improves these metrics from 47.92/61.00/83.02 to 48.44/61.82/83.29.

  10. Knowl 10 — Component ablations and the castling phenomenon

    empirical result

    COCO ablations show that each proposed component contributes to the final detector. In a Castling-ViT-S-320 model with an LCNet backbone, the LCNet baseline obtains 25.5 mAP at 0.48G MACs; replacing the final stages with transformer blocks gives 25.9 mAP at 0.69G MACs; adding linear-angular attention gives 26.2 mAP at 0.62G MACs; adding depthwise convolution gives 26.8 mAP at the same 0.62G MACs; and adding the auxiliary sparse attention gives 27.0 mAP at 0.62G MACs. In the larger Castling-ViT-M-416 setting, the corresponding mAP values are 34.3, 34.9, 34.8, 35.1, and 35.3, while MACs decrease from 3.51G for the transformer-block backbone to 3.15G after linear-angular attention.

    The authors also observe that the nonzero entries in the auxiliary masks increase or remain useful during early and middle COCO training but gradually fall to zero near the end of training in both small and medium models. A synthetic two-layer-DNN experiment supports the proposed explanation: the DNN first fits low-frequency components of the angular similarity and later fits higher-frequency components, so the costly auxiliary attention is useful early but can be removed after the rest of the network has learned the omitted high-order behavior. A stated limitation is that directly replacing convolutional backbones with transformer blocks can increase MACs enough that LCNet-ViT may still have a slightly worse accuracy–efficiency tradeoff than a pure LCNet backbone; designing more efficient ViT-based backbones remains open.

Coverage note — No substantial contributed material was omitted; minor implementation details and the full baseline tables were compressed while retaining the proposed mechanism, architectural choice, evaluation protocol, principal results, ablations, and stated limitation.

References

  1. 1.Alaaeldin Ali, Hugo Touvron, Mathilde Caron, Piotr Bojanowski, Matthijs Douze, Armand Joulin, Ivan Laptev, Natalia Neverova, Gabriel Synnaeve, Jakob Verbeek, et al. Xcit: Cross-covariance image transformers. Advances in neural information processing systems, 34:20014–20027, 2021. 2, 13
  2. 2.Moab Arar, Ariel Shamir, and Amit H Bermano. Learned queries for efficient local attention. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10841–10852, 2022. 1, 2, 13
  3. 3.Iz Beltagy, Matthew E Peters, and Arman Cohan. Longformer: The long-document transformer. arXiv preprint arXiv:2004.05150, 2020. 1, 13
  4. 4.Daniel Bolya, Cheng-Yang Fu, Xiaoliang Dai, Peizhao Zhang, Christoph Feichtenhofer, and Judy Hoffman. Token merging: Your vit but faster. arXiv preprint arXiv:2210.09461, 2022. 2
  5. 5.Daniel Bolya, Cheng-Yang Fu, Xiaoliang Dai, Peizhao Zhang, and Judy Hoffman. Hydra attention: Efficient attention with many heads. arXiv preprint arXiv:2209.07484, 2022. 1, 2, 5, 13
  6. 6.Han Cai, Chuang Gan, and Song Han. Efficientvit: Enhanced linear attention for high-resolution low-computation visual recognition. arXiv preprint arXiv:2205.14756, 2022. 1, 2, 3, 5, 7, 13, 14
  7. 7.Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. End-to-end object detection with transformers. In European conference on computer vision, pages 213–229. Springer, 2020. 13
  8. 8.Beidi Chen, Tri Dao, Eric Winsor, Zhao Song, Atri Rudra, and Christopher Re. Scatterbrain: Unifying sparse and low-rank attention. Advances in Neural Information Processing Systems, 34:17413–17426, 2021. 1, 2, 5, 13
  9. 9.Chun-Fu Chen, Quanfu Fan, and Rameswar Panda. Crossvit: Cross-attention multi-scale vision transformer for image classification. arXiv preprint arXiv:2103.14899, 2021. 2
  10. 10.Minghao Chen, Houwen Peng, Jianlong Fu, and Haibin Ling. Autoformer: Searching transformers for visual recognition. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 12270–12280, 2021. 6
  11. 11.Bowen Cheng, Ishan Misra, Alexander G Schwing, Alexander Kirillov, and Rohit Girdhar. Masked-attention mask transformer for universal image segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1290–1299, 2022. 3, 5, 6, 7, 13
  12. 12.Bowen Cheng, Alex Schwing, and Alexander Kirillov. Per-pixel classification is not all you need for semantic segmentation. Advances in Neural Information Processing Systems, 34:17864–17875, 2021. 13
  13. 13.Krzysztof Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Andreea Gane, Tamas Sarlos, Peter Hawkins, Jared Davis, Afroz Mohiuddin, Lukasz Kaiser, et al. Rethinking attention with performers. arXiv preprint arXiv:2009.14794, 2020. 3
  14. 14.Krzysztof Marcin Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Andreea Gane, Tamas Sarlos, Peter Hawkins, Jared Quincy Davis, Afroz Mohiuddin, Lukasz Kaiser, David Benjamin Belanger, Lucy J Colwell, and Adrian Weller. Rethinking attention with performers. In International Conference on Learning Representations, 2021. 1, 2, 13
  15. 15.Cheng Cui, Tingquan Gao, Shengyu Wei, Yuning Du, Ruoyu Guo, Shuilong Dong, Bin Lu, Ying Zhou, Xueying Lv, Qiwen Liu, et al. Pp-lcnet: A lightweight cpu convolutional neural network. arXiv preprint arXiv:2109.15099, 2021. 5, 6
  16. 16.Xiaoliang Dai, Alvin Wan, P. Zhang, B. Wu, Zijian He, Zhen Wei, K. Chen, Yuandong Tian, Matthew E. Yu, Peter Vajda, and J. Gonzalez. Fbnetv3: Joint architecture-recipe search using neural acquisition function. ArXiv, abs/2006.02049, 2020. 2
  17. 17.Jyotikrishna Dass, Shang Wu, Huihong Shi, Chaojian Li, Zhifan Ye, Zhongfeng Wang, and Yingyan Lin. Vitality: Unifying low-rank and sparse approximation for vision transformer acceleration with a linear taylor attention. In The 29th IEEE International Symposium on High-Performance Computer Architecture (HPCA 2023), 2023. 2, 5, 7, 14
  18. 18.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition, pages 248–255. Ieee, 2009. 1, 5
  19. 19.Xiaoyi Dong, Jianmin Bao, Dongdong Chen, Weiming Zhang, Nenghai Yu, Lu Yuan, Dong Chen, and Baining Guo. Cswin transformer: A general vision transformer backbone with cross-shaped windows. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12124–12134, 2022. 6
  20. 20.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition at scale. In International Conference on Learning Representations, 2021. 1, 2, 3
  21. 21.Haoqi Fan, Bo Xiong, Karttikeya Mangalam, Yanghao Li, Zhicheng Yan, Jitendra Malik, and Christoph Feichtenhofer. Multiscale vision transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 6824–6835, 2021. 2
  22. 22.Peng Gao, Teli Ma, Hongsheng Li, Jifeng Dai, and Yu Qiao. Convmae: Masked convolution meets masked autoencoders. arXiv preprint arXiv:2205.03892, 2022. 13
  23. 23.Zheng Ge, Songtao Liu, Feng Wang, Zeming Li, and Jian Sun. Yolox: Exceeding yolo series in 2021. arXiv preprint arXiv:2107.08430, 2021. 6, 7, 14
  24. 24.Ben Graham, Alaaeldin El-Nouby, Hugo Touvron, Pierre Stock, Armand Joulin, Herve Jegou, and Matthijs Douze. Levit: a vision transformer in convnet’s clothing for faster inference. arXiv preprint arXiv:2104.01136, 2021. 2, 3, 5, 6, 13, 14
  25. 25.Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollar, and Ross Girshick. Masked autoencoders are scalable vision learners. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 16000–16009, 2022. 6, 7, 13
  26. 26.Byeongho Heo, Sangdoo Yun, Dongyoon Han, Sanghyuk Chun, Junsuk Choe, and Seong Joon Oh. Rethinking spatial dimensions of vision transformers. arXiv preprint arXiv:2103.16302, 2021. 2
  27. 27.Paul Honeine and Cedric Richard. The angular kernel in machine learning for hyperspectral data classification. In 2010 2nd Workshop on Hyperspectral Image and Signal Processing: Evolution in Remote Sensing, pages 1–4. IEEE, 2010. 4
  28. 28.Andrew Howard, Mark Sandler, Grace Chu, Liang-Chieh Chen, Bo Chen, Mingxing Tan, Weijun Wang, Yukun Zhu, Ruoming Pang, Vijay Vasudevan, et al. Searching for mobilenetv3. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 1314–1324, 2019. 2
  29. 29.Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and François Fleuret. Transformers are rnns: Fast autoregressive transformers with linear attention. In International Conference on Machine Learning, pages 5156–5165. PMLR, 2020. 1, 2, 3, 4, 13
  30. 30.Nirmal Keshava. Distance metrics and band selection in hyperspectral processing with applications to material identification and spectral libraries. IEEE Transactions on Geoscience and remote sensing, 42(7):1552–1565, 2004. 4
  31. 31.Kyungmin Kim, Bichen Wu, Xiaoliang Dai, Peizhao Zhang, Zhicheng Yan, Peter Vajda, and Seon Joo Kim. Rethinking the self-attention in vision transformers. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3071–3075, 2021. 4, 5
  32. 32.Nikita Kitaev, Łukasz Kaiser, and Anselm Levskaya. Reformer: The efficient transformer. arXiv preprint arXiv:2001.04451, 2020. 1, 13
  33. 33.Yanghao Li, Hanzi Mao, Ross Girshick, and Kaiming He. Exploring plain vision transformer backbones for object detection. arXiv preprint arXiv:2203.16527, 2022. 13
  34. 34.Yanghao Li, Chao-Yuan Wu, Haoqi Fan, Karttikeya Mangalam, Bo Xiong, Jitendra Malik, and Christoph Feichtenhofer. Mvitv2: Improved multiscale vision transformers for classification and detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4804–4814, 2022. 3, 5, 6, 13, 14
  35. 35.Yanyu Li, Geng Yuan, Yang Wen, Eric Hu, Georgios Evangelidis, Sergey Tulyakov, Yanzhi Wang, and Jian Ren. Efficientformer: Vision transformers at mobilenet speed. arXiv preprint arXiv:2206.01191, 2022. 2
  36. 36.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollar, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In European conference on computer vision, pages 740–755. Springer, 2014. 1, 5
  37. 37.Jing Liu, Zizheng Pan, Haoyu He, Jianfei Cai, and Bohan Zhuang. Ecoformer: Energy-saving attention with linear complexity. In NeurIPS, 2022. 1, 2, 4, 13
  38. 38.Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 10012–10022, 2021. 1, 2, 6, 13, 14
  39. 39.Jiachen Lu, Jinghan Yao, Junge Zhang, Xiatian Zhu, Hang Xu, Weiguo Gao, Chunjing Xu, Tao Xiang, and Li Zhang. Soft: softmax-free transformer with linear complexity. Advances in Neural Information Processing Systems, 34:21297–21309, 2021. 1, 2, 13
  40. 40.Sachin Mehta and Mohammad Rastegari. Mobilevit: lightweight, general-purpose, and mobile-friendly vision transformer. arXiv preprint arXiv:2110.02178, 2021. 2, 6
  41. 41.John Moody and Christian Darken. Learning with localized receptive fields. Yale Univ., Department of Computer Science, 1988. 4
  42. 42.Zipeng Qin, Jianbo Liu, Xiaolin Zhang, Maoqing Tian, Aojun Zhou, Shuai Yi, and Hongsheng Li. Pyramid fusion transformer for semantic segmentation. arXiv preprint arXiv:2201.04019, 2022. 3, 13
  43. 43.Zhen Qin, Weixuan Sun, Hui Deng, Dongxu Li, Yunshen Wei, Baohong Lv, Junjie Yan, Lingpeng Kong, and Yiran Zhong. cosformer: Rethinking softmax in attention. In International Conference on Learning Representations, 2022. 1, 2, 13
  44. 44.Ilija Radosavovic, Raj Prateek Kosaraju, Ross Girshick, Kaiming He, and Piotr Dollar. Designing network design spaces. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10428–10436, 2020. 6
  45. 45.Ali Rahimi and Benjamin Recht. Random features for large-scale kernel machines. Advances in neural information processing systems, 20, 2007. 3
  46. 46.Yongming Rao, Wenliang Zhao, Benlin Liu, Jiwen Lu, Jie Zhou, and Cho-Jui Hsieh. Dynamicvit: Efficient vision transformers with dynamic token sparsification. Advances in neural information processing systems, 34:13937–13949, 2021. 2
  47. 47.Mark Sandler, Andrew Howard, Menglong Zhu, Andrey Zhmoginov, and Liang-Chieh Chen. Mobilenetv2: Inverted residuals and linear bottlenecks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 4510–4520, 2018. 6
  48. 48.Robert R. Schaller. Moore’s law: past, present and future. IEEE spectrum, 1997. 14
  49. 49.Zhuoran Shen, Mingyuan Zhang, Haiyu Zhao, Shuai Yi, and Hongsheng Li. Efficient attention: Attention with linear complexities. In Proceedings of the IEEE/CVF winter conference on applications of computer vision, pages 3531–3539, 2021. 3, 7
  50. 50.Chen Sun, Abhinav Shrivastava, Saurabh Singh, and Abhinav Gupta. Revisiting unreasonable effectiveness of data in deep learning era. In ICCV, 2017. 2
  51. 51.Mingxing Tan and Quoc Le. Efficientnet: Rethinking model scaling for convolutional neural networks. In International conference on machine learning, pages 6105–6114. PMLR, 2019. 6
  52. 52.Mingxing Tan, Ruoming Pang, and Quoc V Le. Efficientdet: Scalable and efficient object detection. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10781–10790, 2020. 6, 7
  53. 53.Shitao Tang, Jiahui Zhang, Siyu Zhu, and Ping Tan. Quadtree attention for vision transformers. In International Conference on Learning Representations, 2022. 1
  54. 54.Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Herve Jegou. Training data-efficient image transformers & distillation through attention. In Proceedings of the 38th International Conference on Machine Learning, volume 139 of Proceedings of Machine Learning Research, pages 10347–10357. PMLR, 18–24 Jul 2021. 2, 5, 6, 14
  55. 55.Hugo Touvron, Matthieu Cord, Alexandre Sablayrolles, Gabriel Synnaeve, and Herve Jegou. Going deeper with image transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 32–42, 2021. 6
  56. 56.Zhengzhong Tu, Hossein Talebi, Han Zhang, Feng Yang, Peyman Milanfar, Alan Bovik, and Yinxiao Li. Maxvit: Multi-axis vision transformer. In Computer Vision–ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part XXIV, pages 459–479. Springer, 2022. 2, 13
  57. 57.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Advances in Neural Information Processing Systems, volume 30. Curran Associates, Inc., 2017. 1, 3
  58. 58.Sinong Wang, Belinda Z. Li, Madian Khabsa, Han Fang, and Hao Ma. Linformer: Self-attention with linear complexity, 2020. 1, 2, 3, 13
  59. 59.Wenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan, Kaitao Song, Ding Liang, Tong Lu, Ping Luo, and Ling Shao. Pyramid vision transformer: A versatile backbone for dense prediction without convolutions. arXiv preprint arXiv:2102.12122, 2021. 2, 6, 14
  60. 60.Wenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan, Kaitao Song, Ding Liang, Tong Lu, Ping Luo, and Ling Shao. Pyramid vision transformer: A versatile backbone for dense prediction without convolutions. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 568–578, October 2021. 3, 13
  61. 61.Wenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan, Kaitao Song, Ding Liang, Tong Lu, Ping Luo, and Ling Shao. Pvt v2: Improved baselines with pyramid vision transformer. Computational Visual Media, 8(3):415–424, 2022. 6
  62. 62.Bichen Wu, Chaojian Li, Hang Zhang, Xiaoliang Dai, Peizhao Zhang, Matthew Yu, Jialiang Wang, Yingyan Lin, and Peter Vajda. Fbnetv5: Neural architecture search for multiple tasks in one run. arXiv preprint arXiv:2111.10007, 2021. 6, 7
  63. 63.Haiping Wu, Bin Xiao, Noel Codella, Mengchen Liu, Xiyang Dai, Lu Yuan, and Lei Zhang. Cvt: Introducing convolutions to vision transformers. arXiv preprint arXiv:2103.15808, 2021. 2
  64. 64.Zhanghao Wu, Zhijian Liu, Ji Lin, Yujun Lin, and Song Han. Lite transformer with long-short range attention. In International Conference on Learning Representations, 2020. 1
  65. 65.Yunyang Xiong, Hanxiao Liu, Suyog Gupta, Berkin Akin, Gabriel Bender, Pieter-Jan Kindermans, Mingxing Tan, Vikas Singh, and Bo Chen. Mobiledets: Searching for object detection architectures for mobile accelerators. arXiv preprint arXiv:2004.14525, 2020. 2, 6, 7
  66. 66.Yunyang Xiong, Zhanpeng Zeng, Rudrasis Chakraborty, Mingxing Tan, Glenn Fung, Yin Li, and Vikas Singh. Nystromformer: A nystrom-based algorithm for approximating self-attention. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, pages 14138–14148, 2021. 1, 2, 13
  67. 67.Zhi-Qin John Xu, Yaoyu Zhang, Tao Luo, Yanyang Xiao, and Zheng Ma. Frequency principle: Fourier analysis sheds light on deep neural networks. arXiv preprint arXiv:1901.06523, 2019. 2, 5, 8
  68. 68.Hongxu Yin, Arash Vahdat, Jose M Alvarez, Arun Mallya, Jan Kautz, and Pavlo Molchanov. A-vit: Adaptive tokens for efficient vision transformer. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10809–10818, 2022. 2
  69. 69.Haoran You, Zhanyi Sun, Huihong Shi, Zhongzhi Yu, Yang Zhao, Yongan Zhang, Chaojian Li, Baopu Li, and Yingyan Lin. Vitcod: Vision transformer acceleration via dedicated algorithm and accelerator co-design. In The 29th IEEE International Symposium on High-Performance Computer Architecture (HPCA 2023), 2023. 1
  70. 70.Guanghua Yu, Qinyao Chang, Wenyu Lv, Chang Xu, Cheng Cui, Wei Ji, Qingqing Dang, Kaipeng Deng, Guanzhong Wang, Yuning Du, et al. Pp-picodet: A better real-time object detector on mobile devices. arXiv preprint arXiv:2111.00902, 2021. 5, 6, 7
  71. 71.Zhongzhi Yu, Yonggan Fu, Sicheng Li, Chaojian Li, and Yingyan Lin. Mia-former: Efficient and robust vision transformers via multi-grained input-adaptation. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 36, pages 8962–8970, 2022. 2
  72. 72.Li Yuan, Yunpeng Chen, Tao Wang, Weihao Yu, Yujun Shi, Zi-Hang Jiang, Francis EH Tay, Jiashi Feng, and Shuicheng Yan. Tokens-to-token vit: Training vision transformers from scratch on imagenet. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 558–567, 2021. 2, 3
  73. 73.Yuhui Yuan, Rao Fu, Lang Huang, Weihong Lin, Chao Zhang, Xilin Chen, and Jingdong Wang. Hrformer: High-resolution vision transformer for dense predict. Advances in Neural Information Processing Systems, 34:7281–7293, 2021. 6
  74. 74.Zixiao Zhang, Xiaoqiang Lu, Guojin Cao, Yuting Yang, Licheng Jiao, and Fang Liu. Vit-yolo: Transformer-based yolo for object detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 2799–2808, 2021. 13
  75. 75.Bolei Zhou, Hang Zhao, Xavier Puig, Sanja Fidler, Adela Barriuso, and Antonio Torralba. Scene parsing through ade20k dataset. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 633–641, 2017. 5
  76. 76.Chen Zhu, Wei Ping, Chaowei Xiao, Mohammad Shoeybi, Tom Goldstein, Anima Anandkumar, and Bryan Catanzaro. Long-short transformer: Efficient transformers for language and vision. Advances in Neural Information Processing Systems, 34, 2021. 1, 2
  77. 77.Xizhou Zhu, Weijie Su, Lewei Lu, Bin Li, Xiaogang Wang, and Jifeng Dai. Deformable detr: Deformable transformers for end-to-end object detection. arXiv preprint arXiv:2010.04159, 2020. 3, 13

Citation

MLA
You, H., et al. “Castling-ViT: Compressing Self-Attention via Switching Towards Linear-Angular Attention at Vision Transformer Inference”. arXiv, 2022, http://arxiv.org/abs/2211.10526v5.
APA
You, H., Xiong, Y., Dai, X., Wu, B., Zhang, P., Fan, H., Vajda, P., & Lin, Y. C. (2022). Castling-ViT: Compressing Self-Attention via Switching Towards Linear-Angular Attention at Vision Transformer Inference. arXiv. http://arxiv.org/abs/2211.10526v5
Chicago
You, H., Y. Xiong, X. Dai, et al. 2022. “Castling-ViT: Compressing Self-Attention via Switching Towards Linear-Angular Attention at Vision Transformer Inference”. arXiv. http://arxiv.org/abs/2211.10526v5.
Harvard
You, H. et al. (2022) “Castling-ViT: Compressing Self-Attention via Switching Towards Linear-Angular Attention at Vision Transformer Inference”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2211.10526v5.
Vancouver
1. You H, Xiong Y, Dai X, Wu B, Zhang P, Fan H, Vajda P, Lin YC (2022) Castling-ViT: Compressing Self-Attention via Switching Towards Linear-Angular Attention at Vision Transformer Inference. arXiv

BibTeX

@article{you2022castling,
  title = {Castling-ViT: Compressing Self-Attention via Switching Towards Linear-Angular Attention at Vision Transformer Inference},
  author = {You, Haoran and Xiong, Yunyang and Dai, Xiaoliang and Wu, Bichen and Zhang, Peizhao and Fan, Haoqi and Vajda, Peter and Lin, Yingyan Celine},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2211.10526v5},
  eprint = {2211.10526}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE