Castling-ViT: Compressing Self-Attention via Switching Towards Linear-Angular Attention at Vision Transformer Inference
Haoran YouYunyang XiongXiaoliang DaiBichen WuPeizhao ZhangHaoqi FanPeter VajdaYingyan Celine Lin
Proposes a training framework that pairs linear-angular attention with a decaying auxiliary softmax branch, enabling vision transformers to switch to efficient linear-complexity inference without sacrificing accuracy across classification and detection tasks.
Vision transformers deliver leading accuracy across visual recognition tasks, but their core self-attention mechanisms suffer from computational complexity that grows quadratically with image resolution. While previous efficient designs adopted local windows or linear approximations to reduce computation, they often sacrificed the model's ability to capture global context or fine local details. This runtime inefficiency creates a major deployment bottleneck, particularly for high-resolution vision applications on resource-constrained platforms.
The article evaluates Castling-ViT, a novel framework designed to close the accuracy gap between efficient linear attention and standard quadratic self-attention without introducing additional inference overhead. Castling-ViT trains vision transformers using both an efficient linear-angular attention mechanism and an auxiliary masked quadratic attention branch, then eliminates the auxiliary branch entirely during inference—a switch analogous to the castling move in chess.
To construct this approach, the authors decomposed spectral angular similarity kernels into linear terms and higher-order residual terms. The linear terms are retained for lightweight computation, while the non-linear residuals are approximated using a depthwise convolution alongside an auxiliary masked softmax attention module. During training, a thresholding regularization causes the auxiliary attention masks to naturally decay to zero. The evaluation assessed the framework across standard benchmarks: ImageNet for image classification, COCO for object detection, and ADE20K for semantic segmentation, integrating the mechanism into popular baseline architectures.
The experimental findings show substantial improvements in efficiency and accuracy. In image classification benchmarks, Castling-ViT delivered up to a 1.8% top-1 accuracy improvement under comparable computational budgets or up to a 40% reduction in multiply-accumulate operations while maintaining baseline accuracy. In object detection tasks on the COCO dataset, Castling-ViT improved average precision by up to 1.2 to 6.0 points compared to alternative convolutional and transformer baselines under similar operation counts. In semantic segmentation on ADE20K, integrating Castling-ViT into standard segmentation backbones achieved 15% to 19% reductions in computational operations while matching or exceeding baseline segmentation quality. Furthermore, ablation experiments confirmed that the proposed linear-angular kernel outperformed five other standard linear attention kernels by up to 4.6% in detection accuracy.
These results demonstrate that vision models do not need to choose between global context modeling and operational efficiency. Deploying linear-angular attention enables lower runtime latency, reduced hardware resource consumption, and lower operational costs for high-resolution visual processing. The findings challenge the conventional assumption that linear attention must underperform traditional quadratic attention, proving that auxiliary training mechanisms can effectively transfer complex feature learning into streamlined runtime architectures.
Organizations seeking to deploy vision transformers in resource-limited or low-latency environments should consider adopting linear-angular attention modules as drop-in replacements for standard self-attention. For development pipelines, engineering teams should evaluate the castling training recipe when training or fine-tuning transformer backbones for detection and segmentation tasks.
Confidence in these findings is supported by consistent empirical improvements across multiple architectures and standard vision benchmarks. However, a key limitation highlighted by the authors is that directly swapping transformer blocks into lightweight convolutional backbones can still introduce overhead compared to purely convolutional designs, as overall network structures may not yet be optimal. Further architectural exploration and on-device hardware profiling are recommended before deploying these models in strictly resource-bounded production systems.
- Paper: An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale, Alexey Dosovitskiy et al. (2021). Introduces the standard Vision Transformer architecture whose quadratic self-attention complexity Castling-ViT seeks to compress and optimize.
- Paper: On the Relationship between Self-Attention and Convolutional Layers, Jean-Baptiste Cordonnier et al. (2020). Establishes the theoretical and functional relationship between self-attention and convolutions, providing essential conceptual backing for approximating attention residuals with depthwise convolutions.
- Paper: CvT: Introducing Convolutions to Vision Transformers, Haiping Wu et al. (2021). Demonstrates how integrating depthwise convolutions into transformer attention blocks captures local context, directly motivating hybrid local-global attention approximations.
- Paper: Twins: Revisiting the Design of Spatial Attention in Vision Transformers, Xiangxiang Chu et al. (2021). Explores the separation and balance of local and global attention mechanisms to reduce computational complexity in vision transformers.
- Paper: Training data-efficient image transformers & distillation through attention, Hugo Touvron et al. (2021). Presents core training protocols and distillation methods for data-efficient ViT backbones used as baselines for efficient vision transformer design.
- Paper: TransNeXt: Robust Foveal Visual Perception for Vision Transformers, Dai Shi (2024). Extends efficient vision transformer design by combining multi-scale local and global perception via aggregated attention and convolutional gating.
- Paper: FFT-Based Dynamic Token Mixer for Vision, Yuki Tatsunami et al. (2024). Explores frequency-domain alternatives for linearizing global token mixing to overcome the quadratic cost of standard multi-head self-attention.
- Paper: Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space Model, Lianghui Zhu et al. (2024). Investigates bidirectional state-space models as an alternative paradigm to linearize sequence processing in visual representations without softmax attention.
- Paper: MambaVision: A Hybrid Mamba-Transformer Vision Backbone, Ali Hatamizadeh et al. (2025). Combines linear-complexity sequence modeling with selective self-attention blocks to balance global context with computational efficiency.
