AdaViT: Adaptive Vision Transformers for Efficient Image Recognition
Lingchen MengHengduo LiBor-Chun ChenShiyi LanZuxuan WuYu-Gang JiangSer-Nam Lim
Presents AdaViT, an adaptive computation framework that dynamically selects informative image patches, attention heads, and transformer blocks for each input to cut vision transformer inference costs by over half with negligible accuracy loss on ImageNet.
Vision transformers deliver exceptional performance in image recognition by modeling global context across image patches. However, their reliance on stacked multi-head self-attention layers results in heavy computational costs that grow quadratically with the number of input patches. Deploying standard vision transformers requires using the same intensive model for every input, even though simple images require substantially less processing than complex, cluttered scenes. Reducing this computational burden is essential for deploying modern vision systems efficiently without degrading predictive accuracy.
The article introduces and evaluates AdaViT (Adaptive Vision Transformer), an adaptive computation framework designed to improve the inference efficiency of vision transformers. The approach determines on a per-image basis which image patches to retain, which attention heads to activate, and which network blocks to skip entirely, thereby aligning computational expenditure with the complexity of each input.
The method attaches light-weight decision sub-networks before each transformer block to predict binary retention decisions dynamically during inference. To enable end-to-end training alongside the vision transformer backbone, the authors use a continuous relaxation technique known as Gumbel-Softmax to overcome the non-differentiability of discrete gating choices. The model optimizes a joint objective combining standard classification loss with a usage loss governed by target budget parameters. Credibility was established through extensive experiments on the ImageNet benchmark, encompassing approximately 1.2 million training images and 50,000 validation images, using a standard 19-block transformer backbone.
The findings demonstrate substantial operational efficiency gains. First, AdaViT improves inference efficiency by more than 2x compared to the standard vision transformer backbone, reducing required computation from 8.5 to 3.9 GFLOPs per image while incurring only a 0.8% decrease in Top-1 accuracy (from 81.9% to 81.1%). Second, AdaViT significantly outperforms random pruning and fine-tuned random baselines under identical compute budgets, achieving up to 48.1% higher accuracy than unstructured pruning and confirming the efficacy of its learned decision policies. Third, the framework flexibly accommodates varying computing constraints by adjusting budget hyperparameters, outperforming competing static vision architectures and convolutional models across diverse operating points. Finally, qualitative analyses confirm intuitive resource distribution: early layers retain more patches while later layers filter down to salient regions, and cluttered visual scenes automatically receive higher computational allocations than simple, object-centric images.
These results demonstrate that dynamic inference provides a viable path to halving computational processing requirements in real-world computer vision deployments. This efficiency drop translates directly into reduced cloud computing costs, lower latency, and better feasibility for resource-constrained edge systems. Furthermore, it shifts the design paradigm away from purely static architectures toward adaptive input-dependent networks without sacrificing accuracy.
Organizations evaluating vision transformers should consider implementing adaptive gating mechanisms to optimize runtime costs, tuning the budget hyperparameters to balance speed and accuracy requirements. For head selection, practitioners can choose full deactivation for maximum computational savings or partial deactivation if accuracy retention is paramount. However, decision-makers should note that a slight accuracy gap (0.8%) remains relative to fully unconstrained base models. Further validation on domain-specific downstream tasks, such as dense object detection or video analysis, is recommended before full-scale deployment in non-classification environments.
- Paper: An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale, Alexey Dosovitskiy et al. (2021). Introduces the standard Vision Transformer backbone whose heavy computational quadratic attention AdaViT dynamically prunes and accelerates.
- Paper: Training data-efficient image transformers & distillation through attention, Hugo Touvron et al. (2021). Establishes standard training protocols and compact Vision Transformer baselines that form the foundation for AdaViT's experimental evaluation on ImageNet.
- Paper: Are Sixteen Heads Really Better than One?, Paul Michel et al. (2019). Demonstrates the redundancy of individual attention heads in transformer architectures, directly motivating AdaViT's head-deactivation mechanism.
- Paper: Dynamic Convolution: Attention Over Convolution Kernels, Yinpeng Chen et al. (2019). Explores dynamic, input-dependent computation in deep vision models, providing foundational concepts for AdaViT's per-image adaptive gating strategy.
- Paper: Scaling Vision with Sparse Mixture of Experts, Carlos Riquelme et al. (2021). Pioneers conditional patch-routing in vision transformers, laying important conceptual groundwork for input-dependent token and sub-network selection.
- Paper: CF-ViT: A General Coarse-to-Fine Method for Vision Transformer, Mengzhao Chen et al. (2023). Extends input-adaptive vision transformer efficiency by employing a coarse-to-fine spatial patch splitting and routing paradigm.
- Paper: BiFormer: Vision Transformer with Bi-Level Routing Attention, Lei Zhu et al. (2023). Builds upon dynamic patch selection by introducing bi-level routing attention to adaptively sparsify self-attention based on visual content.
- Paper: Learning Adaptive and View-Invariant Vision Transformer for Real-Time UAV Tracking, Yongxin Li et al. (2024). Applies dynamic block-skipping transformer mechanisms to real-time UAV visual tracking under strict onboard edge constraints.
- Paper: Compressing Transformers: Features Are Low-Rank, but Weights Are Not!, Hao Yu et al. (2023). Explores complementary post-hoc vision transformer compression by leveraging low-rank internal feature redundancies across layers.
- Paper: VisionZip: Longer is Better but Not Necessary in Vision Language Models, Senqiao Yang et al. (2025). Generalizes visual token pruning and merging principles to reduce token redundancy in large multimodal vision-language architectures.
