AdaViT: Adaptive Vision Transformers for Efficient Image Recognition

Lingchen MengHengduo LiBor-Chun ChenShiyi LanZuxuan WuYu-Gang JiangSer-Nam Lim

article2022CVPR349 citations

Presents AdaViT, an adaptive computation framework that dynamically selects informative image patches, attention heads, and transformer blocks for each input to cut vision transformer inference costs by over half with negligible accuracy loss on ImageNet.

Listen

Vision transformers deliver exceptional performance in image recognition by modeling global context across image patches. However, their reliance on stacked multi-head self-attention layers results in heavy computational costs that grow quadratically with the number of input patches. Deploying standard vision transformers requires using the same intensive model for every input, even though simple images require substantially less processing than complex, cluttered scenes. Reducing this computational burden is essential for deploying modern vision systems efficiently without degrading predictive accuracy.

The article introduces and evaluates AdaViT (Adaptive Vision Transformer), an adaptive computation framework designed to improve the inference efficiency of vision transformers. The approach determines on a per-image basis which image patches to retain, which attention heads to activate, and which network blocks to skip entirely, thereby aligning computational expenditure with the complexity of each input.

The method attaches light-weight decision sub-networks before each transformer block to predict binary retention decisions dynamically during inference. To enable end-to-end training alongside the vision transformer backbone, the authors use a continuous relaxation technique known as Gumbel-Softmax to overcome the non-differentiability of discrete gating choices. The model optimizes a joint objective combining standard classification loss with a usage loss governed by target budget parameters. Credibility was established through extensive experiments on the ImageNet benchmark, encompassing approximately 1.2 million training images and 50,000 validation images, using a standard 19-block transformer backbone.

The findings demonstrate substantial operational efficiency gains. First, AdaViT improves inference efficiency by more than 2x compared to the standard vision transformer backbone, reducing required computation from 8.5 to 3.9 GFLOPs per image while incurring only a 0.8% decrease in Top-1 accuracy (from 81.9% to 81.1%). Second, AdaViT significantly outperforms random pruning and fine-tuned random baselines under identical compute budgets, achieving up to 48.1% higher accuracy than unstructured pruning and confirming the efficacy of its learned decision policies. Third, the framework flexibly accommodates varying computing constraints by adjusting budget hyperparameters, outperforming competing static vision architectures and convolutional models across diverse operating points. Finally, qualitative analyses confirm intuitive resource distribution: early layers retain more patches while later layers filter down to salient regions, and cluttered visual scenes automatically receive higher computational allocations than simple, object-centric images.

These results demonstrate that dynamic inference provides a viable path to halving computational processing requirements in real-world computer vision deployments. This efficiency drop translates directly into reduced cloud computing costs, lower latency, and better feasibility for resource-constrained edge systems. Furthermore, it shifts the design paradigm away from purely static architectures toward adaptive input-dependent networks without sacrificing accuracy.

Organizations evaluating vision transformers should consider implementing adaptive gating mechanisms to optimize runtime costs, tuning the budget hyperparameters to balance speed and accuracy requirements. For head selection, practitioners can choose full deactivation for maximum computational savings or partial deactivation if accuracy retention is paramount. However, decision-makers should note that a slight accuracy gap (0.8%) remains relative to fully unconstrained base models. Further validation on domain-specific downstream tasks, such as dense object detection or video analysis, is recommended before full-scale deployment in non-classification environments.

arXiv: 2111.15668
Cover for AdaViT: Adaptive Vision Transformers for Efficient Image Recognition

Abstract

Built on top of self-attention mechanisms, vision transformers have demonstrated remarkable performance on a variety of tasks recently. While achieving excellent performance, they still require relatively intensive computational cost that scales up drastically as the numbers of patches, self-attention heads and transformer blocks increase. In this paper, we argue that due to the large variations among images, their need for modeling long-range dependencies between patches differ. To this end, we introduce AdaViT, an adaptive computation framework that learns to derive usage policies on which patches, self-attention heads and transformer blocks to use throughout the backbone on a per-input basis, aiming to improve inference efficiency of vision transformers with a minimal drop of accuracy for image recognition. Optimized jointly with a transformer backbone in an end-to-end manner, a light-weight decision network is attached to the backbone to produce decisions on-the-fly. Extensive experiments on ImageNet demonstrate that our method obtains more than 2× improvement on efficiency compared to state-of-the-art vision transformers with only 0.8% drop of accuracy, achieving good efficiency/accuracy trade-offs conditioned on different computational budgets. We further conduct quantitative and qualitative analysis on learned usage polices and provide more insights on the redundancy in vision transformers. Code is available at https://github.com/MengLcool/AdaViT.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Approach
  • 3.1. Preliminaries
  • 3.2. Adaptive Vision Transformer
  • 3.3. Objective Function
  • 4. Experiment
  • 4.1. Experimental Setup
  • 4.2. Main Results
  • 4.3. Ablation Study
  • 4.4. Analysis
  • 5. Conclusion
  • References

Knowls

  1. Knowl 1 — Adaptive Vision Transformer Framework for Dynamic Inference

    model/method

    Adaptive Vision Transformer (AdaViT) is an adaptive computation framework that dynamically adjusts inference computation for vision transformers on a per-input basis. Rather than using a fixed computational graph for all images, AdaViT learns instance-specific usage policies across three orthogonal redundancy dimensions:

    1. Patch selection: determining which spatial patch token embeddings to retain or drop at each block.
    2. Head selection: determining which self-attention heads within multi-head self-attention (MSA) layers to activate or deactivate.
    3. Block selection: determining which transformer blocks (or their constituent MSA and feed-forward sublayers) to execute or bypass via residual connections.

    Light-weight decision sub-networks inserted throughout the transformer backbone produce binary routing decisions on-the-fly, allocating minimal computation to easy, object-centric images and full model capacity to visually complex or cluttered images.

  2. Knowl 2 — Light-Weight Decision Network for Dynamic Policy Prediction

    model/method

    To derive usage policies without introducing significant computational overhead, AdaViT inserts a decision network before each transformer block l∈[1,L]l \in [1, L]. The decision network at block ll consists of three linear layers parameterized by Wl={Wlp,Wlh,Wlb}W_l = \{W_l^p, W_l^h, W_l^b\} that take the current block input representation ZlZ_l and predict policy logit vectors:

    (mlp,mlh,mlb)=(Wlp,Wlh,Wlb)Zl(m_l^p, m_l^h, m_l^b) = (W_l^p, W_l^h, W_l^b) Z_l

    where mlp∈RNm_l^p \in \mathbb{R}^N represents keep logits for the NN spatial patch tokens, mlh∈RHm_l^h \in \mathbb{R}^H represents keep logits for the HH self-attention heads, and mlb∈R2m_l^b \in \mathbb{R}^2 represents keep logits for the two sublayers (multi-head self-attention and feed-forward network). Passing each logit through a sigmoid function yields marginal probabilities of keeping the corresponding component. Because the decision network operates directly on the output ZlZ_l of the previous l−1l-1 transformer blocks, it avoids the overhead of a separate feature extractor.

  3. Knowl 3 — Objective Function and Gumbel-Softmax Optimization for AdaViT

    equation

    To optimize discrete binary decision variables M∈{0,1}M \in \{0, 1\} in an end-to-end differentiable manner during training, AdaViT uses the Gumbel-Softmax continuous relaxation:

    Mi,k=exp⁡((log⁡mi,k+Gi,k)/τ)∑j=1Kexp⁡((log⁡mi,j+Gi,j)/τ)for k=1,…,KM_{i,k} = \frac{\exp((\log m_{i,k} + G_{i,k})/\tau)}{\sum_{j=1}^K \exp((\log m_{i,j} + G_{i,j})/\tau)} \quad \text{for } k = 1, \dots, K

    where K=2K=2 for binary decisions, mi,km_{i,k} is the unnormalized class probability, τ\tau is the temperature hyperparameter controlling distribution sharpness, and Gi,k=−log⁡(−log⁡(Ui,k))G_{i,k} = -\log(-\log(U_{i,k})) is sampled from standard Gumbel noise with Ui,k∼Uniform(0,1)U_{i,k} \sim \text{Uniform}(0, 1).

    The full model is optimized jointly by minimizing the combined loss function:

    min⁡θ,WL=Lce+Lusage\min_{\theta, W} \mathcal{L} = \mathcal{L}_{ce} + \mathcal{L}_{usage}

    where Lce=−ylog⁡(F(I;θ))\mathcal{L}_{ce} = -y \log(F(I; \theta)) is the cross-entropy classification loss on input image II with ground-truth label yy, and Lusage\mathcal{L}_{usage} penalizes deviation from predefined computational budgets γp,γh,γb∈(0,1]\gamma_p, \gamma_h, \gamma_b \in (0, 1]:

    Lusage=(1Dp∑d=1DpMdp−γp)2+(1Dh∑d=1DhMdh−γh)2+(1Db∑d=1DbMdb−γb)2\mathcal{L}_{usage} = \left(\frac{1}{D_p}\sum_{d=1}^{D_p} M_d^p - \gamma_p\right)^2 + \left(\frac{1}{D_h}\sum_{d=1}^{D_h} M_d^h - \gamma_h\right)^2 + \left(\frac{1}{D_b}\sum_{d=1}^{D_b} M_d^b - \gamma_b\right)^2

    with Dp=L×ND_p = L \times N total patch decisions, Dh=L×HD_h = L \times H total head decisions, and Db=L×2D_b = L \times 2 total sublayer decisions across all LL blocks of the backbone.

  4. Knowl 4 — Partial and Full Attention Head Deactivation Mechanisms

    model/method

    AdaViT introduces two alternative mechanisms for deactivating self-attention heads based on the binary policy variable Ml,ih∈{0,1}M_{l,i}^h \in \{0, 1\} for head ii in transformer block ll:

    1. Partial Deactivation: Bypasses the computation of the (N+1)×(N+1)(N+1) \times (N+1) attention map by replacing the softmax output with an identity matrix I∈R(N+1)×(N+1)\mathbf{I} \in \mathbb{R}^{(N+1) \times (N+1)}:

    Attn(Q,K,V)l,i={softmax(QKTdk)Vif Ml,ih=1I⋅Vif Ml,ih=0\text{Attn}(Q, K, V)_{l,i} = \begin{cases} \text{softmax}\left(\frac{QK^T}{\sqrt{d_k}}\right) V & \text{if } M_{l,i}^h = 1 \\ \mathbf{I} \cdot V & \text{if } M_{l,i}^h = 0 \end{cases}

    where Q,K,VQ, K, V are the query, key, and value matrices, and dkd_k is the per-head feature dimension. This avoids the O(N2)O(N^2) attention map computation while preserving output tensor shapes.

    1. Full Deactivation: Completely removes deactivated heads and reduces the output projection dimension accordingly:

    MSA(Zl)=Concat([headl,i∣Ml,ih=1])WlO′\text{MSA}(Z_l) = \text{Concat}\left(\left[\text{head}_{l,i} \mid M_{l,i}^h = 1\right]\right) W_l^{O'}

    where WlO′W_l^{O'} is the linear projection matrix adjusted to match the dimension of only the active heads. Full deactivation saves more floating-point operations than partial deactivation for the same drop rate, but alters embedding dimensions dynamically.

  5. Knowl 5 — Adaptive Patch Pruning and Independent Sublayer Gating

    model/method

    AdaViT controls execution flow within each transformer block ll via two gating operations:

    1. Patch Pruning: Gating mask Mlp∈{0,1}NM_l^p \in \{0, 1\}^N filters the patch sequence, removing uninformative spatial tokens while always retaining the class token zl,clsz_{l,cls}:

    Zl=[zl,cls;Ml,1pzl,1;Ml,2pzl,2;… ;Ml,Npzl,N]Z_l = \left[z_{l,cls}; M_{l,1}^p z_{l,1}; M_{l,2}^p z_{l,2}; \dots; M_{l,N}^p z_{l,N}\right]

    Tokens with Ml,jp=0M_{l,j}^p = 0 are discarded from the current block's attention and feed-forward operations.

    1. Independent Sublayer Gating: The block decision vector Mlb=(Ml,0b,Ml,1b)∈{0,1}2M_l^b = (M_{l,0}^b, M_{l,1}^b) \in \{0, 1\}^2 gates the Multi-Head Self-Attention (MSA) and Feed-Forward Network (FFN) sublayers independently across residual connections:

    Zl′=Ml,0b⋅MSA(Zl)+ZlZ'_l = M_{l,0}^b \cdot \text{MSA}(Z_l) + Z_l Zl+1=Ml,1b⋅FFN(Zl′)+Zl′Z_{l+1} = M_{l,1}^b \cdot \text{FFN}(Z'_l) + Z'_l

    This provides the flexibility to execute MSA only, FFN only, both, or skip the entire transformer block.

  6. Knowl 6 — ImageNet Classification Performance and Efficiency of AdaViT

    data/table
    Method Top-1 Acc (%) FLOPs (G) Image Size # Patch # Head # Block
    ResNet-50 79.1 4.1 224×224224\times224 - - -
    ResNet-101 79.9 7.9 224×224224\times224 - - -
    ViT-S/16 78.1 10.1 224×224224\times224 196 12 8
    DeiT-S 79.9 4.6 224×224224\times224 196 6 12
    PVT-Small 79.8 3.8 224×224224\times224 - - 15
    Swin-T 81.3 4.5 224×224224\times224 - - 12
    T2T-ViT-19 81.9 8.5 224×224224\times224 196 7 19
    CrossViT-15 81.5 5.8 224×224224\times224 196 6 15
    LocalViT-S 80.8 4.6 224×224224\times224 196 6 12
    Baseline Upperbound 81.9 8.5 224×224224\times224 196 7 19
    Baseline Random 33.0 4.0 224×224224\times224 ∼118\sim 118 ∼5.6\sim 5.6 ∼16.2\sim 16.2
    Baseline Random+ 71.5 3.9 224×224224\times224 ∼121\sim 121 ∼5.6\sim 5.6 ∼16.2\sim 16.2
    AdaViT (Ours) 81.1 3.9 224×224224\times224 ∼95\sim 95 ∼4.5\sim 4.5 ∼15.5\sim 15.5

    On ImageNet-1K, AdaViT with a T2T-ViT-19 backbone achieves 81.1% Top-1 accuracy requiring 3.9 GFLOPs per image, reducing computational cost by more than 2×2\times relative to the unpruned T2T-ViT-19 upperbound (8.5 GFLOPs, 81.9% accuracy) with only a 0.8% drop in classification accuracy. Under matched computational budgets (~3.9–4.0 GFLOPs), AdaViT substantially outperforms un-finetuned random gating (Baseline Random, 33.0% Top-1) by 48.1% and finetuned random gating (Baseline Random+, 71.5% Top-1) by 9.6%, verifying that policy decisions are instance-adaptive rather than generic model compression.

  7. Knowl 7 — Ablation of Selection Policies and Head Deactivation Variants

    empirical result

    Ablation studies on ImageNet demonstrate the individual necessity of AdaViT's learned policies and quantify the trade-offs of head deactivation schemes:

    1. Policy Ablation: Replacing individual learned usage policies with random policies at equivalent computational costs causes sharp performance drops from AdaViT's 81.1% Top-1 accuracy: random patch selection drops to 49.2% (-31.9%), random head selection drops to 57.4% (-23.7%), and random block selection drops to 64.7% (-16.4%).

    2. Partial vs. Full Head Deactivation: When 50% of self-attention heads are retained, partial head deactivation achieves 81.7% Top-1 accuracy at 6.9 GFLOPs, whereas full head deactivation achieves 80.3% Top-1 accuracy at 5.1 GFLOPs. Increasing the full deactivation retention ratio to 60% and 70% yields 80.8% Top-1 (5.8 GFLOPs) and 81.1% Top-1 (6.6 GFLOPs), respectively, showing that full deactivation attains higher FLOP reductions at the expense of a moderate accuracy penalty.

  8. Knowl 8 — Network-Depth and Semantic-Complexity Computation Allocation in AdaViT

    empirical result

    Empirical analysis of usage policies learned by AdaViT reveals distinct depth-wise and category-wise behaviors:

    1. Depth-wise Computation Distribution:

      • Tokens/Patches: The percentage of retained patch tokens decreases monotonically from early to deep transformer layers, pruning background areas and concentrating on discriminative object regions in later blocks.
      • Heads and Blocks: Attention heads and blocks are retained at higher rates in the final layers of the network compared to middle layers, as deeper representations directly drive the classification decision.
    2. Category-wise Resource Allocation: Computational cost scales dynamically with image complexity. Structurally simple, object-centric categories (such as parachute, dishrag, kite, white stork, and airship) require only ∼1.5\sim 1.5–3.03.0 GFLOPs, whereas complex scene categories with cluttered backgrounds (such as shoe shop, barbershop, restaurant, toyshop, and comic book) require ∼4.0\sim 4.0–5.05.0 GFLOPs.

  9. Knowl 9 — Experimental Configuration for AdaViT on ImageNet

    experimental setup

    AdaViT is evaluated on the ImageNet (ILSVRC2012) dataset (~1.2M training images, 50K validation images):

    • Backbone Architecture: T2T-ViT-19 with L=19L = 19 transformer blocks, H=7H = 7 attention heads per MSA layer, and N=196N = 196 tokens at input resolution 224×224224 \times 224, initialized with pretrained T2T-ViT weights.
    • Decision Network: Attached to every transformer block starting from block 2 (l≥2l \ge 2).
    • Training Details: Trained for 150 epochs using AdamW on 8 GPUs with batch size 512, initial learning rate 0.0005, weight decay 0.065, and a cosine learning rate decay schedule. The Gumbel-Softmax temperature is set to τ=5.0\tau = 5.0.
    • Evaluation Metrics: Top-1 classification accuracy (%) and computational complexity in Giga Floating-Point Operations (GFLOPs) per image.
  10. Knowl 10 — Accuracy Drop Relative to Unpruned Transformer Upperbound

    limitation

    A limitation of AdaViT is a slight degradation in Top-1 classification accuracy compared to the unpruned full-capacity transformer upperbound (81.1% at 3.9 GFLOPs vs. 81.9% at 8.5 GFLOPs on T2T-ViT-19). Dynamic on-the-fly pruning of tokens, heads, and blocks can discard fine-grained cues necessary for classifying difficult boundary cases.

Coverage note — No substantial contributed material was omitted.

References

  1. 1.Babak Ehteshami Bejnordi, Tijmen Blankevoort, and Max Welling. Batch-shaping for learning conditional channel gated networks. In ICLR, 2020. 3
  2. 2.Tolga Bolukbasi, Joseph Wang, Ofer Dekel, and Venkatesh Saligrama. Adaptive neural networks for fast test-time prediction. In ICML, 2017. 3
  3. 3.Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. End-to-end object detection with transformers. In ECCV, 2020. 1
  4. 4.Chun-Fu Chen, Quanfu Fan, and Rameswar Panda. Crossvit: Cross-attention multi-scale vision transformer for image classification. In ICCV, 2021. 1, 2, 4, 6
  5. 5.Xiangxiang Chu, Zhi Tian, Yuqing Wang, Bo Zhang, Haibing Ren, Xiaolin Wei, Huaxia Xia, and Chunhua Shen. Twins: Revisiting the design of spatial attention in vision transformers. In NeurIPS, 2021. 1
  6. 6.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In CVPR, 2009. 2, 5
  7. 7.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. In ICLR, 2021. 1, 2, 3, 4, 6
  8. 8.Alaaeldin El-Nouby, Natalia Neverova, Ivan Laptev, and Herve J ´ egou. Training vision transformers for image retrieval. arXiv preprint arXiv:2102.05644, 2021. 2
  9. 9.Haoqi Fan, Bo Xiong, Karttikeya Mangalam, Yanghao Li, Zhicheng Yan, Jitendra Malik, and Christoph Feichtenhofer. Multiscale vision transformers. arXiv preprint arXiv:2104.11227, 2021. 1, 2
  10. 10.Michael Figurnov, Maxwell D Collins, Yukun Zhu, Li Zhang, Jonathan Huang, Dmitry Vetrov, and Ruslan Salakhutdinov. Spatially adaptive computation time for residual networks. In CVPR, 2017. 3
  11. 11.Ben Graham, Alaaeldin El-Nouby, Hugo Touvron, Pierre Stock, Armand Joulin, Herve J ´ egou, and Matthijs Douze. ´ Levit: a vision transformer in convnet’s clothing for faster inference. In ICCV, 2021. 2, 3
  12. 12.Kai Han, An Xiao, Enhua Wu, Jianyuan Guo, Chunjing Xu, and Yunhe Wang. Transformer in transformer. In NeurIPS, 2021. 1
  13. 13.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In CVPR, 2016. 6
  14. 14.Shuting He, Hao Luo, Pichao Wang, Fan Wang, Hao Li, and Wei Jiang. Transreid: Transformer-based object reidentification. In ICCV, 2021. 2
  15. 15.Andrew Howard, Mark Sandler, Grace Chu, Liang-Chieh Chen, Bo Chen, Mingxing Tan, Weijun Wang, Yukun Zhu, Ruoming Pang, Vijay Vasudevan, et al. Searching for mobilenetv3. In ICCV, 2019. 2
  16. 16.Andrew G Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, Weijun Wang, Tobias Weyand, Marco Andreetto, and Hartwig Adam. Mobilenets: Efficient convolutional neural networks for mobile vision applications. arXiv preprint arXiv:1704.04861, 2017. 2
  17. 17.Gao Huang, Danlu Chen, Tianhong Li, Felix Wu, Laurens van der Maaten, and Kilian Q Weinberger. Multi-scale dense networks for resource efficient image classification. In ICLR, 2018. 3
  18. 18.Hengduo Li, Zuxuan Wu, Abhinav Shrivastava, and Larry S Davis. 2d or not 2d? adaptive 3d convolution selection for efficient video recognition. In CVPR, 2021. 3
  19. 19.Hao Li, Hong Zhang, Xiaojuan Qi, Ruigang Yang, and Gao Huang. Improved techniques for training adaptive deep networks. In ICCV, 2019. 3
  20. 20.Yawei Li, Kai Zhang, Jiezhang Cao, Radu Timofte, and Luc Van Gool. Localvit: Bringing locality to vision transformers. arXiv preprint arXiv:2104.05707, 2021. 1, 2, 6
  21. 21.Ji Lin, Yongming Rao, Jiwen Lu, and Jie Zhou. Runtime neural pruning. In NIPS, 2017. 3
  22. 22.Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. In ICCV, 2021. 1, 2, 3, 4, 6
  23. 23.Ze Liu, Jia Ning, Yue Cao, Yixuan Wei, Zheng Zhang, Stephen Lin, and Han Hu. Video swin transformer. arXiv preprint arXiv:2106.13230, 2021. 1, 2
  24. 24.Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. In ICLR, 2019. 6
  25. 25.Ningning Ma, Xiangyu Zhang, Hai-Tao Zheng, and Jian Sun. Shufflenet v2: Practical guidelines for efficient cnn architecture design. In ECCV, 2018. 2
  26. 26.Chris J Maddison, Andriy Mnih, and Yee Whye Teh. The concrete distribution: A continuous relaxation of discrete random variables. In ICLR, 2017. 2, 4, 5
  27. 27.Jiageng Mao, Yujing Xue, Minzhe Niu, Haoyue Bai, Jiashi Feng, Xiaodan Liang, Hang Xu, and Chunjing Xu. Voxel transformer for 3d object detection. In ICCV, 2021. 2
  28. 28.Yue Meng, Chung-Ching Lin, Rameswar Panda, Prasanna Sattigeri, Leonid Karlinsky, Aude Oliva, Kate Saenko, and Rogerio Feris. Ar-net: Adaptive frame resolution for efficient action recognition. In ECCV, 2020. 3
  29. 29.Mahyar Najibi, Bharat Singh, and Larry S Davis. Autofocus: Efficient multi-scale inference. In ICCV, 2019. 3
  30. 30.Bowen Pan, Rameswar Panda, Yifan Jiang, Zhangyang Wang, Rogerio Feris, and Aude Oliva. Ia-red2 : Interpretability-aware redundancy reduction for vision transformers. In NeurIPS, 2021. 3
  31. 31.Xuran Pan, Zhuofan Xia, Shiji Song, Li Erran Li, and Gao Huang. 3d object detection with pointformer. In CVPR, 2021. 2
  32. 32.Rene Ranftl, Alexey Bochkovskiy, and Vladlen Koltun. Vision transformers for dense prediction. In ICCV, 2021. 2
  33. 33.Yongming Rao, Wenliang Zhao, Benlin Liu, Jiwen Lu, Jie Zhou, and Cho-Jui Hsieh. Dynamicvit: Efficient vision transformers with dynamic token sparsification. In NeurIPS, 2021. 3
  34. 34.Mark Sandler, Andrew Howard, Menglong Zhu, Andrey Zhmoginov, and Liang-Chieh Chen. Mobilenetv2: Inverted residuals and linear bottlenecks. In CVPR, 2018. 2
  35. 35.Chen Sun, Abhinav Shrivastava, Saurabh Singh, and Abhinav Gupta. Revisiting unreasonable effectiveness of data in deep learning era. In ICCV, 2017. 2
  36. 36.Richard S Sutton and Andrew G Barto. Reinforcement learning: An introduction. MIT press Cambridge, 1998. 5
  37. 37.Mingxing Tan and Quoc Le. Efficientnet: Rethinking model scaling for convolutional neural networks. In ICML, 2019. 2
  38. 38.Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Herve J ´ egou. Training ´ data-efficient image transformers & distillation through attention. In ICML, 2021. 1, 2, 3, 4, 6
  39. 39.Burak Uzkent and Stefano Ermon. Learning when and where to zoom with deep reinforcement learning. In CVPR, 2020. 3
  40. 40.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In NeurIPS, 2017. 1, 2, 3, 4
  41. 41.Andreas Veit and Serge Belongie. Convolutional networks with adaptive inference graphs. In ECCV, 2018. 3
  42. 42.Junke Wang, Zuxuan Wu, Jingjing Chen, Xintong Han, Abhinav Shrivastava, Yu-Gang Jiang, and Ser-Nam Lim. Objectformer for image manipulation detection and localization. In CVPR, 2022. 2
  43. 43.Rui Wang, Dongdong Chen, Zuxuan Wu, Yinpeng Chen, Xiyang Dai, Mengchen Liu, Yu-Gang Jiang, Luowei Zhou, and Lu Yuan. Bevt: Bert pretraining of video transformers. In CVPR, 2022. 1
  44. 44.Wenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan, Kaitao Song, Ding Liang, Tong Lu, Ping Luo, and Ling Shao. Pyramid vision transformer: A versatile backbone for dense prediction without convolutions. In ICCV, 2021. 1, 2, 6
  45. 45.Xin Wang, Fisher Yu, Zi-Yi Dou, Trevor Darrell, and Joseph E Gonzalez. Skipnet: Learning dynamic routing in convolutional networks. In ECCV, 2018. 3
  46. 46.Xiaohan Wang, Linchao Zhu, Yu Wu, and Yi Yang. Symbiotic attention for egocentric action recognition with objectcentric alignment. IEEE TPAMI, 2020. 2
  47. 47.Yulin Wang, Rui Huang, Shiji Song, Zeyi Huang, and Gao Huang. Not all images are worth 16x16 words: Dynamic vision transformers with adaptive sequence length. In NeurIPS, 2021. 3
  48. 48.Yuqing Wang, Zhaoliang Xu, Xinlong Wang, Chunhua Shen, Baoshan Cheng, Hao Shen, and Huaxia Xia. End-to-end video instance segmentation with transformers. In CVPR, 2021. 2
  49. 49.Haiping Wu, Bin Xiao, Noel Codella, Mengchen Liu, Xiyang Dai, Lu Yuan, and Lei Zhang. Cvt: Introducing convolutions to vision transformers. arXiv preprint arXiv:2103.15808, 2021. 1
  50. 50.Zuxuan Wu, Hengduo Li, Caiming Xiong, Yu-Gang Jiang, and Larry Steven Davis. A dynamic frame selection framework for fast video recognition. IEEE TPAMI, 2022. 3
  51. 51.Zuxuan Wu, Hengduo Li, Yingbin Zheng, Caiming Xiong, Yu-Gang Jiang, and Larry S. Davis. A coarse-to-fine framework for resource efficient video recognition. IJCV, 2021. 3
  52. 52.Zuxuan Wu, Tushar Nagarajan, Abhishek Kumar, Steven Rennie, Larry S Davis, Kristen Grauman, and Rogerio Feris. Blockdrop: Dynamic inference paths in residual networks. In CVPR, 2018. 3
  53. 53.Tete Xiao, Mannat Singh, Eric Mintun, Trevor Darrell, Piotr Dollar, and Ross Girshick. Early convolutions help trans- ´ formers see better. In NeurIPS, 2021. 2
  54. 54.Enze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar, Jose M Alvarez, and Ping Luo. Segformer: Simple and efficient design for semantic segmentation with transformers. In NeurIPS, 2021. 2
  55. 55.Le Yang, Yizeng Han, Xi Chen, Shiji Song, Jifeng Dai, and Gao Huang. Resolution adaptive networks for efficient inference. In CVPR, 2020. 3
  56. 56.Li Yuan, Yunpeng Chen, Tao Wang, Weihao Yu, Yujun Shi, Zihang Jiang, Francis EH Tay, Jiashi Feng, and Shuicheng Yan. Tokens-to-token vit: Training vision transformers from scratch on imagenet. In ICCV, 2021. 1, 2, 3, 4, 5, 6
  57. 57.Xiangyu Zhang, Xinyu Zhou, Mengxiao Lin, and Jian Sun. Shufflenet: An extremely efficient convolutional neural network for mobile devices. In CVPR, 2018. 2
  58. 58.Yanyi Zhang, Xinyu Li, Chunhui Liu, Bing Shuai, Yi Zhu, Biagio Brattoli, Hao Chen, Ivan Marsic, and Joseph Tighe. Vidtr: Video transformer without convolutions. In ICCV, 2021. 1
  59. 59.Linchao Zhu and Yi Yang. Label independent memory for semi-supervised few-shot video classification. IEEE TPAMI, 2020. 2

Citation

MLA
Meng, L., et al. “AdaViT: Adaptive Vision Transformers for Efficient Image Recognition”. arXiv, 2021, http://arxiv.org/abs/2111.15668v1.
APA
Meng, L., Li, H., Chen, B.-C., Lan, S., Wu, Z., Jiang, Y.-G., & Lim, S.-N. (2021). AdaViT: Adaptive Vision Transformers for Efficient Image Recognition. arXiv. http://arxiv.org/abs/2111.15668v1
Chicago
Meng, L., H. Li, B.-C. Chen, et al. 2021. “AdaViT: Adaptive Vision Transformers for Efficient Image Recognition”. arXiv. http://arxiv.org/abs/2111.15668v1.
Harvard
Meng, L. et al. (2021) “AdaViT: Adaptive Vision Transformers for Efficient Image Recognition”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2111.15668v1.
Vancouver
1. Meng L, Li H, Chen B-C, Lan S, Wu Z, Jiang Y-G, Lim S-N (2021) AdaViT: Adaptive Vision Transformers for Efficient Image Recognition. arXiv

BibTeX

@article{meng2021adavit,
  title = {AdaViT: Adaptive Vision Transformers for Efficient Image Recognition},
  author = {Meng, Lingchen and Li, Hengduo and Chen, Bor-Chun and Lan, Shiyi and Wu, Zuxuan and Jiang, Yu-Gang and Lim, Ser-Nam},
  year = {2021},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2111.15668v1},
  eprint = {2111.15668}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE