Zero-TPrune: Zero-Shot Token Pruning Through Leveraging of the Attention Graph in Pre-Trained Transformers

Hongjie WangBhishma DedhiaNiraj K. Jha

article2024CVPR68 citations

Proposes a training-free token pruning framework that uses attention graphs and PageRank-guided similarity grouping to accelerate vision Transformer inference with negligible accuracy loss.

Listen

Deploying modern Vision Transformers on edge devices with constrained compute, memory, and energy budgets is difficult because their computational complexity grows quadratically with the input sequence length. Token pruning is an effective way to drop unnecessary visual data and accelerate inference without altering the core model architecture. However, existing pruning methods typically require computationally expensive re-training or fine-tuning for every distinct hardware target, creating a major barrier for practical edge deployment.

The article introduces and evaluates Zero-TPrune, a zero-shot, training-free token pruning framework that exploits both token importance and token similarity in pre-trained Transformers without requiring downstream fine-tuning.

To evaluate Zero-TPrune, the authors conducted extensive experiments across standard vision architectures (including DeiT, LV-ViT, MAE, AugReg, and SWAG) using the ImageNet dataset. Zero-TPrune models the internal attention matrix as a directed graph and derives token importance using a custom Weighted PageRank algorithm. It then combines this with an importance-guided similarity stage that partitions tokens and eliminates redundant features using cosine similarity on key embedding vectors, all evaluated on standardized hardware benchmarks.

Key findings show that Zero-TPrune delivers substantial speedups while preserving accuracy without any fine-tuning. On the DeiT-S backbone, it reduces computational cost by 34.7% and increases throughput by 45.3% on an NVIDIA A100 GPU with only a 0.4% loss in top-1 accuracy. When compared with state-of-the-art methods that require hundreds of GPU hours of fine-tuning (such as DynamicViT), Zero-TPrune matches their accuracy within 0.1% while completely removing post-pruning training overhead. Compared with other training-free pruning techniques, Zero-TPrune cuts accuracy loss by up to 49% at similar computational budgets and demonstrates superior transfer learning capability across downstream tasks.

These results demonstrate that edge deployments can eliminate costly retraining cycles and dynamically switch pruning configurations at zero computational cost, significantly reducing infrastructure expenses, development timelines, and energy consumption.

Organizations seeking to deploy Vision Transformers to resource-constrained environments should adopt zero-shot pruning pipelines like Zero-TPrune, prioritizing smaller pre-trained models with moderate pruning over aggressively pruned large models to achieve optimal accuracy and efficiency. Future efforts should evaluate the framework on additional computer vision tasks, such as object detection, segmentation, and image generation.

A primary limitation is that Zero-TPrune exhibits diminishing returns when large models are pruned aggressively (e.g., reducing computational cost by 50% or more), where alternative base architectures may be preferable. Nevertheless, confidence in the reported zero-shot performance and throughput improvements across evaluated vision benchmarks remains high.

arXiv: 2305.17328
Cover for Zero-TPrune: Zero-Shot Token Pruning Through Leveraging of the Attention Graph in Pre-Trained Transformers

Abstract

Deployment of Transformer models on edge devices is becoming increasingly challenging due to the exponentially growing inference cost that scales quadratically with the number of tokens in the input sequence. Token pruning is an emerging solution to address this challenge due to its ease of deployment on various Transformer backbones. However, most token pruning methods require computationally expensive fine-tuning, which is undesirable in many edge deployment cases. In this work, we propose Zero-TPrune, the first zero-shot method that considers both the importance and similarity of tokens in performing token pruning. It leverages the attention graph of pre-trained Transformer models to produce an importance distribution for tokens via our proposed Weighted Page Rank (WPR) algorithm. This distribution further guides token partitioning for efficient similarity-based pruning. Due to the elimination of the fine-tuning overhead, Zero-TPrune can prune large models at negligible computational cost, switch between different pruning configurations at no computational cost, and perform hyperparameter tuning efficiently. We evaluate the performance of Zero-TPrune on vision tasks by applying it to various vision Transformer backbones and testing them on ImageNet. Without any fine-tuning, Zero-TPrune reduces the FLOPs cost of DeiT-S by 34.7% and improves its throughput by 45.3% with only 0.4% accuracy loss. Compared with state-of-the-art pruning methods that require fine-tuning, Zero-TPrune not only eliminates the need for fine-tuning after pruning but also does so with only 0.1% accuracy loss. Compared with state-of-the-art fine-tuning-free pruning methods, Zero-TPrune reduces accuracy loss by up to 49% with similar FLOPs budgets. Project webpage: https://jha-lab.github.io/zerotprune.

Table of Contents

  • 1. Introduction
  • 2. Related Works
  • 3. Methodology
  • 3.1. Overview: Zero-TPrune
  • 3.2. I-stage: Importance-based Pruning
  • 3.3. S-stage: Similarity-based Pruning
  • 4. Experimental Results
  • 4.1. Ablation Experiments
  • 4.2. Comparison with State-of-the-Art Methods
  • 5. Conclusion
  • References

Knowls

  1. Knowl 1 — Multi-Stage Pruning Layer Architecture in Zero-TPrune

    model/method

    Zero-TPrune inserts modular pruning layers between Transformer blocks to reduce sequence length dynamically without requiring post-pruning fine-tuning. Each pruning layer consists of three sequential sub-stages:

    1. I′I'-stage (Pre-ranking): Assigns initial importance scores to all current tokens through a single iteration of voting on the multi-head attention graph. No tokens are pruned in this stage; its sole purpose is to establish an importance ranking to guide subsequent token partitioning.
    2. SS-stage (Similarity-based Pruning): Partitions tokens into two groups based on the importance ranking from the I′I'-stage, computes pair-wise similarity between the groups, and discards the less important token in each of the top-rr most similar pairs to eliminate feature redundancy.
    3. II-stage (Importance-based Pruning): Applies full Weighted Page Rank (WPR), Emphasizing Informative Region (EIR) aggregation, and Variance-based Head Filtering (VHF) on the remaining tokens' attention graph to compute converged importance scores and retain only the top-kk tokens.

    Executing similarity-based pruning (SS-stage) prior to full importance pruning (II-stage) avoids a failure mode where semantically unimportant tokens with high self-similarity or shared background attention collectively accumulate high graph centrality and crowd out salient foreground object tokens.

  2. Knowl 2 — Weighted Page Rank for Token Importance Assignment

    algorithm

    The Weighted Page Rank (WPR) algorithm computes an importance score distribution over image tokens by interpreting the self-attention matrix as a weighted, directed adjacency matrix of a complete attention graph. Each directed edge (xj,xi)(x_j, x_i) indicates information routing from token xjx_j to token xix_i, weighted by the attention value A(xi,xj)A(x_i, x_j).

    Input: Number of tokens N>0N > 0, attention matrix A∈RN×NA \in \mathbb{R}^{N \times N}, initial graph signal s0∈RNs^0 \in \mathbb{R}^N, convergence tolerance ϵ>0\epsilon > 0 or maximum iterations TT
    Output: Importance score vector s∈RNs \in \mathbb{R}^N
    t←0t \leftarrow 0
    while (t<Tt < T and ∣st−st−1∣>ϵ|s^t - s^{t-1}| > \epsilon) or (t=0t = 0) do
        t←t+1t \leftarrow t + 1
        st←ATst−1s^t \leftarrow A^T s^{t-1}
    end while
    s←sts \leftarrow s^t
    return ss

    For head hh in layer ll, the recursive stationary score for token xix_i satisfies:

    s(h,l)(xi)=1N∑j=1NA(h,l)(xi,xj)⋅s(h,l)(xj)s^{(h,l)}(x_i) = \frac{1}{N} \sum_{j=1}^N A^{(h,l)}(x_i, x_j) \cdot s^{(h,l)}(x_j)

    To ensure convergence without per-step convergence testing overhead during inference, fixed iteration counts TT are assigned by layer depth: 30–5030\text{--}50 iterations in the initial three layers, 5–105\text{--}10 iterations in intermediate layers, and 11 iteration in the final three layers. The computational complexity of the WPR procedure is O(N2)O(N^2) per head.

  3. Knowl 3 — Multi-Head Importance Score Aggregation with EIR and VHF

    equation

    To aggregate token importance scores across attention heads while suppressing uninformative heads and emphasizing localized salient features, Zero-TPrune combines Emphasizing Informative Region (EIR) root-mean-square aggregation with a Variance-based Head Filter (VHF):

    s(l)(xi)=∑h=1Nh(s(h,l)(xi))2⋅η(vmin⁡≤Varh≤vmax⁡)∑h=1Nhη(vmin⁡≤Varh≤vmax⁡)s^{(l)}(x_i) = \sqrt{\frac{\sum_{h=1}^{N_h} \left(s^{(h,l)}(x_i)\right)^2 \cdot \eta(v_{\min} \le \mathrm{Var}_h \le v_{\max})}{\sum_{h=1}^{N_h} \eta(v_{\min} \le \mathrm{Var}_h \le v_{\max})}}

    where:

    • s(h,l)(xi)∈Rs^{(h,l)}(x_i) \in \mathbb{R} is the importance score of token xix_i in head hh of layer ll, derived via Weighted Page Rank;
    • NhN_h is the total number of attention heads in layer ll;
    • Varh=1N∑i=1N(s(h,l)(xi)−sˉ(h,l))2\mathrm{Var}_h = \frac{1}{N} \sum_{i=1}^N \left(s^{(h,l)}(x_i) - \bar{s}^{(h,l)}\right)^2 is the variance of the token importance score distribution in head hh;
    • vmin⁡v_{\min} and vmax⁡v_{\max} are the lower and upper variance thresholds (empirically set to vmin⁡=0.01v_{\min} = 0.01 and vmax⁡=0.7v_{\max} = 0.7);
    • η(⋅)\eta(\cdot) is an indicator function returning 11 if the head's variance falls within [vmin⁡,vmax⁡][v_{\min}, v_{\max}] and 00 otherwise, filtering out degenerate heads whose score distributions are either nearly uniform (low variance) or concentrated entirely on edge/boundary tokens (excessive variance);
    • EIR uses the quadratic sum to favor tokens that exhibit high importance in a subset of heads over tokens that have mediocre importance across all heads.

    The overall computational complexity of the II-stage including WPR, EIR, and VHF is O(N2)O(N^2), where NN is the number of tokens.

  4. Knowl 4 — Importance-Guided Similarity-Based Token Pruning in the S-Stage

    model/method

    The SS-stage removes redundant tokens that encode similar representations by combining importance rank guidance with feature similarity:

    1. Importance Partitioning: Tokens are ranked by their preliminary importance scores from the I′I'-stage and split sequentially into two roughly equal-sized bipartite partitions, Group AA (lower importance) and Group BB (higher importance).
    2. Feature Representation: Each token is represented by its projection vector from the self-attention Key matrix K∈RN×dK \in \mathbb{R}^{N \times d}, where dd is the embedding dimension.
    3. Bipartite Similarity Matching: For each token in Group AA, its cosine similarity against all tokens in Group BB is computed. Tokens within the same group are never paired, eliminating redundant intra-group comparisons.
    4. Pruning without Merging: The top-rr pairs with the highest cosine similarities across the bipartite matching are identified. For each selected pair, the token belonging to Group AA is pruned (discarded), while the token in Group BB is preserved unaltered. Discarding rather than merging retains uniform token weighting and preserves structural compatibility with sparse Transformer backbones.

    The computational complexity of the SS-stage is O(N2⋅d)O(N^2 \cdot d).

  5. Knowl 5 — Prior Initialization for Classification Tokens in Zero-TPrune

    model/method

    The initialization of the graph signal s0∈RNs^0 \in \mathbb{R}^N for the Weighted Page Rank algorithm depends on whether the model is configured for task-agnostic universal feature processing or classification:

    • Universal Mode (Zero-TPrune-uni): The graph signal is initialized uniformly across all NN tokens: s0=1N1Ns^0 = \frac{1}{N} \mathbf{1}_N
    • Classification Mode (Zero-TPrune): Because the classification ([CLS][\mathrm{CLS}]) token acts as the primary sink for global class representations, its initial importance score is assigned a prior scaled by N\sqrt{N} relative to image patch tokens: s0(CLS)=NN+N−1,s0(xi)=1N+N−1∀xi≠CLSs^0(\mathrm{CLS}) = \frac{\sqrt{N}}{\sqrt{N} + N - 1}, \quad s^0(x_i) = \frac{1}{\sqrt{N} + N - 1} \quad \forall x_i \ne \mathrm{CLS}

    This inductive bias amplifies the voting weight of the [CLS][\mathrm{CLS}] token's incoming and outgoing attention links during iterative importance propagation.

  6. Knowl 6 — Ablation Breakdown of Zero-TPrune Components on DeiT-S

    data/table

    Ablation of individual components in Zero-TPrune evaluated on ImageNet-1k using the DeiT-S backbone at image resolution 224×224224 \times 224 and batch size 512.

    Method Acc@1 (%) Params FLOPs/img Throughput (img/s)
    Unpruned DeiT-S 79.8 22M 4.55G 1505.9
    Random drop 76.8 22M 3.08G 2164.4
    WPR 78.6 22M 3.08G 2136.5
    WPR + EIR 78.8 22M 3.08G 2132.6
    WPR + EIR + VHF (II-stage) 78.9 22M 3.08G 2103.1
    II-stage + SS-stage (Zero-TPrune) 79.4 22M 3.08G 2063.9

    At a matched inference budget of 3.08 GFLOPs3.08\text{ GFLOPs} (~32.3%32.3\% FLOPs reduction):

    • Pure Weighted Page Rank (WPR) improves Top-1 accuracy over naive random token dropping by +1.8%+1.8\%.
    • Adding Emphasizing Informative Region (EIR) aggregation yields +0.2%+0.2\%.
    • Adding the Variance-based Head Filter (VHF) provides an additional +0.1%+0.1\%, bringing the isolated II-stage accuracy to 78.9%78.9\%.
    • Combining the complete II-stage with the similarity-based SS-stage increases accuracy to 79.4%79.4\%, cutting the accuracy drop from unpruned baseline (79.8%79.8\%) down to only 0.4%0.4\% while achieving a 37.0%37.0\% throughput speedup (2063.9 img/s2063.9\text{ img/s} vs. 1505.9 img/s1505.9\text{ img/s}).
  7. Knowl 7 — Off-the-Shelf Performance Comparison on DeiT-S Without Fine-Tuning

    data/table

    Evaluation of fine-tuning-free token reduction methods applied off-the-shelf to the DeiT-S backbone on ImageNet-1k validation (224×224224 \times 224 images, throughput measured on a single NVIDIA A100 GPU).

    Method Acc@1 (%) GFLOPs Throughput (img/s)
    DeiT-S (Baseline) 79.8 4.55 1505.9
    + ATS 79.2 (-0.6%) 3.00 (-33.4%) 2062.3 (+36.9%)
    + ToMe 78.9 (-0.9%) 2.95 (-35.2%) 2263.9 (+50.3%)
    + Zero-TP-a 79.4 (-0.4%) 2.97 (-34.7%) 2188.4 (+45.3%)
    + Zero-TP-b 79.1 (-0.7%) 2.50 (-45.1%) 2458.4 (+63.2%)
    + Zero-TP-c 79.8 (-0.0%) 3.97 (-12.7%) 1673.2 (+11.1%)

    At comparable compute reductions (~3.0 GFLOPs3.0\text{ GFLOPs}, 34–35%34\text{--}35\% reduction):

    • Zero-TPrune (Zero-TP-a) achieves 79.4%79.4\% Top-1 accuracy, outperforming Adaptive Token Sampling (ATS, 79.2%79.2\%) and Token Merging (ToMe, 78.9%78.9\%), reducing the accuracy loss of ToMe by 55.6%55.6\%.
    • At an aggressive 45.1%45.1\% FLOP reduction (Zero-TP-b, 2.50 GFLOPs2.50\text{ GFLOPs}), Zero-TPrune retains 79.1%79.1\% Top-1 accuracy (+63.2%+63.2\% throughput).
    • At a mild 12.7%12.7\% FLOP reduction (Zero-TP-c, 3.97 GFLOPs3.97\text{ GFLOPs}), Zero-TPrune preserves full baseline accuracy (79.8%79.8\%) with an 11.1%11.1\% throughput gain.
  8. Knowl 8 — Zero-Shot Pruning Across Diverse Vision Transformer Backbones

    data/table

    Top-1 validation accuracy on ImageNet-1k across different Vision Transformer backbones without post-pruning fine-tuning, comparing Zero-TPrune (Zero-TP) against ATS and ToMe.

    Method Acc@top1 (%) GFLOPs Method Acc@top1 (%) GFLOPs
    AugReg 81.41 4.55 MAE 83.62 55.4
    + ATS 79.21 2.80 + ATS 82.07 42.3
    + ToMe 79.30 2.78 + ToMe 82.69 42.2
    + Zero-TP 80.22 2.79 + Zero-TP 82.93 42.3
    LV-ViT-S 83.3 6.6 SWAG (384px) 85.30 55.6
    + ATS 80.4 3.5 + ATS 84.21 43.8
    + ToMe 79.8 3.6 + ToMe 85.09 43.8
    + Zero-TP 81.5 3.5 + Zero-TP 85.17 43.8

    On medium-sized backbones (AugReg and LV-ViT-S), Zero-TPrune reduces accuracy loss relative to fine-tuning-free baselines by up to 49%49\%. For example, on LV-ViT-S at 3.5 GFLOPs3.5\text{ GFLOPs}, Zero-TPrune achieves 81.5%81.5\% accuracy compared to 80.4%80.4\% (ATS) and 79.8%79.8\% (ToMe). On large-scale backbones (MAE and SWAG) under moderate pruning (~20–24%20\text{--}24\% FLOP reduction), Zero-TPrune maintains the highest accuracy among training-free methods (82.93%82.93\% on MAE, 85.17%85.17\% on SWAG).

  9. Knowl 9 — Comparative Trade-offs Between Zero-Shot and Fine-Tuned Token Pruning

    empirical result

    State-of-the-art token pruning methods such as DynamicViT and A-ViT insert parameterized prediction/halting modules that require computationally expensive joint retraining or fine-tuning (e.g., DynamicViT requires approximately 150 hours of fine-tuning on an NVIDIA A100 GPU for DeiT-S, and thousands of GPU hours for DeiT-B/L). When deployed without fine-tuning, DynamicViT and A-ViT degrade to random token dropping, losing approximately 3.0%3.0\% Top-1 accuracy at 3.08 GFLOPs3.08\text{ GFLOPs} on DeiT-S.

    In contrast, Zero-TPrune operates completely training-free:

    • It reduces the accuracy drop of untrained DynamicViT/A-ViT by more than 60%60\%.
    • Without any fine-tuning, Zero-TPrune matches the accuracy of fully fine-tuned DynamicViT and A-ViT within 0.1%0.1\% at 3.5 GFLOPs3.5\text{ GFLOPs} on DeiT-S.
    • When fine-tuning after pruning is optionally applied, Zero-TPrune outperforms both fine-tuned DynamicViT and fine-tuned A-ViT.
  10. Knowl 10 — Scaling Limits of Aggressive Zero-Shot Pruning on Large Transformers

    limitation

    When applying Zero-TPrune without fine-tuning to large Vision Transformer architectures (e.g., LV-ViT-M, DeiT-B) under aggressive compression budgets (reducing GFLOPs by ≥50%\ge 50\%), token-merging methods like ToMe retain slightly higher accuracy than pure pruning. Pruning removes half of the spatial feature representations entirely, causing non-recoverable information loss in deep representations unless tokens are merged or fine-tuned.

    However, aggressively pruning a large model is computationally suboptimal compared to moderately pruning a smaller base model: an unpruned LV-ViT-S pruned with Zero-TPrune achieves 81.5%81.5\% ImageNet Top-1 accuracy at 3.5 GFLOPs3.5\text{ GFLOPs}, whereas an LV-ViT-M aggressively compressed with ToMe achieves 81.6%81.6\% accuracy but requires 6.3 GFLOPs6.3\text{ GFLOPs} (1.8×1.8\times the computational cost for a 0.1%0.1\% accuracy difference).

Coverage note — None was omitted. All core methodological stages (WPR, EIR, VHF, S-stage, CLS prior), quantitative tables (Table 1, 2, 3), and empirical analyses were captured.

References

  1. 1.Andrea Banino, Jan Balaguer, and Charles Blundell. PonderNet: Learning to Ponder. arXiv preprint arXiv:2107.05407, 2021.
  2. 2.Daniel Bolya and Judy Hoffman. Token Merging for Fast Stable Diffusion. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4598–4602, 2023.
  3. 3.Daniel Bolya, Cheng-Yang Fu, Xiaoliang Dai, Peizhao Zhang, Christoph Feichtenhofer, and Judy Hoffman. Token Merging: Your ViT But Faster. arXiv preprint arXiv:2210.09461, 2022.
  4. 4.Sergey Brin. The PageRank Citation Ranking: Bringing Order to the Web. Proceedings of ASIS, 1998, 98:161–172, 1998.
  5. 5.Yun-Hao Cao, Hao Yu, and Jianxin Wu. Training Vision Transformers with Only 2040 Images. In Proceedings of the European Conference on Computer Vision, pages 220–237. Springer, 2022.
  6. 6.Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever. Generating Long Sequences with Sparse Transformers. arXiv preprint arXiv:1904.10509, 2019.
  7. 7.Mircea Cimpoi, Subhransu Maji, Iasonas Kokkinos, Sammy Mohamed, and Andrea Vedaldi. Describing Textures in the Wild. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 3606–3613, 2014.
  8. 8.Marco Cuturi, Olivier Teboul, and Jean-Philippe Vert. Differentiable Ranking and Sorting Using Optimal Transport. Advances in Neural Information Processing Systems, 32, 2019.
  9. 9.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. ImageNet: A Large-Scale Hierarchical Image Database. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 248–255, 2009.
  10. 10.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT: Pre-Training of Deep Bidirectional Transformers for Language Understanding. arXiv preprint arXiv:1810.04805, 2018.
  11. 11.Peiyan Dong, Mengshu Sun, Alec Lu, Yanyue Xie, Kenneth Liu, Zhenglun Kong, Xin Meng, Zhengang Li, Xue Lin, Zhenman Fang, et al. HeatViT: Hardware-Efficient Adaptive Token Pruning for Vision Transformers. In Proceedings of the IEEE International Symposium on High-Performance Computer Architecture, pages 442–455, 2023.
  12. 12.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. arXiv preprint arXiv:2010.11929, 2020.
  13. 13.Maha Elbayad, Jiatao Gu, Edouard Grave, and Michael Auli. Depth-Adaptive Transformer. arXiv preprint, 2019.
  14. 14.Mohsen Fayyaz, Soroush Abbasi Koohpayegani, Farnoush Rezaei Jafari, Sunando Sengupta, Hamid Reza Vaezi Joze, Eric Sommerlade, Hamed Pirsiavash, and Jürgen Gall. Adaptive Token Sampling for Efficient Vision Transformers. In Proceedings of the European Conference on Computer Vision, pages 396–414. Springer, 2022.
  15. 15.Saurabh Goyal, Anamitra Roy Choudhury, Saurabh Raje, Venkatesan Chakaravarthy, Yogish Sabharwal, and Ashish Verma. PoWER-BERT: Accelerating BERT Inference via Progressive Word-Vector Elimination. In Proceedings of the International Conference on Machine Learning, pages 3690–3699. PMLR, 2020.
  16. 16.Alex Graves. Adaptive Computation Time for Recurrent Neural Networks. arXiv preprint arXiv:1603.08983, 2016.
  17. 17.Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross Girshick. Masked Autoencoders are Scalable Vision Learners. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 16000–16009, 2022.
  18. 18.Zi-Hang Jiang, Qibin Hou, Li Yuan, Daquan Zhou, Yujun Shi, Xiaojie Jin, Anran Wang, and Jiashi Feng. All Tokens Matter: Token Labeling for Training Better Vision Transformers. Advances in Neural Information Processing Systems, 34:18590–18602, 2021.
  19. 19.Gyuwan Kim and Kyunghyun Cho. Length-Adaptive Transformer: Train Once with Length Drop, Use Anytime with Search. arXiv preprint arXiv:2010.07003, 2020.
  20. 20.Sehoon Kim, Sheng Shen, David Thorsley, Amir Gholami, Woosuk Kwon, Joseph Hassoun, and Kurt Keutzer. Learned Token Pruning for Transformers. In Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, pages 784–794, 2022.
  21. 21.Zhenglun Kong, Peiyan Dong, Xiaolong Ma, Xin Meng, Wei Niu, Mengshu Sun, Xuan Shen, Geng Yuan, Bin Ren, Hao Tang, et al. SPViT: Enabling Faster Vision Transformers Latency-Aware Soft Token Pruning. In Proceedings of the European Conference on Computer Vision, pages 620–640. Springer, 2022.
  22. 22.Jonathan Krause, Michael Stark, Jia Deng, and Li Fei-Fei. 3D Object Representations for Fine-Grained Categorization. In Proceedings of the IEEE International Conference on Computer Vision Workshops, pages 554–561, 2013.
  23. 23.Weijie Liu, Peng Zhou, Zhe Zhao, Zhiruo Wang, Haotang Deng, and Qi Ju. FastBERT: a Self-Distilling BERT with Adaptive Inference Time. arXiv preprint arXiv:2004.02178, 2020.
  24. 24.Xiangcheng Liu, Tianyi Wu, and Guodong Guo. Adaptive Sparse ViT: Towards Learnable Adaptive Token Pruning by Fully Exploiting Self-Attention. arXiv preprint arXiv:2209.13802, 2022.
  25. 25.Subhransu Maji, Esa Rahtu, Juho Kannala, Matthew Blaschko, and Andrea Vedaldi. Fine-Grained Visual Classification of Aircraft. arXiv preprint arXiv:1306.5151, 2013.
  26. 26.Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A. Rusu, Joel Veness, Marc G. Bellemare, Alex Graves, Martin Riedmiller, Andreas K. Fidjeland, Georg Ostrovski, et al. Human-Level Control through Deep Reinforcement Learning. Nature, 518(7540):529–533, 2015.
  27. 27.Maria-Elena Nilsback and Andrew Zisserman. A Visual Vocabulary for Flower Classification. In Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition, pages 1447–1454, 2006.
  28. 28.Omkar M. Parkhi, Andrea Vedaldi, Andrew Zisserman, and C.V. Jawahar. Cats and Dogs. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 3498–3505, 2012.
  29. 29.Ariadna Quattoni and Antonio Torralba. Recognizing Indoor Scenes. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 413–420, 2009.
  30. 30.Yongming Rao, Wenliang Zhao, Benlin Liu, Jiwen Lu, Jie Zhou, and Cho-Jui Hsieh. DynamicViT: Efficient Vision Transformers with Dynamic Token Sparsification. Advances in Neural Information Processing Systems, 34:13937–13949, 2021.
  31. 31.Stefan Schaal and Christopher G. Atkeson. Learning Control in Robotics. IEEE Robotics & Automation Magazine, 17(2): 20–29, 2010.
  32. 32.Mannat Singh, Laura Gustafson, Aaron Adcock, Vinicius de Freitas Reis, Bugra Gedik, Raj Prateek Kosaraju, Dhruv Mahajan, Ross Girshick, Piotr Dollár, and Laurens Van Der Maaten. Revisiting Weakly Supervised Pre-Training of Visual Perception Models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 804–814, 2022.
  33. 33.Andreas Steiner, Alexander Kolesnikov, Xiaohua Zhai, Ross Wightman, Jakob Uszkoreit, and Lucas Beyer. How to Train Your ViT? Data, Augmentation, and Regularization in Vision Transformers. arXiv preprint arXiv:2106.10270, 2021.
  34. 34.Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian Goodfellow, and Rob Fergus. Intriguing Properties of Neural Networks. arXiv preprint arXiv:1312.6199, 2013.
  35. 35.Quan Tang, Bowen Zhang, Jiajun Liu, Fagui Liu, and Yifan Liu. Dynamic Token Pruning in Plain Vision Transformers for Semantic Segmentation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 777–786, 2023.
  36. 36.Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Hervé Jégou. Training Data-Efficient Image Transformers & Distillation through Attention. In Proceedings of the International Conference on Machine Learning, pages 10347–10357. PMLR, 2021.
  37. 37.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is All You Need. Advances in Neural Information Processing Systems, 30, 2017.
  38. 38.Catherine Wah, Steve Branson, Peter Welinder, Pietro Perona, and Serge Belongie. The Caltech-UCSD Birds-200-2011 Dataset. California Institute of Technology, https://www.vision.caltech.edu/datasets/cub_200_2011/, 2011.
  39. 39.Hanrui Wang, Zhekai Zhang, and Song Han. SpAtten: Efficient Sparse Attention Architecture with Cascade Token and Head Pruning. In Proceedings of the IEEE International Symposium on High-Performance Computer Architecture, pages 97–110, 2021.
  40. 40.Siyuan Wei, Tianzhu Ye, Shen Zhang, Yao Tang, and Jiajun Liang. Joint Token Pruning and Squeezing Towards More Aggressive Compression of Vision Transformers. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2092–2101, 2023.
  41. 41.Xiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, and Lucas Beyer. Scaling Vision Transformers. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12104–12113, 2022.
  42. 42.Deming Ye, Yankai Lin, Yufei Huang, and Maosong Sun. Tr-BERT: Dynamic Token Reduction for Accelerating BERT Inference. arXiv preprint arXiv:2105.11618, 2021.
  43. 43.Hongxu Yin, Arash Vahdat, Jose M. Alvarez, Arun Mallya, Jan Kautz, and Pavlo Molchanov. A-ViT: Adaptive Tokens for Efficient Vision Transformer. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10809–10818, 2022.

Citation

MLA
Wang, H., et al. “Zero-TPrune: Zero-Shot Token Pruning Through Leveraging of the Attention Graph in Pre-Trained Transformers”. arXiv, 2023, http://arxiv.org/abs/2305.17328v3.
APA
Wang, H., Dedhia, B., & Jha, N. K. (2023). Zero-TPrune: Zero-Shot Token Pruning through Leveraging of the Attention Graph in Pre-Trained Transformers. arXiv. http://arxiv.org/abs/2305.17328v3
Chicago
Wang, H., B. Dedhia, and N. K. Jha. 2023. “Zero-TPrune: Zero-Shot Token Pruning Through Leveraging of the Attention Graph in Pre-Trained Transformers”. arXiv. http://arxiv.org/abs/2305.17328v3.
Harvard
Wang, H., Dedhia, B. and Jha, N.K. (2023) “Zero-TPrune: Zero-Shot Token Pruning through Leveraging of the Attention Graph in Pre-Trained Transformers”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2305.17328v3.
Vancouver
1. Wang H, Dedhia B, Jha NK (2023) Zero-TPrune: Zero-Shot Token Pruning through Leveraging of the Attention Graph in Pre-Trained Transformers. arXiv

BibTeX

@article{wang2023zero,
  title = {Zero-TPrune: Zero-Shot Token Pruning through Leveraging of the Attention Graph in Pre-Trained Transformers},
  author = {Wang, Hongjie and Dedhia, Bhishma and Jha, Niraj K.},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2305.17328v3},
  eprint = {2305.17328}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE