Joint Token Pruning and Squeezing Towards More Aggressive Compression of Vision Transformers

Siyuan WeiTianzhu YeShen ZhangYao TangJiajun Liang

article2023CVPR141 citations

Proposes a joint token pruning and squeezing framework that preserves critical visual context by fusing discarded tokens into nearest-neighbor host tokens, enabling aggressive compression of vision transformers while boosting classification accuracy over existing baselines.

Listen

Vision transformers deliver exceptional visual recognition accuracy across many artificial intelligence applications, but their high computational and memory demands create substantial deployment bottlenecks. Existing compression techniques primarily discard redundant visual patches or collapse discarded elements into a single aggregate feature. However, these conventional approaches lose vital subject details and environmental context, causing sharp drops in accuracy when aggressive compression is applied.

The article aims to introduce and evaluate a compression framework that enables aggressive reductions in computational cost without incurring severe accuracy loss. Specifically, the article demonstrates how discarded visual features can be systematically matched and integrated into retained features across both standard and hybrid vision transformer architectures.

To accomplish this, the authors evaluated a two-stage method known as token pruning and squeezing. After ranking and partitioning visual tokens into reserved and pruned subsets, the system matches each discarded token to its most similar retained host token using standard similarity measures. The discarded features are then blended into these host tokens using similarity-based weights, preserving a constant data shape that supports rapid, hardware-friendly execution. The performance of this technique was evaluated through image classification benchmarks on large-scale datasets, including ImageNet1K and iNaturalist 2019, comparing multiple compression levels against existing state-of-the-art baselines.

The findings demonstrate substantial improvements across several operational benchmarks. First, when shrinking standard models to 35% of their original computational budget, the proposed method improved top-1 accuracy by 1% to 6% compared to existing baselines. Second, the approach accelerated throughput on standard hardware, enabling a mid-sized architecture to run faster than a baseline small model (1,745 versus 1,686 images per second) while achieving a 4.78% higher accuracy. Third, across longer fine-tuning periods, compressed variants outperformed the original uncompressed architectures while requiring only 65% of the original computational workload. Fourth, intentional perturbation tests showed that the method suffered roughly half the accuracy degradation of competing techniques under sub-optimal selection policies, confirming stronger robustness against token-selection errors.

These results indicate that organizations can deploy higher-accuracy vision transformer architectures on constrained computing hardware, meaningfully lowering infrastructure and energy costs without sacrificing reliability. Unlike previous methods that force a severe trade-off between model throughput and predictive accuracy, feature squeezing effectively preserves background context and fine-grained visual details. The architecture also allows plug-and-play adoption across standard model families with minimal fine-tuning overhead.

Decision-makers evaluating model compression should consider integrating feature-squeezing modules into their current transformer deployment pipelines, particularly where fixed computational constraints prevent using full-scale models. When selecting compression configurations, teams should weigh simpler attention-based scoring for fast deployment against learnable scoring heads that yield slightly higher peak accuracy after longer fine-tuning. For hybrid architectures that rely on rigid spatial convolutions, practitioners must evaluate compatibility adjustments before wide deployment.

The primary limitations include lower flexibility when integrating the framework into hybrid vision transformer layers that require strict spatial grid structures, as well as a reliance on fine-tuning pre-trained models rather than training efficiently from scratch. Nevertheless, the extensive empirical evidence provides high confidence in the method's effectiveness for general image classification and hardware acceleration.

Cover for Joint Token Pruning and Squeezing Towards More Aggressive Compression of Vision Transformers

Abstract

Although vision transformers (ViTs) have shown promising results in various computer vision tasks recently, their high computational cost limits their practical applications. Previous approaches that prune redundant tokens have demonstrated a good trade-off between performance and computation costs. Nevertheless, errors caused by pruning strategies can lead to significant information loss. Our quantitative experiments reveal that the impact of pruned tokens on performance should be noticeable. To address this issue, we propose a novel joint Token Pruning & Squeezing module (TPS) for compressing vision transformers with higher efficiency. Firstly, TPS adopts pruning to get the reserved and pruned subsets. Secondly, TPS squeezes the information of pruned tokens into partial reserved tokens via the unidirectional nearest-neighbor matching and similarity-based fusing steps. Compared to state-of-the-art methods, our approach outperforms them under all token pruning intensities. Especially while shrinking DeiT-tiny&small computational budgets to 35%, it improves the accuracy by 1%-6% compared with baselines on ImageNet classification. The proposed method can accelerate the throughput of DeiT-small beyond DeiT-tiny, while its accuracy surpasses DeiT-tiny by 4.78%. Experiments on various transformers demonstrate the effectiveness of our method, while analysis experiments prove our higher robustness to the errors of the token pruning policy. Code is available at https://github.com/megvii-research/TPS-CVPR2023.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Method
  • 3.1. Motivation
  • 3.2. Token Pruning
  • 3.3. Token Squeezing
  • 3.4. TPS on Hybrid ViTs
  • 4. Experiment
  • 4.1. Main Results
  • 4.2. Ablation Study
  • 4.3. Robustness Experiments
  • 5. Conclusions and Limitations
  • References

Knowls

  1. Knowl 1 — Token Pruning and Squeezing preserves discarded-token information

    model/method

    The paper proposes a joint Token Pruning & Squeezing (TPS) module for reducing the computational cost of vision transformers while retaining information that ordinary token pruning discards. Given an input sequence of visual tokens, TPS first partitions the tokens into a reserved subset SrS^r and a pruned subset SpS^p according to token-importance scores. Instead of deleting SpS^p, TPS assigns every pruned token to a similar reserved token, called its host, and fuses the pruned token into that host. Multiple pruned tokens may share one host, while reserved tokens that receive no pruned tokens remain unchanged.

    The resulting sequence contains exactly the reserved tokens, so the number of tokens processed by subsequent transformer layers is fixed and equal to ∣Sr∣|S^r|. This preserves constant-shape inference while recovering information from discarded subject and contextual tokens. A DynamicViT toy analysis further motivated TPS: exchanging the reserved and pruned subsets in the first pruning layer showed that pruned tokens alone could correctly classify examples missed by the original policy, and the additional accuracy from these cases increased as pruning became more aggressive.

  2. Knowl 2 — Unidirectional nearest-neighbor matching assigns pruned tokens to hosts

    equation

    Let SrS^r and SpS^p be the reserved and pruned token sets, with corresponding index sets IrI^r and IpI^p, and let each token feature be a vector in Rd\mathbb{R}^d. TPS computes a similarity ci,jc_{i,j} between every pruned token xi\boldsymbol{x}_i, i∈Ipi\in I^p, and every reserved token xj\boldsymbol{x}_j, j∈Irj\in I^r. Each pruned token independently selects one reserved host by

    xihost=arg⁡max⁡xj∈Srci,j.\boldsymbol{x}^{\mathrm{host}}_i=\mathop{\arg\max}_{\boldsymbol{x}_j\in S^r} c_{i,j}.

    The matching is unidirectional and many-to-one: every pruned token has one host, multiple pruned tokens may select the same host, and some reserved tokens may not be selected. The matching mask M∈{0,1}Np×NrM\in\{0,1\}^{N_p\times N_r}, where Np=∣Sp∣N_p=|S^p| and Nr=∣Sr∣N_r=|S^r|, is defined by

    mi,j={1,if xj is the host of xi,0,otherwise.m_{i,j}=\begin{cases} 1,&\text{if }\boldsymbol{x}_j\text{ is the host of }\boldsymbol{x}_i,\\ 0,&\text{otherwise.} \end{cases}

    In the reported experiments, TPS uses the cosine similarity of the current token features:

    ci,j=xiTxj∥xi∥ ∥xj∥,i∈Ip, j∈Ir,c_{i,j}=\frac{\boldsymbol{x}_i^{\mathsf T}\boldsymbol{x}_j}{\|\boldsymbol{x}_i\|\,\|\boldsymbol{x}_j\|},\qquad i\in I^p,\ j\in I^r,

    where ∥⋅∥\|\cdot\| is the Euclidean norm. This feature-derived similarity introduces no additional matching parameters.

  3. Knowl 3 — Similarity-weighted fusion replaces each host with a compressed feature

    equation

    For every reserved token xj∈Sr\boldsymbol{x}_j\in S^r, TPS produces an updated host feature yj\boldsymbol{y}_j by mixing the original reserved feature with the pruned features assigned to it:

    yj=wjxj+∑xi∈Spwixi.\boldsymbol{y}_j=w_j\boldsymbol{x}_j+\sum_{\boldsymbol{x}_i\in S^p}w_i\boldsymbol{x}_i.

    Here mi,j∈{0,1}m_{i,j}\in\{0,1\} is the matching mask, ci,jc_{i,j} is the cosine similarity between pruned token ii and reserved token jj, and e\mathrm e is Euler's number. For a fixed reserved token jj, the weight of a matched pruned token is

    wi=exp⁡(ci,j)mi,j∑xi∈Spexp⁡(ci,j)mi,j+e,w_i=\frac{\exp(c_{i,j})m_{i,j}}{\sum_{\boldsymbol{x}_i\in S^p}\exp(c_{i,j})m_{i,j}+\mathrm e},

    and the weight of the reserved token itself is

    wj=e∑xi∈Spexp⁡(ci,j)mi,j+e.w_j=\frac{\mathrm e}{\sum_{\boldsymbol{x}_i\in S^p}\exp(c_{i,j})m_{i,j}+\mathrm e}.

    Consequently, more similar pruned tokens receive greater influence, while the original reserved token always receives the largest baseline weight because its self-similarity is treated as 11. A reserved token with no assigned pruned token has wj=1w_j=1 and remains unchanged. The fusing step uses regular matrix operations controlled by MM and outputs only NrN_r tokens, preserving the fixed inference shape.

  4. Knowl 4 — dTPS and eTPS provide inter-block and intra-block TPS variants

    model/method

    The paper instantiates TPS in two plug-and-play forms. The inter-block variant, dTPS, follows DynamicViT: a learnable token-scoring head predicts importance scores, and a binary pruning decision is sampled during training with Straight-Through Gumbel-Softmax so that the score-based selection remains differentiable. The intra-block variant, eTPS, follows EViT and uses the attention values from the class token to score the importance of image tokens.

    During inference, both variants retain a fixed number of tokens determined by a prescribed token-reduction or keeping ratio ρ\rho using a Top-kk selection operation. The selected tokens form SrS^r, and the remaining tokens form SpS^p. Thus both variants maintain constant token dimensions and are compatible with hardware-oriented inference optimizations. The same TPS matching and similarity-based fusing mechanism is then applied to the two subsets.

  5. Knowl 5 — TPS adapts to hybrid transformers with spatially structured operations

    model/method

    TPS can be inserted into hybrid vision transformers, but operations requiring a complete spatial arrangement need special handling. In PVT, the module is placed before the first transformer block of each stage, applies token pruning, and produces a token-selection decision policy. For attention, TPS reduces the token dimension of the input and the corresponding query tensor QQ. If a spatial-reduction layer follows and requires a structured spatial tensor, features of dropped tokens are filled with zeros so that the spatial layout remains valid.

    This adaptation allows TPS to reduce computation in hybrid architectures such as PVT and CvT, while acknowledging that convolutional, pooling, and spatial-reduction operations prevent completely straightforward token removal.

  6. Knowl 6 — TPS improves the accuracy–compute trade-off on ImageNet-1K

    data/table

    The main comparison fine-tuned pretrained DeiT models on ImageNet-1K with 224×224224\times224 inputs for 30 epochs under multiple pruning locations and token-keeping ratios. Training used DeiT augmentations, AdamW, a cosine learning-rate schedule, and a base learning rate of (B/1024)×2.5×10−4(B/1024)\times2.5\times10^{-4} for batch size BB. The following reported results compare Top-1 accuracy, parameter count, and computational cost.

    Backbone Method Param (M) GFLOPs Top-1 Acc. (%)
    DeiT-S DeiT-S 22.05 4.6 79.8
    DeiT-S DynamicViT 22.77 2.9 79.3
    DeiT-S EViT 22.05 3.0 79.5
    DeiT-S eTPS 22.05 3.0 79.7
    DeiT-S dTPS∗^{*} 22.77 3.0 80.1
    DeiT-T DeiT-T 5.72 1.3 72.2
    DeiT-T DynamicViT (re-impl.) 5.90 0.8 71.4
    DeiT-T EViT (re-impl.) 5.72 0.8 71.9
    DeiT-T eTPS 5.72 0.8 72.3
    DeiT-T dTPS∗^{*} 5.90 0.8 72.9

    Here ∗^{*} denotes 100 fine-tuning epochs; other listed pruning results use 30 epochs unless otherwise specified. Across the tested pruning settings, TPS consistently outperformed DynamicViT and EViT, with the advantage increasing under more aggressive pruning. When the DeiT computational budget was reduced to 35% of the original, TPS avoided approximately 1%−6%1\%-6\% of the accuracy loss incurred by the baselines. TPS-equipped DeiT-small reached 1745 images/s on one NVIDIA RTX 2080Ti with batch size 32, exceeding DeiT-tiny's 1686 images/s, while its accuracy was 4.78 percentage points higher than DeiT-tiny.

  7. Knowl 7 — TPS generalizes across vanilla, hybrid, and fine-grained classification backbones

    data/table

    The authors integrated TPS into multiple pretrained vanilla and hybrid vision transformers. The reported hybrid-backbone results show that dTPS can reduce computation while preserving much of the original accuracy.

    Backbone Method Param (M) GFLOPs Top-1 Acc. (%)
    PVT-T PVT-T 13.23 1.94 75.1
    PVT-T dTPS∗^{*} 13.85 1.69 75.2
    PVT-S PVT-S 24.49 3.83 79.8
    PVT-S dTPS∗^{*} 25.11 3.14 79.2
    CvT-13 CvT-13 20.00 4.58 81.6
    CvT-13 dTPS∗^{*} 20.72 3.04 80.8
    CvT-21 CvT-21 31.62 7.21 82.5
    CvT-21 dTPS∗^{*} 32.35 4.10 80.9

    The method also improved or matched several vanilla-transformer pruning baselines: on LV-ViT-small, dTPS∗^{*} achieved 82.6% Top-1 accuracy at 3.8 GFLOPs versus 83.3% at 6.6 GFLOPs for the unpruned model; on LV-ViT-tiny, dTPS∗^{*} achieved 78.7% at 2.0 GFLOPs versus 79.1% at 2.9 GFLOPs; and on PS-ViT-B/14, dTPS∗^{*} achieved 81.5% at 3.7 GFLOPs versus 81.7% at 5.4 GFLOPs.

    On the fine-grained iNaturalist 2019 dataset, dTPS improved over the DynamicViT reimplementation at matched compute. For DeiT-small, 30-epoch training gave 74.2% for dTPS versus 74.0% for DynamicViT, while 100-epoch dTPS reached 74.7%. For DeiT-tiny, the corresponding results were 71.7%, 71.4%, and 72.4%. The 100-epoch DeiT-small dTPS model was only 0.1 percentage points below the unpruned DeiT-small while using 65% of its computational budget.

  8. Knowl 8 — Current full token features and cosine similarity are the best matching choices

    empirical result

    Ablation experiments on DeiT-tiny used pruning layers {4,7,8}\{4,7,8\} and a token-keeping ratio of 0.70.7. The matching representation performed best when it used the full token embedding, which includes both content and positional information. Top-1 accuracies were 71.90% with the full feature, 71.73% with the content feature obtained by subtracting the positional embedding, and 70.92% with the positional feature alone.

    The experiments also compared cosine similarity computed from current token features with reuse of the preceding attention map. For dTPS, cosine similarity gave 0.810 GFLOPs and 71.90% accuracy, while reused attention gave 0.807 GFLOPs and 71.35%. For eTPS, cosine similarity gave 0.821 GFLOPs and 72.26%, while reused attention gave 0.818 GFLOPs and 71.67%. Thus current-feature cosine matching produced a small computational increase but higher accuracy for both TPS variants.

  9. Knowl 9 — TPS is more robust than pruning and reorganization under incorrect token policies

    empirical result

    To simulate errors from a suboptimal token-scoring policy, the authors replaced the learned token-selection policy with random selection. All models used DeiT-small, identical pruning settings, and 30 fine-tuning epochs; each random-policy result was averaged over 100 trials. The original-policy and random-policy accuracies were:

    Method Policy Top-1 Acc. (%) Reported relative drop
    DynamicViT Original 79.42 –
    DynamicViT Random 76.51 -3.7
    dTPS Original 79.68 –
    dTPS Random 78.19 -1.9
    EViT Original 79.51 –
    EViT Random 77.47 -2.6
    eTPS Original 79.66 –
    eTPS Random 78.06 -2.0

    The smaller degradation of dTPS and eTPS indicates that fusing information from pruned tokens makes the compressed models less sensitive to errors in the token-selection policy than direct token pruning or single-token reorganization.

  10. Knowl 10 — Limitations concern spatial structure and pruning-aware training

    limitation

    The paper identifies two limitations. First, hybrid vision transformers contain convolutional, pooling, or spatial-reduction operations that require spatially structured inputs, so token pruning cannot always be integrated as directly as it can in plain transformer blocks. The zero-complementation and other structural adjustments used for PVT do not eliminate this integration constraint.

    Second, the reported method relies on fine-tuning pretrained models after inserting token compression. More advanced pruning-aware training-from-scratch strategies might reduce the total training time, but they were not developed in this work. The authors also identify greater adaptability to hybrid transformers and application to dense prediction tasks as future directions.

Coverage note — No substantial contributed material was deliberately omitted; illustrative qualitative prediction examples were subsumed by the quantitative robustness and accuracy results.

References

  1. 1.Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. End-to-end object detection with transformers. In European conference on computer vision, pages 213–229. Springer, 2020. 2
  2. 2.Chun-Fu Richard Chen, Quanfu Fan, and Rameswar Panda. Crossvit: Cross-attention multi-scale vision transformer for image classification. In Proceedings of the IEEE/CVF international conference on computer vision, pages 357–366, 2021. 6
  3. 3.Bowen Cheng, Alex Schwing, and Alexander Kirillov. Per-pixel classification is not all you need for semantic segmentation. Advances in Neural Information Processing Systems, 34:17864–17875, 2021. 2
  4. 4.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition, pages 248–255. Ieee, 2009. 1, 2, 5
  5. 5.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929, 2020. 1, 2
  6. 6.Alaaeldin El-Nouby, Natalia Neverova, Ivan Laptev, and Herve J ´ egou. Training vision transformers for image retrieval. arXiv preprint arXiv:2102.05644, 2021. 2
  7. 7.Haoqi Fan, Bo Xiong, Karttikeya Mangalam, Yanghao Li, Zhicheng Yan, Jitendra Malik, and Christoph Feichtenhofer. Multiscale vision transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 6824–6835, 2021. 2
  8. 8.Mohsen Fayyaz, Soroush Abbasi Koohpayegani, Farnoush Rezaei, and Sommerlade1 Hamed Pirsiavash2 Juergen Gall. Adaptive token sampling for efficient vision transformers. 1, 2, 3, 6, 7
  9. 9.Kai Han, An Xiao, Enhua Wu, Jianyuan Guo, Chunjing Xu, and Yunhe Wang. Transformer in transformer. Advances in Neural Information Processing Systems, 34:15908–15919, 2021. 6
  10. 10.Shuting He, Hao Luo, Pichao Wang, Fan Wang, Hao Li, and Wei Jiang. Transreid: Transformer-based object re-identification. In Proceedings of the IEEE/CVF international conference on computer vision, pages 15013–15022, 2021. 2
  11. 11.Geoffrey Hinton, Oriol Vinyals, Jeff Dean, et al. Distilling the knowledge in a neural network. arXiv preprint arXiv:1503.02531, 2(7), 2015. 1
  12. 12.Eric Jang, Shixiang Gu, and Ben Poole. Categorical reparameterization with gumbel-softmax. arXiv preprint arXiv:1611.01144, 2016. 4
  13. 13.Zi-Hang Jiang, Qibin Hou, Li Yuan, Daquan Zhou, Yujun Shi, Xiaojie Jin, Anran Wang, and Jiashi Feng. All tokens matter: Token labeling for training better vision transformers. Advances in Neural Information Processing Systems, 34:18590–18602, 2021. 2, 6, 7
  14. 14.Zhenglun Kong, Peiyan Dong, Xiaolong Ma, Xin Meng, Wei Niu, Mengshu Sun, Xuan Shen, Geng Yuan, Bin Ren, Hao Tang, et al. Spvit: Enabling faster vision transformers via latency-aware soft token pruning. In European Conference on Computer Vision, pages 620–640. Springer, 2022. 1, 3, 4, 7
  15. 15.Yawei Li, Kai Zhang, Jiezhang Cao, Radu Timofte, and Luc Van Gool. Localvit: Bringing locality to vision transformers. arXiv preprint arXiv:2104.05707, 2021. 2
  16. 16.Youwei Liang, Chongjian Ge, Zhan Tong, Yibing Song, Jue Wang, and Pengtao Xie. Not all patches are what you need: Expediting vision transformers via token reorganizations. arXiv preprint arXiv:2202.07800, 2022. 1, 2, 3, 4, 5, 6, 7, 8
  17. 17.Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 10012–10022, 2021. 1, 2, 6
  18. 18.Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101, 2017. 5
  19. 19.Jiageng Mao, Yujing Xue, Minzhe Niu, Haoyue Bai, Jiashi Feng, Xiaodan Liang, Hang Xu, and Chunjing Xu. Voxel transformer for 3d object detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 3164–3173, 2021. 2
  20. 20.Lingchen Meng, Hengduo Li, Bor-Chun Chen, Shiyi Lan, Zuxuan Wu, Yu-Gang Jiang, and Ser-Nam Lim. Adavit: Adaptive vision transformers for efficient image recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12309–12318, 2022. 2, 3, 6
  21. 21.Bowen Pan, Rameswar Panda, Yifan Jiang, Zhangyang Wang, Rogerio Feris, and Aude Oliva. Ia-red ˆ2: Interpretability-aware redundancy reduction for vision transformers. Advances in Neural Information Processing Systems, 34:24898–24911, 2021. 1, 7
  22. 22.Xuran Pan, Zhuofan Xia, Shiji Song, Li Erran Li, and Gao Huang. 3d object detection with pointformer. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 7463–7472, 2021. 2
  23. 23.Zizheng Pan, Bohan Zhuang, Jing Liu, Haoyu He, and Jianfei Cai. Scalable vision transformers with hierarchical pooling. In Proceedings of the IEEE/cvf international conference on computer vision, pages 377–386, 2021. 6
  24. 24.Rene Ranftl, Alexey Bochkovskiy, and Vladlen Koltun. Vision transformers for dense prediction. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 12179–12188, 2021. 2
  25. 25.Yongming Rao, Wenliang Zhao, Benlin Liu, Jiwen Lu, Jie Zhou, and Cho-Jui Hsieh. Dynamicvit: Efficient vision transformers with dynamic token sparsification. Advances in neural information processing systems, 34:13937–13949, 2021. 1, 2, 3, 4, 5, 6, 7, 8
  26. 26.Yongming Rao, Wenliang Zhao, Zheng Zhu, Jiwen Lu, and Jie Zhou. Global filter networks for image classification. Advances in Neural Information Processing Systems, 34:980–993, 2021. 2
  27. 27.Yehui Tang, Kai Han, Yunhe Wang, Chang Xu, Jianyuan Guo, Chao Xu, and Dacheng Tao. Patch slimming for efficient vision transformers. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12165–12174, 2022. 1, 3
  28. 28.Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Herve J ´ egou. Training data-efficient image transformers & distillation through attention. In International Conference on Machine Learning, pages 10347–10357. PMLR, 2021. 2, 3, 5, 6, 7, 8
  29. 29.Grant Van Horn, Oisin Mac Aodha, Yang Song, Yin Cui, Chen Sun, Alex Shepard, Hartwig Adam, Pietro Perona, and Serge Belongie. The inaturalist challenge 2019 dataset. arXiv preprint arXiv:1707.06642, 2019. 2, 5, 7, 8
  30. 30.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in neural information processing systems, 30, 2017. 2
  31. 31.Wenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan, Kaitao Song, Ding Liang, Tong Lu, Ping Luo, and Ling Shao. Pyramid vision transformer: A versatile backbone for dense prediction without convolutions. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 568–578, 2021. 1, 2, 5, 6, 7
  32. 32.Yuqing Wang, Zhaoliang Xu, Xinlong Wang, Chunhua Shen, Baoshan Cheng, Hao Shen, and Huaxia Xia. End-to-end video instance segmentation with transformers. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8741–8750, 2021. 2
  33. 33.Haiping Wu, Bin Xiao, Noel Codella, Mengchen Liu, Xiyang Dai, Lu Yuan, and Lei Zhang. Cvt: Introducing convolutions to vision transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 22–31, 2021. 1, 2, 5, 7
  34. 34.Weijian Xu, Yifan Xu, Tyler Chang, and Zhuowen Tu. Co-scale conv-attentional image transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 9981–9990, 2021. 6
  35. 35.Yifan Xu, Zhijie Zhang, Mengdan Zhang, Kekai Sheng, Ke Li, Weiming Dong, Liqing Zhang, Changsheng Xu, and Xing Sun. Evo-vit: Slow-fast token evolution for dynamic vision transformer. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 36, pages 2964–2972, 2022. 1, 2, 3, 6, 7
  36. 36.Hongxu Yin, Arash Vahdat, Jose M Alvarez, Arun Mallya, Jan Kautz, and Pavlo Molchanov. A-vit: Adaptive tokens for efficient vision transformer. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10809–10818, 2022. 1, 2, 3, 7
  37. 37.Xumin Yu, Yongming Rao, Ziyi Wang, Zuyan Liu, Jiwen Lu, and Jie Zhou. Pointr: Diverse point cloud completion with geometry-aware transformers. In Proceedings of the IEEE/CVF international conference on computer vision, pages 12498–12507, 2021. 2
  38. 38.Li Yuan, Yunpeng Chen, Tao Wang, Weihao Yu, Yujun Shi, Zi-Hang Jiang, Francis EH Tay, Jiashi Feng, and Shuicheng Yan. Tokens-to-token vit: Training vision transformers from scratch on imagenet. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 558–567, 2021. 2, 6
  39. 39.Xiaoyu Yue, Shuyang Sun, Zhanghui Kuang, Meng Wei, Philip HS Torr, Wayne Zhang, and Dahua Lin. Vision transformer with progressive sampling. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 387–396, 2021. 2, 6, 7
  40. 40.Wang Zeng, Sheng Jin, Wentao Liu, Chen Qian, Ping Luo, Wanli Ouyang, and Xiaogang Wang. Not all tokens are equal: Human-centric visual analysis via token clustering transformer. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 11101–11111, 2022. 2, 6

Citation

MLA
Wei, S., et al. “Joint Token Pruning and Squeezing Towards More Aggressive Compression of Vision Transformers”. arXiv, 2023, http://arxiv.org/abs/2304.10716v1.
APA
Wei, S., Ye, T., Zhang, S., Tang, Y., & Liang, J. (2023). Joint Token Pruning and Squeezing Towards More Aggressive Compression of Vision Transformers. arXiv. http://arxiv.org/abs/2304.10716v1
Chicago
Wei, S., T. Ye, S. Zhang, Y. Tang, and J. Liang. 2023. “Joint Token Pruning and Squeezing Towards More Aggressive Compression of Vision Transformers”. arXiv. http://arxiv.org/abs/2304.10716v1.
Harvard
Wei, S. et al. (2023) “Joint Token Pruning and Squeezing Towards More Aggressive Compression of Vision Transformers”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2304.10716v1.
Vancouver
1. Wei S, Ye T, Zhang S, Tang Y, Liang J (2023) Joint Token Pruning and Squeezing Towards More Aggressive Compression of Vision Transformers. arXiv

BibTeX

@article{wei2023joint,
  title = {Joint Token Pruning and Squeezing Towards More Aggressive Compression of Vision Transformers},
  author = {Wei, Siyuan and Ye, Tianzhu and Zhang, Shen and Tang, Yao and Liang, Jiajun},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2304.10716v1},
  eprint = {2304.10716}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE