Joint Token Pruning and Squeezing Towards More Aggressive Compression of Vision Transformers
Siyuan WeiTianzhu YeShen ZhangYao TangJiajun Liang
Proposes a joint token pruning and squeezing framework that preserves critical visual context by fusing discarded tokens into nearest-neighbor host tokens, enabling aggressive compression of vision transformers while boosting classification accuracy over existing baselines.
Vision transformers deliver exceptional visual recognition accuracy across many artificial intelligence applications, but their high computational and memory demands create substantial deployment bottlenecks. Existing compression techniques primarily discard redundant visual patches or collapse discarded elements into a single aggregate feature. However, these conventional approaches lose vital subject details and environmental context, causing sharp drops in accuracy when aggressive compression is applied.
The article aims to introduce and evaluate a compression framework that enables aggressive reductions in computational cost without incurring severe accuracy loss. Specifically, the article demonstrates how discarded visual features can be systematically matched and integrated into retained features across both standard and hybrid vision transformer architectures.
To accomplish this, the authors evaluated a two-stage method known as token pruning and squeezing. After ranking and partitioning visual tokens into reserved and pruned subsets, the system matches each discarded token to its most similar retained host token using standard similarity measures. The discarded features are then blended into these host tokens using similarity-based weights, preserving a constant data shape that supports rapid, hardware-friendly execution. The performance of this technique was evaluated through image classification benchmarks on large-scale datasets, including ImageNet1K and iNaturalist 2019, comparing multiple compression levels against existing state-of-the-art baselines.
The findings demonstrate substantial improvements across several operational benchmarks. First, when shrinking standard models to 35% of their original computational budget, the proposed method improved top-1 accuracy by 1% to 6% compared to existing baselines. Second, the approach accelerated throughput on standard hardware, enabling a mid-sized architecture to run faster than a baseline small model (1,745 versus 1,686 images per second) while achieving a 4.78% higher accuracy. Third, across longer fine-tuning periods, compressed variants outperformed the original uncompressed architectures while requiring only 65% of the original computational workload. Fourth, intentional perturbation tests showed that the method suffered roughly half the accuracy degradation of competing techniques under sub-optimal selection policies, confirming stronger robustness against token-selection errors.
These results indicate that organizations can deploy higher-accuracy vision transformer architectures on constrained computing hardware, meaningfully lowering infrastructure and energy costs without sacrificing reliability. Unlike previous methods that force a severe trade-off between model throughput and predictive accuracy, feature squeezing effectively preserves background context and fine-grained visual details. The architecture also allows plug-and-play adoption across standard model families with minimal fine-tuning overhead.
Decision-makers evaluating model compression should consider integrating feature-squeezing modules into their current transformer deployment pipelines, particularly where fixed computational constraints prevent using full-scale models. When selecting compression configurations, teams should weigh simpler attention-based scoring for fast deployment against learnable scoring heads that yield slightly higher peak accuracy after longer fine-tuning. For hybrid architectures that rely on rigid spatial convolutions, practitioners must evaluate compatibility adjustments before wide deployment.
The primary limitations include lower flexibility when integrating the framework into hybrid vision transformer layers that require strict spatial grid structures, as well as a reliance on fine-tuning pre-trained models rather than training efficiently from scratch. Nevertheless, the extensive empirical evidence provides high confidence in the method's effectiveness for general image classification and hardware acceleration.
- Paper: AdaViT: Adaptive Vision Transformers for Efficient Image Recognition, Lingchen Meng et al. (2022). AdaViT introduces dynamic patch/token pruning techniques in Vision Transformers that establish the baseline token-dropping paradigms improved upon by TPS.
- Paper: Training data-efficient image transformers & distillation through attention, Hugo Touvron et al. (2021). This paper introduces DeiT and its training recipes, which serve as the primary baseline architecture and evaluation testbed for the token pruning and squeezing module.
- Paper: Tokens-to-Token ViT: Training Vision Transformers from Scratch on ImageNet, Li Yuan et al. (2021). T2T-ViT investigates progressive token aggregation and length reduction in vision transformers, motivating structured token fusion and pruning strategies.
- Paper: An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale, Alexey Dosovitskiy et al. (2021). This seminal work establishes the patch-token architecture of Vision Transformers, the foundation upon which sequence length reduction and token compression operate.
- Paper: Zero-TPrune: Zero-Shot Token Pruning Through Leveraging of the Attention Graph in Pre-Trained Transformers, Hongjie Wang et al. (2024). Zero-TPrune extends token pruning and similarity-based aggregation methodologies into a zero-shot framework without requiring extensive retraining.
- Paper: VisionZip: Longer is Better but Not Necessary in Vision Language Models, Senqiao Yang et al. (2025). VisionZip applies dominant token selection and similarity-based visual token merging concepts to accelerate multimodal vision-language models.
- Paper: Filter, Correlate, Compress: Training-Free Token Reduction for MLLM Acceleration, Yuhang Han et al. (2026). FiCoCo generalizes token reduction by routing information from pruned visual tokens to preserved ones across multimodal large language models.
- Paper: MADTP: Multimodal Alignment-Guided Dynamic Token Pruning for Accelerating Vision-Language Transformer, Jianjian Cao et al. (2024). MADTP expands dynamic token pruning beyond unimodal vision backbones to cross-modal alignment in vision-language transformers.
- Paper: Global Compression Commander: Plug-and-Play Inference Acceleration for High-Resolution Large Vision-Language Models, Xuyang Liu et al. (2026). GlobalCom2 builds on token reduction principles to adaptively retain and compress visual tokens in dynamic high-resolution vision-language models.
