Zero-TPrune: Zero-Shot Token Pruning Through Leveraging of the Attention Graph in Pre-Trained Transformers
Hongjie WangBhishma DedhiaNiraj K. Jha
Proposes a training-free token pruning framework that uses attention graphs and PageRank-guided similarity grouping to accelerate vision Transformer inference with negligible accuracy loss.
Deploying modern Vision Transformers on edge devices with constrained compute, memory, and energy budgets is difficult because their computational complexity grows quadratically with the input sequence length. Token pruning is an effective way to drop unnecessary visual data and accelerate inference without altering the core model architecture. However, existing pruning methods typically require computationally expensive re-training or fine-tuning for every distinct hardware target, creating a major barrier for practical edge deployment.
The article introduces and evaluates Zero-TPrune, a zero-shot, training-free token pruning framework that exploits both token importance and token similarity in pre-trained Transformers without requiring downstream fine-tuning.
To evaluate Zero-TPrune, the authors conducted extensive experiments across standard vision architectures (including DeiT, LV-ViT, MAE, AugReg, and SWAG) using the ImageNet dataset. Zero-TPrune models the internal attention matrix as a directed graph and derives token importance using a custom Weighted PageRank algorithm. It then combines this with an importance-guided similarity stage that partitions tokens and eliminates redundant features using cosine similarity on key embedding vectors, all evaluated on standardized hardware benchmarks.
Key findings show that Zero-TPrune delivers substantial speedups while preserving accuracy without any fine-tuning. On the DeiT-S backbone, it reduces computational cost by 34.7% and increases throughput by 45.3% on an NVIDIA A100 GPU with only a 0.4% loss in top-1 accuracy. When compared with state-of-the-art methods that require hundreds of GPU hours of fine-tuning (such as DynamicViT), Zero-TPrune matches their accuracy within 0.1% while completely removing post-pruning training overhead. Compared with other training-free pruning techniques, Zero-TPrune cuts accuracy loss by up to 49% at similar computational budgets and demonstrates superior transfer learning capability across downstream tasks.
These results demonstrate that edge deployments can eliminate costly retraining cycles and dynamically switch pruning configurations at zero computational cost, significantly reducing infrastructure expenses, development timelines, and energy consumption.
Organizations seeking to deploy Vision Transformers to resource-constrained environments should adopt zero-shot pruning pipelines like Zero-TPrune, prioritizing smaller pre-trained models with moderate pruning over aggressively pruned large models to achieve optimal accuracy and efficiency. Future efforts should evaluate the framework on additional computer vision tasks, such as object detection, segmentation, and image generation.
A primary limitation is that Zero-TPrune exhibits diminishing returns when large models are pruned aggressively (e.g., reducing computational cost by 50% or more), where alternative base architectures may be preferable. Nevertheless, confidence in the reported zero-shot performance and throughput improvements across evaluated vision benchmarks remains high.
- Paper: An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale, Alexey Dosovitskiy et al. (2021). This seminal work establishes the foundational Vision Transformer (ViT) architecture whose quadratic attention complexity and visual token mechanics Zero-TPrune directly aims to compress.
- Paper: Training data-efficient image transformers & distillation through attention, Hugo Touvron et al. (2021). This paper introduces DeiT and distillation tokens, serving as one of the primary Vision Transformer baseline architectures used to evaluate Zero-TPrune's zero-shot token pruning performance.
- Paper: Joint Token Pruning and Squeezing Towards More Aggressive Compression of Vision Transformers, Siyuan Wei et al. (2023). This work establishes the dual concepts of token importance ranking and similarity-based token merging for Vision Transformer compression that Zero-TPrune builds upon in a training-free framework.
- Paper: AdaViT: Adaptive Vision Transformers for Efficient Image Recognition, Lingchen Meng et al. (2022). This paper presents adaptive per-image token pruning for vision transformers, motivating the token redundancy problem that Zero-TPrune solves without requiring learned gating sub-networks or retraining.
- Paper: VisionZip: Longer is Better but Not Necessary in Vision Language Models, Senqiao Yang et al. (2025). This work extends training-free visual token selection and semantic merging strategies from pure vision transformers to the vision encoders of large multimodal vision-language models.
- Paper: Global Compression Commander: Plug-and-Play Inference Acceleration for High-Resolution Large Vision-Language Models, Xuyang Liu et al. (2026). This paper generalizes training-free visual token compression to dynamic high-resolution vision-language models by combining global contextual scoring with local patch pruning.
