SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
Yuan ZhangChun-Kai FanJunpeng MaWenzhao ZhengTao HuangKuan ChengDenis A. GudovskiyTomoyuki OkunoYohei NakataKurt Keutzer
Proposes a training-free visual token pruning and recycling framework that uses question-relevant text attention to dramatically accelerate vision-language model inference across image and video benchmarks without sacrificing task accuracy.
Modern vision-language artificial intelligence models, which process both images and text, face major computational bottlenecks. Encoding high-resolution images or video frames generates thousands of visual data units, known as tokens. These visual sequences consume substantial memory and processing power, even though visual information is often sparse and redundant compared to concise textual prompts. While prior efficiency methods either require expensive model retraining or discard visual tokens without considering the user prompt, the article demonstrates that text instructions must guide which visual parts are kept.
The main objective of the article is to introduce and evaluate SparseVLM, a training-free framework designed to accelerate vision-language model inference. The approach dynamically identifies and prunes redundant visual data by using the context of textual prompts without requiring additional model parameters or fine-tuning.
To evaluate this framework, the authors conducted extensive empirical tests across multiple vision-language architectures, including LLaVA, Mini-Gemini, Qwen2-VL, and Video-LLaVA. Testing spanned eight standard image-understanding benchmarks and four video question-answering datasets. The SparseVLM mechanism operates directly within the model's self-attention layers by first selecting visually relevant text words to act as raters. It then calculates the importance of each visual token against these text raters, determines layer-by-layer pruning ratios based on attention matrix redundancy, and aggregates discarded tokens into compact summary representations to minimize information loss.
The findings confirm substantial performance and efficiency gains across multiple benchmarks. When applied to LLaVA, SparseVLM eliminated 66.7% to 77.8% of visual tokens, reducing processing latency by 37% to 43.1% and computational floating-point operations by up to 62.8%, while retaining 97% to 99.1% of the baseline model's accuracy. Furthermore, under heavy sparsification retaining only 64 tokens, SparseVLM outperformed prior leading acceleration methods by 17.3%. In video understanding benchmarks, SparseVLM pruned 90.5% of visual tokens and achieved an average accuracy of 95.0% relative to the uncompressed baseline, outperforming the competing FastV method by 14.7%.
These results demonstrate that vision-language models can be deployed much more affordably and quickly on edge devices and cloud infrastructure. Because SparseVLM requires no additional training, organizations can instantly integrate it into existing model pipelines to lower server operating costs, cut memory cache requirements by 67%, and improve response latency with minimal impact on accuracy.
Engineering teams and technology decision-makers should consider piloting SparseVLM as a plug-and-play optimization for latency-critical and high-volume multimodal applications. While the technique exhibits high reliability across varied benchmarks, practitioners should conduct validation on task-specific domains where extreme token reduction might risk discarding fine visual details.
- Paper: MADTP: Multimodal Alignment-Guided Dynamic Token Pruning for Accelerating Vision-Language Transformer, Jianjian Cao et al. (2024). Introduces multimodal alignment-guided dynamic token pruning for vision-language models, establishing the foundational paradigm of leveraging cross-modal attention to score and prune tokens that SparseVLM builds upon with training-free mechanisms.
- Paper: Joint Token Pruning and Squeezing Towards More Aggressive Compression of Vision Transformers, Siyuan Wei et al. (2023). Develops a token pruning and squeezing framework that merges discarded visual features into retained tokens, providing the conceptual basis for token recycling methods used in SparseVLM.
- Paper: Zero-TPrune: Zero-Shot Token Pruning Through Leveraging of the Attention Graph in Pre-Trained Transformers, Hongjie Wang et al. (2024). Presents a training-free token pruning approach leveraging attention matrices and graph centrality in transformers, preceding SparseVLM's text-guided attention scoring strategy.
- Paper: AdaViT: Adaptive Vision Transformers for Efficient Image Recognition, Lingchen Meng et al. (2022). Establishes adaptive token and layer pruning mechanisms in Vision Transformers to dynamically reduce computational overhead per input instance.
- Paper: Video-LLaVA: Learning United Visual Representation by Alignment Before Projection, Bin Lin et al. (2023). Introduces the Video-LLaVA architecture and visual token projection baseline evaluated extensively within SparseVLM's video understanding experiments.
- Paper: Global Compression Commander: Plug-and-Play Inference Acceleration for High-Resolution Large Vision-Language Models, Xuyang Liu et al. (2026). Extends training-free vision token compression to high-resolution multimodal models by dynamically coordinating global thumbnail attention with local crop token pruning.
- Paper: Filter, Correlate, Compress: Training-Free Token Reduction for MLLM Acceleration, Yuhang Han et al. (2026). Advances training-free token reduction across encoder and decoder stages in multimodal LLMs by integrating inter-token correlation routing with self-preserving compression.
- Paper: VisionZip: Longer is Better but Not Necessary in Vision Language Models, Senqiao Yang et al. (2025). Generalizes visual token sparsification in vision-language models via a text-agnostic framework that identifies dominant proxy tokens and merges redundant features.
