SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Yuan ZhangChun-Kai FanJunpeng MaWenzhao ZhengTao HuangKuan ChengDenis A. GudovskiyTomoyuki OkunoYohei NakataKurt Keutzer

article2025ICML433 citations

Proposes a training-free visual token pruning and recycling framework that uses question-relevant text attention to dramatically accelerate vision-language model inference across image and video benchmarks without sacrificing task accuracy.

Listen

Modern vision-language artificial intelligence models, which process both images and text, face major computational bottlenecks. Encoding high-resolution images or video frames generates thousands of visual data units, known as tokens. These visual sequences consume substantial memory and processing power, even though visual information is often sparse and redundant compared to concise textual prompts. While prior efficiency methods either require expensive model retraining or discard visual tokens without considering the user prompt, the article demonstrates that text instructions must guide which visual parts are kept.

The main objective of the article is to introduce and evaluate SparseVLM, a training-free framework designed to accelerate vision-language model inference. The approach dynamically identifies and prunes redundant visual data by using the context of textual prompts without requiring additional model parameters or fine-tuning.

To evaluate this framework, the authors conducted extensive empirical tests across multiple vision-language architectures, including LLaVA, Mini-Gemini, Qwen2-VL, and Video-LLaVA. Testing spanned eight standard image-understanding benchmarks and four video question-answering datasets. The SparseVLM mechanism operates directly within the model's self-attention layers by first selecting visually relevant text words to act as raters. It then calculates the importance of each visual token against these text raters, determines layer-by-layer pruning ratios based on attention matrix redundancy, and aggregates discarded tokens into compact summary representations to minimize information loss.

The findings confirm substantial performance and efficiency gains across multiple benchmarks. When applied to LLaVA, SparseVLM eliminated 66.7% to 77.8% of visual tokens, reducing processing latency by 37% to 43.1% and computational floating-point operations by up to 62.8%, while retaining 97% to 99.1% of the baseline model's accuracy. Furthermore, under heavy sparsification retaining only 64 tokens, SparseVLM outperformed prior leading acceleration methods by 17.3%. In video understanding benchmarks, SparseVLM pruned 90.5% of visual tokens and achieved an average accuracy of 95.0% relative to the uncompressed baseline, outperforming the competing FastV method by 14.7%.

These results demonstrate that vision-language models can be deployed much more affordably and quickly on edge devices and cloud infrastructure. Because SparseVLM requires no additional training, organizations can instantly integrate it into existing model pipelines to lower server operating costs, cut memory cache requirements by 67%, and improve response latency with minimal impact on accuracy.

Engineering teams and technology decision-makers should consider piloting SparseVLM as a plug-and-play optimization for latency-critical and high-volume multimodal applications. While the technique exhibits high reliability across varied benchmarks, practitioners should conduct validation on task-specific domains where extreme token reduction might risk discarding fine visual details.

Cover for SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Abstract

In vision-language models (VLMs), visual tokens usually bear a significant amount of computational overhead despite sparsity of information in them when compared to text tokens. To address this, most existing methods learn a network to prune redundant visual tokens using certain training data. Differently, we propose a text-guided training-free token optimization mechanism dubbed SparseVLM that eliminates the need of extra parameters or fine-tuning costs. Given that visual tokens complement text tokens in VLM’s linguistic reasoning, we select relevant text tokens to rate the significance of visual tokens using self-attention matrices and, then, prune visual tokens using the proposed strategy to maximize sparsity while retaining information. In particular, we introduce a rank-based strategy to adaptively determine the sparsification ratio for each layer, alongside a token recycling method that compresses pruned tokens into more compact representations. Experimental results show that SparseVLM increases the efficiency of various VLMs in a number of image and video understanding tasks. Our code is available at https://github.com/Gumpest/SparseVLMs.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Method
  • 3.1. Preliminary: Attention in VLM Decoders
  • 3.2. Sparsification Guidance from Text to Vision
  • 3.3. Visual Token Recycling
  • 3.4. Theoretical Analysis of Computational Complexity
  • 4. Experiments
  • 4.1. Image Understanding Tasks
  • 4.2. Video Understanding Tasks
  • 5. Analysis
  • 5.1. Relevant Text Token Selection
  • 5.2. Recycling of Pruned Tokens
  • 5.3. Computational Efficiency
  • 5.4. Qualitative Visualization
  • 6. Conclusion
  • Acknowledgments
  • Impact Statement
  • References
  • Appendix
  • A. The Redundancy of Visual Tokens in VLMs
  • B. Compatibility with FlashAttention
  • C. Computing Budget Detailed Estimation
  • D. Dataset
  • E. Implementation Details
  • F. Efficiency Details
  • G. More Detailed Efficiency Analysis
  • H. More Sparsification Visualization

Knowls

  1. Knowl 1 — SparseVLM Framework Overview

    model/method

    SparseVLM is a training-free, text-guided visual token sparsification framework designed to accelerate vision-language model (VLM) decoder inference without requiring fine-tuning or auxiliary learnable parameters. The framework operates in two main phases:

    1. Pre-LLM Text Rater Selection: Before feeding sequences into the LLM decoder, text prompt tokens are evaluated against visual embedding features. Only text tokens exhibiting strong cross-modal correlation are retained as "raters" to guide visual pruning.
    2. In-Decoder Adaptive Sparsification and Recycling: Within each transformer decoder layer, visual token significance is scored using the self-attention weights between the selected text raters and visual tokens. The number of visual tokens to prune per layer is determined dynamically using the matrix rank of this cross-modal attention slice. Pruned visual tokens with moderate relevance are not simply discarded; instead, they are clustered via density peak aggregation and merged into compact reconstructed tokens to mitigate visual information loss.
  2. Knowl 2 — Relevant Text Rater Selection Mechanism

    model/method

    To prevent irrelevant textual prompt tokens (such as pronouns and prepositions) from corrupting the evaluation of visual token importance, SparseVLM pre-selects visually relevant text tokens as raters prior to decoder layer execution.

    Let an input image xvx_v have visual features Zv=g(xv)∈RLv×DinZ_v = g(x_v) \in \mathbb{R}^{L_v \times D_{in}} from visual encoder g(⋅)g(\cdot), projected into vision embedding tokens Hv=WZv∈RLv×DH_v = W Z_v \in \mathbb{R}^{L_v \times D} via projection matrix WW. Let the language prompt xqx_q be tokenized into text embeddings Hq∈RLt×DH_q \in \mathbb{R}^{L_t \times D}, where DD is the embedding dimension, LvL_v is the number of visual tokens, and LtL_t is the number of text tokens.

    The relevance score rir_i for each text token i∈{1,2,…,Lt}i \in \{1, 2, \dots, L_t\} across all visual tokens is computed as: r=1Lv∑j=1Lv(Softmax(HvHqT))jr = \frac{1}{L_v} \sum_{j=1}^{L_v} \left( \text{Softmax}\left( H_v H_q^T \right) \right)_j

    The threshold is defined as the mean score across text tokens, m=mean(r)=1Lt∑i=1Ltrim = \text{mean}(r) = \frac{1}{L_t} \sum_{i=1}^{L_t} r_i. The subset of selected text raters ss is: s={i∣ri≥m, i∈{1,2,…,Lt}}s = \{i \mid r_i \ge m, \, i \in \{1, 2, \dots, L_t\}\}

    This selection is computed once before the decoder layers at a cost of Lt×Lv×2DL_t \times L_v \times 2D FLOPs.

  3. Knowl 3 — Visual Token Significance Scoring via Attention Priority Matrix

    equation

    Given causal self-attention logits matrix A∈RL×LA \in \mathbb{R}^{L \times L} in a VLM decoder layer (where LL is sequence length including visual and text tokens), the sparsification priority matrix P∈RLt×LvP \in \mathbb{R}^{L_t \times L_v} is extracted as the sub-matrix corresponding to queries of the text instruction tokens L\mathcal{L} and keys of the visual tokens I\mathcal{I}: P=Ai,j,(i,j)∈{L,I}P = A_{i,j}, \quad (i, j) \in \{\mathcal{L}, \mathcal{I}\}

    The visual token significance indicator vector p~=[p~1,p~2,…,p~Lv]∈RLv\tilde{p} = [\tilde{p}_1, \tilde{p}_2, \dots, \tilde{p}_{L_v}] \in \mathbb{R}^{L_v} is calculated by averaging attention scores over the selected text rater subset ss: p~=1∣s∣∑i∈sPi\tilde{p} = \frac{1}{|s|} \sum_{i \in s} P_i

    A higher value in p~j\tilde{p}_j indicates that visual token jj has stronger cross-modal correlation with the question prompt and should be prioritized for retention. Computing p~\tilde{p} requires ∣s∣×Lv|s| \times L_v FLOPs given pre-computed attention logits.

  4. Knowl 4 — Rank-Based Adaptive Layer-Wise Visual Token Pruning

    model/method

    Because different image inputs exhibit varying levels of visual redundancy, SparseVLM determines the number of visual tokens to prune at each decoder layer adaptively using the matrix rank of the sparsification priority matrix P∈RLt×LvP \in \mathbb{R}^{L_t \times L_v}.

    The difference between the number of visual tokens LvL_v and rank(P)\text{rank}(P) represents information redundancy. The number of tokens to delete, NN, is defined as: N=λ×(Lv−rank(P))N = \lambda \times \left( L_v - \text{rank}(P) \right) where λ\lambda is a predefined scaling hyperparameter. The rank is computed via singular value decomposition (SVD) counting singular values exceeding a numerical threshold, with computational complexity Lt×Lv×min⁡(Lt,Lv)L_t \times L_v \times \min(L_t, L_v) FLOPs.

    If N=0N = 0, pruning is skipped for that layer. Otherwise, the NN visual tokens with the lowest significance scores in p~\tilde{p} are removed from the active sequence and placed into a deletion pool.

  5. Knowl 5 — Visual Token Recycling and Semantic Reconstruction

    algorithm

    To prevent loss of visual detail during progressive sparsification, the top-τ\tau fraction of pruned visual tokens from the deleted pool are clustered using kk-nearest neighbor density peak aggregation and merged into compact representative tokens.

    Input: Pruned visual tokens hˉv∈RLr×D\bar{h}_v \in \mathbb{R}^{L_r \times D} (top-τ\tau highest attention values among NN pruned tokens, Lr=τNL_r = \tau N), neighborhood size kk, center ratio θ\theta, hidden dimension DD
    Output: Reconstructed tokens {T1,T2,…,TC}\{T_1, T_2, \dots, T_C\} with C=θLrC = \theta L_r
    for each token hˉiv\bar{h}_i^v in hˉv\bar{h}_v do
        Find kk-nearest neighbors K(hˉiv)\mathcal{K}(\bar{h}_i^v)
        Compute local density ρi=exp⁡(−1k∑hˉjv∈K(hˉiv)∥hˉiv−hˉjv∥22)\rho_i = \exp\left( -\frac{1}{k} \sum_{\bar{h}_j^v \in \mathcal{K}(\bar{h}_i^v)} \|\bar{h}_i^v - \bar{h}_j^v\|_2^2 \right)
    end for
    for each token hˉiv\bar{h}_i^v in hˉv\bar{h}_v do
        if there exists jj such that ρj>ρi\rho_j > \rho_i then
            δi=min⁡j:ρj>ρi∥hˉiv−hˉjv∥2\delta_i = \min_{j: \rho_j > \rho_i} \|\bar{h}_i^v - \bar{h}_j^v\|_2
        else
            δi=max⁡j∥hˉiv−hˉjv∥2\delta_i = \max_j \|\bar{h}_i^v - \bar{h}_j^v\|_2
        end if
        scorei=ρi×δi\text{score}_i = \rho_i \times \delta_i
    end for
    Select indices of top CC scores as cluster centers {c1,c2,…,cC}\{c_1, c_2, \dots, c_C\}
    Assign each remaining token in hˉv\bar{h}_v to its nearest center via cosine similarity, forming groups G1,…,GC\mathcal{G}_1, \dots, \mathcal{G}_C
    for k=1k = 1 to CC do
        Tk=∑t∈GktT_k = \sum_{t \in \mathcal{G}_k} t
    end for
    return {T1,T2,…,TC}\{T_1, T_2, \dots, \mathcal{T}_C\}

    The aggregation step requires Lr(3Lr−1)×2D+LrL_r(3L_r - 1) \times 2D + L_r FLOPs, and token reconstruction requires D(Lr−C)D(L_r - C) FLOPs.

  6. Knowl 6 — Dual-Pass FlashAttention Integration for Training-Free Token Pruning

    model/method

    Because FlashAttention avoids materializing the full attention matrix AA, SparseVLM extracts average attention scores relative to text raters via a dual-pass FlashAttention mechanism:

    1. First Forward Pass: Standard FlashAttention is executed to produce hidden states.
    2. Second Forward Pass: A customized value matrix V∈Rn×dV \in \mathbb{R}^{n \times d} is supplied where rows corresponding to the selected text raters s={i1,i2,…,i∣s∣}s = \{i_1, i_2, \dots, i_{|s|}\} contain entries Vij=1∣s∣V_{ij} = \frac{1}{|s|}, and Vij=0V_{ij} = 0 otherwise. The block-wise FlashAttention inner product OB=PBVBO_B = P_B V_B incrementally accumulates the mean attention score across raters directly within GPU SRAM without explicit full-matrix storage.
    3. Mask Generation and Pruning: From the output OvO_v corresponding to visual tokens, a top-kk selection is performed (Ik={i∣xi∈Ov,rank(xi,Ov)≤k}I_k = \{i \mid x_i \in O_v, \text{rank}(x_i, O_v) \le k\}). Pruned tokens are masked out from the hidden states computed during the first forward pass.
  7. Knowl 7 — Theoretical FLOPs Reduction of SparseVLM

    theoretical result

    In a transformer decoder with Ω\Omega layers, hidden dimension DD (matching the intermediate size of the feed-forward network), LtL_t text tokens, and LvL_v visual tokens, pruning NiN_i visual tokens and reconstructing CiC_i tokens at layer i∈{1,2,…,Ω}i \in \{1, 2, \dots, \Omega\} yields a net FLOPs reduction defined as:

    ΔFLOPs=∑i=1Ω(6(Ni−Ci)D2+2(Ni−Ci)2D)−Overhead\Delta \text{FLOPs} = \sum_{i=1}^\Omega \left( 6(N_i - C_i)D^2 + 2(N_i - C_i)^2 D \right) - \text{Overhead}

    where the overhead across layers and rater selection is: Overhead=2LtLvD+∑i=1Ω(LtiLvi(1+min⁡(Lti,Lvi))+(6(Lri)2+2Lri)D+Lri+D(Lri−Ci))\text{Overhead} = 2 L_t L_v D + \sum_{i=1}^\Omega \left( L_t^i L_v^i \left( 1 + \min(L_t^i, L_v^i) \right) + (6(L_r^i)^2 + 2L_r^i)D + L_r^i + D(L_r^i - C_i) \right) with Lri=τNiL_r^i = \tau N_i and Ci=θLriC_i = \theta L_r^i. Assuming x=τ×θ≪1x = \tau \times \theta \ll 1, this simplifies to: ΔFLOPs≈−2LtLvD+∑i=1ΩDNi(6D+2Ni)−(Lti)2Lvi\Delta \text{FLOPs} \approx -2 L_t L_v D + \sum_{i=1}^\Omega D N_i (6D + 2 N_i) - (L_t^i)^2 L_v^i

  8. Knowl 8 — SparseLLaVA Performance Across Image Understanding Benchmarks

    data/table

    The table below evaluates SparseLLaVA (LLaVA-1.5 with SparseVLM) on eight image understanding benchmarks under varying numbers of retained visual tokens (vanilla = 576 tokens) on an NVIDIA A100-80GB GPU. Performance is shown as raw benchmark scores and relative percentage to the vanilla model.

    Method GQA MMB MME POPE SQA SEED VQAText^{\text{Text}} MMVet Acc. (%) FLOPs (T) Latency (ms)
    Vanilla (576 Tok) 61.9 64.6 1864 85.9 69.5 60.3 58.3 30.9 100.0 4.62 57.82
    Retain 192 Tokens (66.7% Pruned)
    ToMe 54.3 60.5 1563 72.4 65.2 53.1 52.1 27.9 88.9 2.05 34.06
    FastV 52.6 61.0 1605 64.8 69.1 52.1 52.5 26.7 87.9 2.11 34.87
    PDrop 57.1 63.2 1766 82.3 70.2 54.7 56.1 30.5 95.9 2.03 36.74
    SparseVLM 59.5 64.1 1787 85.3 68.7 58.7 57.8 33.1 99.1 2.14 36.50
    Retain 128 Tokens (77.8% Pruned)
    ToMe 52.4 53.3 1343 62.8 59.6 50.9 49.1 27.2 81.9 1.62 30.00
    FastV 49.6 56.1 1490 53.4 68.6 48.1 50.5 26.3 82.4 1.70 30.70
    PDrop 56.0 61.1 1664 82.3 69.9 53.3 55.1 30.8 94.3 1.62 37.77
    SparseVLM 58.4 64.5 1746 85.0 68.6 58.2 56.7 29.0 96.7 1.72 33.28
    Retain 64 Tokens (88.9% Pruned)
    ToMe 48.6 43.7 1138 52.5 50.0 44.0 45.3 24.1 71.1 1.19 26.52
    FastV 46.1 47.2 1255 38.2 68.7 43.7 47.8 19.6 72.0 1.29 27.30
    PDrop 41.9 33.3 1092 55.9 69.2 40.0 45.9 30.7 73.4 1.18 43.41
    SparseVLM 53.8 60.1 1589 77.5 69.8 52.2 53.4 24.9 89.3 1.30 29.89

    SparseVLM retains 99.1% average accuracy when keeping 192 tokens (a 36.9% latency decrease), and 89.3% accuracy when retaining only 64 tokens (an 88.9% compression ratio), outperforming FastV by 17.3% at 64 tokens.

  9. Knowl 9 — Video Question Answering Performance of SparseVideoLLaVA

    data/table

    SparseVLM applied to Video-LLaVA evaluates video QA benchmarks where 2048 initial video tokens are pruned down to 194 tokens (a 90.5% token reduction). Evaluation uses accuracy (Acc.) and GPT evaluation score (Score) across TGIF-QA, MSVD-QA, MSRVTT-QA, and ActivityNet-QA on the first 1000 samples per benchmark.

    Method TGIF MSVD MSRVTT ActivityNet Average
    Acc. Score Acc. Score Acc. Score Acc. Score Acc. Score
    Video-LLaVA (2048 Tok) 18.9 2.54 72.0 3.95 57.1 3.45 43.6 3.81 47.9 3.44
    FastV (194 Tok) 10.2 2.29 58.3 3.62 52.3 3.42 41.3 3.76 40.5 3.27
    SparseVLM (194 Tok) 14.9 2.41 71.7 3.94 56.1 3.43 45.1 3.81 47.0 3.40

    SparseVLM preserves 95.0% of Video-LLaVA's original average accuracy (losing only 0.04 GPT evaluation points), outperforming FastV by 14.7% in relative average accuracy under identical token budgets.

  10. Knowl 10 — Ablation Analysis of Text Rater Guidance and Token Reconstruction

    empirical result

    Ablation experiments on LLaVA-7B isolate the contributions of relevant text rater selection and visual token reconstruction (TR):

    1. Text Rater Guidance vs. Unfiltered Tokens: Retaining 64 visual tokens on LLaVA-7B:
      • On TextVQA, using selected text raters achieves 53.4% accuracy, compared to 53.0% using all text tokens and 52.6% using all tokens (a +0.8% gain over unguided baseline).
      • On POPE, selected text raters achieve 77.5% accuracy, compared to 74.8% using all text tokens and 66.1% using all tokens (+11.4% gain over all tokens, +2.7% over all text tokens).
    2. Token Reconstruction Across Pruning Levels: On GQA and POPE benchmarks, adding token reconstruction (+TR) improves accuracy across all retained token counts (64, 96, 128, 192):
      • GQA accuracy improves from 52.2% to 53.8% at 64 tokens, 55.2% to 56.4% at 96 tokens, 58.1% to 58.4% at 128 tokens, and 59.4% to 59.5% at 192 tokens (average +0.8%).
      • POPE accuracy improves from 72.8% to 77.5% at 64 tokens (+4.7%), 77.5% to 81.9% at 96 tokens (+4.4%), 83.7% to 85.0% at 128 tokens (+1.3%), and 85.2% to 85.3% at 192 tokens (+0.1%) (average +2.6%). The performance benefit of token recycling increases as the pruning ratio grows.

Coverage note — No substantial contributed material was omitted. All primary methodology components (rater selection, rank-based adaptation, token recycling, and FlashAttention integration) and core empirical results (LLaVA-1.5, Video-LLaVA, Qwen2-VL, complexity analysis, and ablations) are fully covered.

References

  1. 1.Alayrac, J.-B., Donahue, J., Luc, P., Miech, A., Barr, I., Hasson, Y., Lenc, K., Mensch, A., Millican, K., Reynolds, M., et al. Flamingo: a visual language model for few-shot learning. Advances in Neural Information Processing Systems, 2022.
  2. 2.Bai, J., Bai, S., Yang, S., Wang, S., Tan, S., Wang, P., Lin, J., Zhou, C., and Zhou, J. Qwen-VL: A frontier large vision-language model with versatile abilities. arXiv:2308.12966, 2023.
  3. 3.Bolya, D., Fu, C.-Y., Dai, X., Zhang, P., Feichtenhofer, C., and Hoffman, J. Token merging: Your vit but faster. In International Conference on Learning Representations, 2023.
  4. 4.Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. Language models are few-shot learners. Advances in Neural Information Processing Systems, 2020.
  5. 5.Cai, M., Yang, J., Gao, J., and Lee, Y. J. Matryoshka multimodal models. In International Conference on Learning Representations, 2025.
  6. 6.Cha, J., Kang, W., Mun, J., and Roh, B. Honeybee: Locality-enhanced projector for multimodal llm. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2024.
  7. 7.Chen, L., Zhao, H., Liu, T., Bai, S., Lin, J., Zhou, C., and Chang, B. An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models. In Proceedings of the European Conference on Computer Vision, 2024a.
  8. 8.Chen, Z., Wu, J., Wang, W., Su, W., Chen, G., Xing, S., Zhong, M., Zhang, Q., Zhu, X., Lu, L., et al. Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2024b.
  9. 9.Dai, W., Li, J., Li, D., Tiong, A., Zhao, J., Wang, W., Li, B., Fung, P., and Hoi, S. InstructBLIP: Towards general-purpose vision-language models with instruction tuning. Advances in Neural Information Processing Systems, 2023.
  10. 10.Dao, T., Fu, D., Ermon, S., Rudra, A., and Ré, C. FlashAttention: Fast and memory-efficient exact attention with io-awareness. Advances in Neural Information Processing Systems, 2022.
  11. 11.Du, Z., Qian, Y., Liu, X., Ding, M., Qiu, J., Yang, Z., and Tang, J. Glm: General language model pretraining with autoregressive blank infilling. In Proceedings of the Annual Meeting of the Association for Computational Linguistics, 2022.
  12. 12.Fu, C., Chen, P., Shen, Y., Qin, Y., Zhang, M., Lin, X., Yang, J., Zheng, X., Li, K., Sun, X., et al. MME: A comprehensive evaluation benchmark for multimodal large language models. arXiv:2306.13394, 2023.
  13. 13.Goyal, Y., Khot, T., Summers-Stay, D., Batra, D., and Parikh, D. Making the v in vqa matter: Elevating the role of image understanding in visual question answering. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2017.
  14. 14.Hudson, D. A. and Manning, C. D. GQA: A new dataset for real-world visual reasoning and compositional question answering. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2019.
  15. 15.Jang, Y., Song, Y., Yu, Y., Kim, Y., and Kim, G. Tgif-qa: Toward spatio-temporal reasoning in visual question answering. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2017.
  16. 16.Kondratyuk, D., Yu, L., Gu, X., Lezama, J., Huang, J., Hornung, R., Adam, H., Akbari, H., Alon, Y., Birodkar, V., et al. Videopoet: A large language model for zero-shot video generation. In International Conference on Machine Learning, 2024.
  17. 17.Li, B., Ge, Y., Ge, Y., Wang, G., Wang, R., Zhang, R., and Shan, Y. Seed-bench: Benchmarking multimodal large language models. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2024a.
  18. 18.Li, J., Li, D., Savarese, S., and Hoi, S. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In International Conference on Machine Learning, 2023a.
  19. 19.Li, Y., Du, Y., Zhou, K., Wang, J., Zhao, W. X., and Wen, J.-R. Evaluating object hallucination in large vision-language models. In Proceedings of the Conference on Empirical Methods in Natural Language Processing, 2023b.
  20. 20.Li, Y., Wang, C., and Jia, J. LLaMA-VID: An image is worth 2 tokens in large language models. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2024b.
  21. 21.Li, Y., Zhang, Y., Wang, C., Zhong, Z., Chen, Y., Chu, R., Liu, S., and Jia, J. Mini-gemini: Mining the potential of multi-modality vision language models. arXiv:2403.18814, 2024c.
  22. 22.Lin, B., Ye, Y., Zhu, B., Cui, J., Ning, M., Jin, P., and Yuan, L. Video-llava: Learning united visual representation by alignment before projection. In Proceedings of the Conference on Empirical Methods in Natural Language Processing, 2024.
  23. 23.Liu, H., Li, C., Li, Y., and Lee, Y. J. Improved baselines with visual instruction tuning. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2024a.
  24. 24.Liu, H., Li, C., Wu, Q., and Lee, Y. J. Visual instruction tuning. Advances in Neural Information Processing Systems, 2024b.
  25. 25.Liu, Y., Duan, H., Zhang, Y., Li, B., Zhang, S., Zhao, W., Yuan, Y., Wang, J., He, C., Liu, Z., et al. Mmbench: Is your multi-modal model an all-around player? In Proceedings of the European Conference on Computer Vision, 2024c.
  26. 26.Liu, Z., Mao, H., Wu, C.-Y., Feichtenhofer, C., Darrell, T., and Xie, S. A convnet for the 2020s. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2022.
  27. 27.Lu, P., Mishra, S., Xia, T., Qiu, L., Chang, K.-W., Zhu, S.-C., Tafjord, O., Clark, P., and Kalyan, A. Learn to explain: Multimodal reasoning via thought chains for science question answering. Advances in Neural Information Processing Systems, 2022.
  28. 28.Maaz, M., Rasheed, H., Khan, S., and Khan, F. Video-chatgpt: Towards detailed video understanding via large vision and language models. In Proceedings of the Annual Meeting of the Association for Computational Linguistics, 2024.
  29. 29.Marr, D. Vision: A computational investigation into the human representation and processing of visual information. MIT press, 2010.
  30. 30.Peng, B., Li, C., He, P., Galley, M., and Gao, J. Instruction tuning with gpt-4. arXiv:2304.03277, 2023.
  31. 31.Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al. Language models are unsupervised multitask learners. OpenAI blog, 2019.
  32. 32.Rodriguez, A. Clustering by fast search and find of density peaks. Science, 2014.
  33. 33.Shang, Y., Cai, M., Xu, B., Lee, Y. J., and Yan, Y. Llava-prumerge: Adaptive token reduction for efficient large multimodal models. arXiv preprint arXiv:2403.15388, 2024.
  34. 34.Singh, A., Natarjan, V., Shah, M., Jiang, Y., Chen, X., Parikh, D., and Rohrbach, M. Towards VQA models that can read. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2019.
  35. 35.Stewart, G. W. On the early history of the singular value decomposition. SIAM review, 1993.
  36. 36.Team, G., Anil, R., Borgeaud, S., Wu, Y., Alayrac, J.-B., Yu, J., Soricut, R., Schalkwyk, J., Dai, A. M., Hauth, A., et al. Gemini: a family of highly capable multimodal models. arXiv:2312.11805, 2023.
  37. 37.Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al. Llama: Open and efficient foundation language models. arXiv:2302.13971, 2023.
  38. 38.Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. Attention is all you need. Advances in Neural Information Processing Systems, 2017.
  39. 39.Wu, S., Chen, J., Lin, K. Q., Wang, Q., Gao, Y., Xu, Q., Xu, T., Hu, Y., Chen, E., and Shou, M. Z. Videollm-mod: Efficient video-language streaming with mixture-of-depths vision computation. Advances in Neural Information Processing Systems, 2024.
  40. 40.Xing, L., Huang, Q., Dong, X., Lu, J., Zhang, P., Zang, Y., Cao, Y., He, C., Wang, J., Wu, F., et al. Pyramiddrop: Accelerating your large vision-language models via pyramid visual redundancy reduction. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2025.
  41. 41.Xu, D., Zhao, Z., Xiao, J., Wu, F., Zhang, H., He, X., and Zhuang, Y. Video question answering via gradually refined attention over appearance and motion. In Proceedings of the ACM international conference on Multimedia, 2017.
  42. 42.Yao, L., Li, L., Ren, S., Wang, L., Liu, Y., Sun, X., and Hou, L. DeCo: Decoupling token compression from semantic abstraction in multimodal large language models. arXiv:2405.20985, 2024.
  43. 43.Ye, X., Gan, Y., Huang, X., Ge, Y., Shan, Y., and Tang, Y. VoCo-LLaMA: Towards vision compression with large language models. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2025.
  44. 44.Yu, W., Yang, Z., Li, L., Wang, J., Lin, K., Liu, Z., Wang, X., and Wang, L. Mm-vet: Evaluating large multimodal models for integrated capabilities. In International Conference on Machine Learning, 2024.
  45. 45.Yu, Z., Xu, D., Yu, J., Yu, T., Zhao, Z., Zhuang, Y., and Tao, D. Activitynet-qa: A dataset for understanding complex web videos via question answering. In AAAI, 2019.
  46. 46.Zhang, Y., Huang, T., Fan, C.-K., Dong, H., Li, J., Wang, J., Cheng, K., Zhang, S., Guo, H., et al. Unveiling the tapestry of consistency in large vision-language models. Advances in Neural Information Processing Systems, 2024a.
  47. 47.Zhang, Y., Huang, T., Liu, J., Jiang, T., Cheng, K., and Zhang, S. Freekd: Knowledge distillation via semantic frequency prompt. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2024b.
  48. 48.Zhu, B., Lin, B., Ning, M., Yan, Y., Cui, J., HongFa, W., Pang, Y., Jiang, W., Zhang, J., Li, Z., et al. Languagebind: Extending video-language pretraining to n-modality by language-based semantic alignment. In International Conference on Learning Representations, 2024a.
  49. 49.Zhu, D., Chen, J., Shen, X., Li, X., and Elhoseiny, M. Minigpt-4: Enhancing vision-language understanding with advanced large language models. In International Conference on Learning Representations, 2024b.

Citation

MLA
Zhang, Y., et al. “SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference”. arXiv, 2024, http://arxiv.org/abs/2410.04417v4.
APA
Zhang, Y., Fan, C.-K., Ma, J., Zheng, W., Huang, T., Cheng, K., Gudovskiy, D., Okuno, T., Nakata, Y., Keutzer, K., & Zhang, S. (2024). SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference. arXiv. http://arxiv.org/abs/2410.04417v4
Chicago
Zhang, Y., C.-K. Fan, J. Ma, et al. 2024. “SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference”. arXiv. http://arxiv.org/abs/2410.04417v4.
Harvard
Zhang, Y. et al. (2024) “SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2410.04417v4.
Vancouver
1. Zhang Y, Fan C-K, Ma J, et al (2024) SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference. arXiv

BibTeX

@article{zhang2024sparsevlm,
  title = {SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference},
  author = {Zhang, Yuan and Fan, Chun-Kai and Ma, Junpeng and Zheng, Wenzhao and Huang, Tao and Cheng, Kuan and Gudovskiy, Denis and Okuno, Tomoyuki and Nakata, Yohei and Keutzer, Kurt and Zhang, Shanghang},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2410.04417v4},
  eprint = {2410.04417}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/