VisionZip: Longer is Better but Not Necessary in Vision Language Models

Senqiao YangYukang ChenZhuotao TianChengyao WangJingyao LiBei YuJiaya Jia

article2025CVPR336 citations

Proposes VisionZip, a visual token reduction method that selects dominant tokens and merges redundant ones from vision encoders to achieve an 8x prefilling speedup with minimal performance loss across image and video benchmarks.

Listen

Modern vision-language models achieve impressive visual and reasoning performance, but they do so at steep computational and memory costs. These architectures convert high-resolution images and video frames into thousands of visual tokens—often vastly outnumbering text tokens by a factor of twenty or more. This creates severe inference latency and resource bottlenecks that restrict deployment in latency-sensitive, resource-constrained environments such as robotics, edge devices, and autonomous systems.

The article evaluates whether all generated visual tokens are truly necessary, demonstrating that standard vision encoders suffer from extreme feature redundancy. It introduces VisionZip, a text-agnostic framework designed to aggressively reduce visual token counts while preserving model performance across diverse multimodal benchmarks.

To demonstrate this, the researchers analyzed token attention patterns within common vision encoders and evaluated VisionZip across eleven image understanding benchmarks and four video question-answering benchmarks. The approach was tested across leading open-source model families—including LLaVA-1.5, LLaVA-NeXT, and Mini-Gemini—under both training-free inference settings and an efficient fine-tuning mode that adapts only the cross-modality projector using a tiny subset of data.

The findings show that attention within vision encoders heavily concentrates into a small subset of dominant proxy tokens, leaving most visual tokens with near-zero attention. VisionZip capitalizes on this by selecting dominant tokens based on attention weights and merging remaining non-dominant tokens based on semantic similarity. Across all evaluated settings, VisionZip retained approximately 95% of original model performance using only 5% to 10% of the visual tokens. On the POPE benchmark with LLaVA-NeXT 7B, it reduced initial token prefilling time by 7.8 times and improved overall inference runtime threefold. Furthermore, because VisionZip selects tokens independently of input text prompts, it outperformed text-dependent pruning baselines by over 5% and maintained robust performance across multi-turn conversational caching without discarding context needed for follow-up questions.

These results demonstrate that increasing visual token length delivers diminishing returns while dramatically driving up operational compute costs. By extracting concentrated visual representations at the encoder stage, organizations can deploy larger, more capable language models—such as a 13-billion parameter model—at faster speeds and lower hardware costs than standard smaller models. Prior text-guided token pruning methods suffer from spatial feature misalignment, whereas prompt-agnostic token reduction provides a more reliable path for production-ready vision-language systems.

Engineering and deployment teams should consider adopting plug-and-play visual token reduction techniques like VisionZip to improve inference throughput and reduce serving infrastructure costs. For teams seeking maximal fidelity, performing a lightweight, 30-minute projector fine-tuning step is recommended to bridge the space between reduced vision tokens and language decoders. Future research and development should focus on training natively dense vision encoders that eliminate feature redundancy upstream rather than relying on oversized token sequences.

While confidence in the empirical results is high across standard vision-language benchmarks and common vision architectures, the evaluations rely on existing benchmark suites and standard self-attention encoders. Stakeholders should conduct pilot validations within their specific domain-specific tasks and production data streams to verify that aggressive token reduction maintains sufficient fine-grained visual accuracy.

Cover for VisionZip: Longer is Better but Not Necessary in Vision Language Models

Abstract

Recent advancements in vision-language models have enhanced performance by increasing the length of visual tokens, making them much longer than text tokens and significantly raising computational costs. However, we observe that the visual tokens generated by popular vision encoders, such as CLIP and SigLIP, contain significant redundancy. To address this, we introduce VisionZip, a simple yet effective method that selects a set of informative tokens for input to the language model, reducing visual token redundancy and improving efficiency while maintaining model performance. The proposed VisionZip can be widely applied to image and video understanding tasks and is well-suited for multi-turn dialogues in real-world scenarios, where previous methods tend to underperform. Experimental results show that VisionZip outperforms the previous state-of-the-art method by at least 5% performance gains across nearly all settings. Moreover, our method significantly enhances model inference speed, improving the prefilling time by 8× and enabling the LLaVA-Next 13B model to infer faster than the LLaVA-Next 7B model while achieving better results. Furthermore, we analyze the causes of this redundancy and encourage the community to focus on extracting better visual features rather than merely increasing token length. Our code is available at https://github.com/dvlab-research/VisionZip.

Table of Contents

  • 1. Introduction
  • 2. VisionZip
  • 2.1. Preliminary
  • 2.2. Redundancy Observation
  • 2.3. Informative Visual Token Zip
  • 2.4. Efficient Tuning
  • 2.5. Usage of VisionZip
  • 3. Experiments
  • 3.1. Effectiveness on Image Understanding
  • 3.2. Effectiveness on Video Understanding
  • 3.3. Efficiency Analysis
  • 4. Analysis and Discussion
  • 4.1. Reasons of Redundancy in Visual Tokens
  • 4.2. Why VisionZip Outperforms Previous Work?
  • 4.3. The Advantage of the VisionZip
  • 5. Related Work
  • 6. Conclusion
  • 7. Acknowledgments
  • References

Knowls

  1. Knowl 1 — VisionZip Token Reduction Framework

    model/method

    VisionZip is a text-agnostic visual token compression framework designed to accelerate vision-language models (VLMs) by eliminating redundant visual tokens prior to feeding them into the large language model (LLM) decoder.

    Given an image processed by a vision encoder (such as CLIP or SigLIP), the visual tokens at a selected layer (typically the second-to-last layer, −2-2) are partitioned and compressed into two complementary sets:

    1. Dominant Tokens: A small subset of KK visual tokens identified as receiving the highest attention in the vision encoder self-attention layers, thereby preserving critical foreground and core visual features.
    2. Contextual Tokens: The remaining non-dominant tokens are compressed into a smaller set of MM contextual tokens by clustering and merging tokens based on key similarity in the self-attention space, capturing broad context and background details without maintaining redundant individual tokens.

    The combined set of K+MK + M tokens (plus the CLS\text{CLS} token if present) replaces the dense visual sequence. Because the token selection and merging are executed purely using vision encoder internal states, the process is text-agnostic, preserving general visual semantics across multi-turn dialogues and enabling compatibility with arbitrary LLM acceleration algorithms.

  2. Knowl 2 — Dominant Token Selection via Vision Tower Self-Attention

    algorithm

    Dominant token selection extracts the most informative visual tokens based on the self-attention weights generated in the vision encoder.

    Let the vision encoder produce attention weight tensors A∈RB×H×S×S\mathbf{A} \in \mathbb{R}^{B \times H \times S \times S} and hidden states X∈RB×S×D\mathbf{X} \in \mathbb{R}^{B \times S \times D} at the target feature layer, where BB is the batch size, HH is the number of attention heads, SS is the visual sequence length, and DD is the hidden dimension. For models with a class token (e.g., CLIP), the class token index is idxcls\text{idx}_{\text{cls}}. For models without a class token (e.g., SigLIP), attention received is averaged across sequence positions.

    Input: Visual hidden states X∈RB×S×DX \in \mathbb{R}^{B \times S \times D}, attention tensor A∈RB×H×S×SA \in \mathbb{R}^{B \times H \times S \times S}, target dominant token count KK, class token index idxcls\text{idx}_{\text{cls}} (if applicable)
    Output: Dominant visual tokens Xdom∈RB×(K+1)×DX_{\text{dom}} \in \mathbb{R}^{B \times (K+1) \times D} (or RB×K×D\mathbb{R}^{B \times K \times D} if no class token)
    if vision encoder has a class token then
        Arec←∑h=1HA[:,h,idxcls,idxcls+1:]A_{\text{rec}} \leftarrow \sum_{h=1}^H A[:, h, \text{idx}_{\text{cls}}, \text{idx}_{\text{cls}}+1:] # attention received from CLS token
        topkidx←topk(Arec,K,dim=1)\text{topk}_{\text{idx}} \leftarrow \text{topk}(A_{\text{rec}}, K, \text{dim}=1) # indices of top K patch tokens
        dominantidx←concat(idxcls,topkidx+1)\text{dominant}_{\text{idx}} \leftarrow \text{concat}(\text{idx}_{\text{cls}}, \text{topk}_{\text{idx}} + 1)
    else
        Arec←∑h=1H∑j=1SA[:,h,j,:]/SA_{\text{rec}} \leftarrow \sum_{h=1}^H \sum_{j=1}^S A[:, h, j, :] / S # mean attention received across all tokens
        dominantidx←topk(Arec,K,dim=1)\text{dominant}_{\text{idx}} \leftarrow \text{topk}(A_{\text{rec}}, K, \text{dim}=1)
    end if
    Xdom←select(X,dominantidx)X_{\text{dom}} \leftarrow \text{select}(X, \text{dominant}_{\text{idx}})
    return XdomX_{\text{dom}}
  3. Knowl 3 — Contextual Token Merging via Key Similarity

    algorithm

    Contextual token merging compresses the non-dominant visual tokens into a compact set of MM tokens by clustering them according to the cosine/dot-product similarity of their key representations from the vision encoder's self-attention mechanism.

    Input: Non-dominant visual tokens Xrem∈RB×Nrem×DX_{\text{rem}} \in \mathbb{R}^{B \times N_{\text{rem}} \times D}, key projection tensor Krem∈RB×Nrem×DkK_{\text{rem}} \in \mathbb{R}^{B \times N_{\text{rem}} \times D_k}, target contextual token count MM
    Output: Merged contextual tokens Xctx∈RB×M×DX_{\text{ctx}} \in \mathbb{R}^{B \times M \times D}
    # Step 1: Split non-dominant tokens uniformly into target centroids and merge candidates
    Xtgt,Xmrg←uniform_split(Xrem,M)X_{\text{tgt}}, X_{\text{mrg}} \leftarrow \text{uniform\_split}(X_{\text{rem}}, M) # Xtgt∈RB×M×DX_{\text{tgt}} \in \mathbb{R}^{B \times M \times D}, Xmrg∈RB×(Nrem−M)×DX_{\text{mrg}} \in \mathbb{R}^{B \times (N_{\text{rem}}-M) \times D}
    Ktgt,Kmrg←uniform_split(Krem,M)K_{\text{tgt}}, K_{\text{mrg}} \leftarrow \text{uniform\_split}(K_{\text{rem}}, M)
    # Step 2: Compute pairwise dot-product similarity matrix
    Sim←Kmrg⋅Ktgt⊤\text{Sim} \leftarrow K_{\text{mrg}} \cdot K_{\text{tgt}}^\top # shape: (B,Nrem−M,M)(B, N_{\text{rem}} - M, M)
    # Step 3: Assign each merge token to the most similar target centroid
    assignidx←arg⁡max⁡(Sim,dim=2)\text{assign}_{\text{idx}} \leftarrow \arg\max(\text{Sim}, \text{dim}=2) # shape: (B,Nrem−M)(B, N_{\text{rem}} - M)
    # Step 4: Average pool assigned merge tokens into their respective target centroids
    Xctx←average_merge(assignidx,Xtgt,Xmrg)X_{\text{ctx}} \leftarrow \text{average\_merge}(\text{assign}_{\text{idx}}, X_{\text{tgt}}, X_{\text{mrg}})
    return XctxX_{\text{ctx}}
  4. Knowl 4 — Visual Token Redundancy via Softmax Gradient Dynamics in Deep Vision Encoders

    theoretical result

    In vision transformer encoders (such as ViT in CLIP and SigLIP), self-attention layers aggregate information across sequence positions via the standard softmax function:

    softmax(zi)=ezi∑j=1nezj\text{softmax}(z_i) = \frac{e^{z_i}}{\sum_{j=1}^n e^{z_j}}

    where zi∈Rz_i \in \mathbb{R} represents the attention logit for token position ii, and nn is the sequence length. The partial derivative of the softmax output with respect to the input logit ziz_i is:

    ∂softmax(zi)∂zi=softmax(zi)⋅(1−softmax(zi))\frac{\partial \text{softmax}(z_i)}{\partial z_i} = \text{softmax}(z_i) \cdot (1 - \text{softmax}(z_i))

    This derivative exhibits an exponential increase in gradient magnitude for high-logit regions, while remaining near zero for low-logit regions. During training, this dynamic creates a self-reinforcing effect: tokens that initially receive slightly higher attention rapidly attract increasing attention, while low-attention positions receive negligible gradient updates.

    As layer depth increases, the vision encoder shortcuts global information aggregation by concentrating image semantics into a few "proxy tokens" (or "attention sinks"). Consequently, by layer −2-2 (layer 23 in a 24-layer ViT), attention is focused almost entirely on a small subset of dominant tokens, rendering the majority of remaining visual patch tokens redundant.

  5. Knowl 5 — Semantic Misalignment of Text-Conditioned Pruning with Vision Encoder Representations

    empirical result

    Text-conditioned visual token pruning approaches (such as FastV and SparseVLM) select visual tokens in the LLM based on cross-attention between text query tokens and image tokens. However, because deep vision transformer encoders aggregate semantic information into peripheral or background "proxy tokens" rather than maintaining detailed representation solely on foreground object patches, text-conditioned attention selects patches directly overlapping semantic objects that actually contain minimal aggregated information after deep vision encoding.

    This misalignment is quantitatively demonstrated on the TextVQA benchmark using SparseVLM constrained to retain 64 visual tokens out of 576:

    1. Standard SparseVLM: Pruning 576 tokens directly down to 64 achieves an accuracy of 51.1%51.1\%.
    2. Ex1 (Masking Dominant Tokens): Removing the top 50 dominant tokens identified by the vision encoder and having SparseVLM select 64 tokens from the remaining 526 tokens causes performance to drop to 46.4%46.4\% (a relative decrease of 9.2%9.2\%).
    3. Ex2 (Pre-filtering to Dominant Tokens): Providing SparseVLM with only the top 128 vision-encoder-dominant tokens selected by VisionZip to pick the final 64 tokens improves accuracy to 52.5%52.5\% (a relative increase of +2.7%+2.7\%).

    These experiments confirm that text-guided token selection is misaligned with the proxy tokens where vision encoders physically aggregate visual information.

  6. Knowl 6 — Efficient Multimodal Projector Tuning for Token-Reduced VLMs

    model/method

    While training-free visual token compression significantly reduces visual sequence length, reducing token count from dense representations (e.g., 576 or 2880 tokens) down to 64 or 160 tokens can cause a representation shift between the reduced visual token space and the LLM embedding space.

    To bridge this alignment gap without high computational cost, VisionZip employs efficient projector fine-tuning (denoted as VisionZip‡\text{VisionZip}^\ddagger):

    • Frozen Components: The pre-trained vision encoder and the large language model (LLM) weights remain entirely frozen.
    • Tuned Component: Only the cross-modality projector (e.g., linear or multi-layer perceptron) is fine-tuned.
    • Data and Training Budget: Fine-tuning requires only 10%10\% of the standard LLaVA-1.5 visual instruction tuning dataset and completes in approximately 30 minutes on 8 NVIDIA A800 GPUs (or standard RTX 3090 GPUs).

    This brief projector adaptation recovers 1%1\% to 3%3\% of benchmark accuracy, bringing performance close to that of vanilla VLMs using 10×10\times more tokens.

  7. Knowl 7 — Limitation of Text-Conditioned Visual Pruning in Multi-Turn Conversations

    limitation

    In multi-turn conversational VLMs, Key-Value (KV) cache states of past dialogue turns are retained to avoid re-encoding history for successive user interactions.

    Text-conditioned token pruning techniques (such as FastV and SparseVLM) dynamically prune visual tokens based on attention weights between the image tokens and the prompt tokens of the current question. Consequently, the visual tokens retained and saved into the KV cache are heavily biased toward the first user turn and lack visual features required to answer subsequent questions regarding different regions or objects in the image. This causes severe degradation and hallucination on subsequent turns.

    In contrast, text-agnostic token reduction methods that identify dominant and contextual tokens purely from the vision encoder representations preserve universal scene semantics in the KV cache across arbitrary conversational turns.

  8. Knowl 8 — Benchmark Performance of VisionZip on LLaVA-1.5 Image Understanding

    data/table

    VisionZip and its projector fine-tuned variant (VisionZip‡\text{VisionZip}^\ddagger) were evaluated on LLaVA-1.5 across 11 multimodal benchmarks at token reduction levels of 192, 128, and 64 tokens from the original 576 visual tokens. The baseline methods FastV and SparseVLM progressively reduce visual tokens during the LLM forward pass.

    Method GQA MMB MME POPE SQA VQAV2^{\text{V2}} VQAText^{\text{Text}} MMMU SEED MMVet LLaVA-B Avg.
    Upper Bound, 576 Tokens (100%)
    Vanilla 61.9 64.7 1862 85.9 69.5 78.5 58.2 36.3 58.6 31.1 66.8 100%
    Retain 192 Tokens (↓66.7%\downarrow 66.7\%)
    FastV 52.7 61.2 1612 64.8 67.3 67.1 52.5 34.3 57.1 27.7 49.4 88.2%
    SparseVLM 57.6 62.5 1721 83.6 69.1 75.6 56.1 33.8 55.8 31.5 66.1 96.4%
    VisionZip 59.3 63.0 1782.6 85.3 68.9 76.8 57.3 36.6 56.4 31.7 67.7 98.5%
    VisionZip‡^\ddagger 60.1 63.4 1834 84.9 68.2 77.4 57.8 36.2 57.1 32.6 66.7 99.1%
    Retain 128 Tokens (↓77.8%\downarrow 77.8\%)
    FastV 49.6 56.1 1490 59.6 60.2 61.8 50.6 34.9 55.9 28.1 52.0 83.5%
    SparseVLM 56.0 60.0 1696 80.5 67.1 73.8 54.9 33.8 53.4 30.0 62.7 93.4%
    VisionZip 57.6 62.0 1761.7 83.2 68.9 75.6 56.8 37.9 54.9 32.6 64.8 97.6%
    VisionZip‡^\ddagger 58.9 62.6 1823 83.7 68.3 76.6 57.0 37.3 55.8 32.9 64.8 98.4%
    Retain 64 Tokens (↓88.9%\downarrow 88.9\%)
    FastV 46.1 48.0 1256 48.0 51.1 55.0 47.8 34.0 51.9 25.8 46.1 75.6%
    SparseVLM 52.7 56.2 1505 75.1 62.2 68.2 51.8 32.7 51.1 23.3 57.5 85.8%
    VisionZip 55.1 60.1 1690 77.0 69.0 72.4 55.5 36.2 52.2 31.7 62.9 94.0%
    VisionZip‡^\ddagger 57.0 61.5 1756 80.9 68.8 74.2 56.0 35.6 53.4 30.2 63.6 95.2%

    With only 64 tokens remaining (an 88.9%88.9\% reduction), training-free VisionZip achieves 94.0%94.0\% of the vanilla model's performance, outperforming SparseVLM by 8.2%8.2\% and FastV by 18.4%18.4\%. Projector fine-tuning (VisionZip‡\text{VisionZip}^\ddagger) further elevates performance to 95.2%95.2\%.

  9. Knowl 9 — Benchmark Performance of VisionZip on High-Resolution LLaVA-NeXT

    data/table

    LLaVA-NeXT encodes a 672×672672 \times 672 image by splitting it into 5 sub-images (4 patches plus 1 resized global image), producing 576×5=2880576 \times 5 = 2880 visual tokens. VisionZip compresses these 2880 tokens down to 640, 320, and 160 tokens without and with projector fine-tuning (VisionZip‡\text{VisionZip}^\ddagger).

    Method GQA MMB MME SQA VQAV2^{\text{V2}} VQAText^{\text{Text}} MMMU Avg.
    Upper Bound, 2880 Tokens (100%)
    Vanilla 64.2 67.9 1842 70.2 80.1 61.3 35.1 100%
    Retain 640 Tokens (↓77.8%\downarrow 77.8\%)
    SparseVLM 60.3 65.7 1772 67.7 77.1 57.8 34.6 96.1%
    VisionZip 61.3 66.3 1787 68.1 79.1 60.2 34.7 97.6%
    VisionZip‡^\ddagger 62.4 65.9 1778 67.9 79.9 60.8 37.2 98.9%
    Retain 320 Tokens (↓88.9%\downarrow 88.9\%)
    SparseVLM 57.7 64.3 1694 67.3 73.4 55.9 34.4 93.3%
    VisionZip 59.3 63.1 1702 67.3 76.2 58.9 35.3 95.0%
    VisionZip‡^\ddagger 61.0 64.4 1770 67.5 78.4 59.3 38.0 97.9%
    Retain 160 Tokens (↓94.4%\downarrow 94.4\%)
    SparseVLM 51.2 63.1 1542 67.5 66.3 46.4 32.8 86.4%
    VisionZip 55.5 60.1 1630 68.3 71.4 56.2 36.1 92.0%
    VisionZip‡^\ddagger 58.2 63.9 1699 67.5 75.6 57.3 37.7 95.5%

    When compressing the visual token budget by 94.4%94.4\% (from 2880 down to 160 tokens), training-free VisionZip achieves 92.0%92.0\% of the vanilla model accuracy compared to SparseVLM's 86.4%86.4\%. With projector tuning, VisionZip‡\text{VisionZip}^\ddagger maintains 95.5%95.5\% relative accuracy.

  10. Knowl 10 — Performance of VisionZip on Video-LLaVA Video QA Benchmarks

    data/table

    Video-LLaVA encodes 8 video frames with LanguageBind, producing 8×256=20488 \times 256 = 2048 visual tokens. VisionZip compresses tokens within each frame from 256 to 17 tokens, yielding a total of 136 video tokens for the sequence (compared to 135 tokens in baseline pruning methods).

    Method TGIF-QA MSVD-QA MSRVTT-QA ActivityNet-QA Avg.
    Vanilla (2048 Tokens) 47.1 (100.0%) 69.8 (100.0%) 56.7 (100.0%) 43.1 (100.0%) 100.0%
    FastV (135 Tokens) 23.1 (49.0%) 38.0 (54.4%) 19.3 (34.0%) 30.6 (71.0%) 52.1%
    SparseVLM (135 Tokens) 44.7 (94.9%) 68.2 (97.7%) 31.0 (54.7%) 42.6 (98.8%) 86.5%
    VisionZip (136 Tokens) 42.4 (90.0%) 63.5 (91.0%) 52.1 (91.9%) 43.0 (99.8%) 93.2%

    VisionZip achieves an overall average accuracy of 93.2%93.2\%, outperforming SparseVLM (86.5%86.5\%) and FastV (52.1%52.1\%). On the MSRVTT-QA dataset, VisionZip scores 52.152.1 (91.9%91.9\% of vanilla), whereas SparseVLM drops to 31.031.0 (54.7%54.7\% of vanilla).

  11. Knowl 11 — Inference Efficiency and Prefilling Latency on LLaVA-NeXT 7B

    data/table

    Inference efficiency was evaluated on a single NVIDIA A800-80GB GPU using the POPE benchmark with LLaVA-NeXT 7B, measuring both total execution time and first-token prefilling latency.

    Method Visual Tokens Total Time ↓\downarrow Speedup Δ\Delta Prefilling Latency ↓\downarrow Reduction Ratio Δ\Delta
    Baseline (Vanilla) 2880 2293s – 218ms –
    FastV 160 1792s 1.3×\times 119ms 1.8×\times
    SparseVLM 160 1895s 1.2×\times 128ms 1.7×\times
    VisionZip 160 756s 3.0×\times 27.8ms 7.8×\times

    Because VisionZip prunes tokens within the vision encoder before entering the LLM, it achieves a 7.8×7.8\times reduction in prefilling latency (down to 27.8 ms27.8\text{ ms} from 218 ms218\text{ ms}) and a 3.0×3.0\times reduction in overall inference time, compared to 1.2×1.2\times--1.3×1.3\times total speedup for text-conditioned LLM pruning methods.

Coverage note — Omitted qualitative visualization figures, related work literature summaries, and the Mini-Gemini curve plot values (Figure 4) which replicate the core empirical findings already fully detailed in the LLaVA-1.5, LLaVA-NeXT, and Video-LLaVA tables.

References

  1. 1.Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. Gpt-4 technical report. arXiv:2303.08774, 2023. 1, 8, 20
  2. 2.Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, et al. Qwen technical report. arXiv:2309.16609, 2023. 1, 8, 20
  3. 3.Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou. Qwen-VL: A frontier large vision-language model with versatile abilities. arXiv:2308.12966, 2023. 1
  4. 4.Fei-Long Chen, Du-Zhen Zhang, Ming-Lun Han, Xiu-Yi Chen, Jing Shi, Shuang Xu, and Bo Xu. Vlp: A survey on vision-language pre-training. Machine Intelligence Research, 20(1):38–56, 2023. 1
  5. 5.Lin Chen, Jisong Li, Xiaoyi Dong, Pan Zhang, Conghui He, Jiaqi Wang, Feng Zhao, and Dahua Lin. Sharegpt4v: Improving large multi-modal models with better captions. arXiv:2311.12793, 2023. 1
  6. 6.Liang Chen, Haozhe Zhao, Tianyu Liu, Shuai Bai, Junyang Lin, Chang Zhou, and Baobao Chang. An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models, 2024. 4, 6, 8, 12, 20
  7. 7.Yukang Chen, Shengju Qian, Haotian Tang, Xin Lai, Zhijian Liu, Song Han, and Jiaya Jia. Longlora: Efficient fine-tuning of long-context large language models. In ICLR, 2024. 20
  8. 8.Zhe Chen, Jiannan Wu, Wenhai Wang, Weijie Su, Guo Chen, Sen Xing, Muyan Zhong, Qinglong Zhang, Xizhou Zhu, Lewei Lu, Bin Li, Ping Luo, Tong Lu, Yu Qiao, and Jifeng Dai. InternVL: Scaling up vision foundation models and aligning for generic visual-linguistic tasks. arXiv:2312.14238, 2023. 8
  9. 9.Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E Gonzalez, et al. Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality. See https://vicuna. lmsys. org (accessed 14 April 2023), 2(3):6, 2023. 20
  10. 10.Alexey Dosovitskiy. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929, 2020. 1
  11. 11.Chaoyou Fu, Peixian Chen, Yunhang Shen, Yulei Qin, Mengdan Zhang, Xu Lin, Jinrui Yang, Xiawu Zheng, Ke Li, Xing Sun, et al. MME: A comprehensive evaluation benchmark for multimodal large language models. arXiv:2306.13394, 2023. 4, 15
  12. 12.Suyu Ge, Yunan Zhang, Liyuan Liu, Minjia Zhang, Jiawei Han, and Jianfeng Gao. Model tells you what to discard: Adaptive kv cache compression for llms. arXiv preprint arXiv:2310.01801, 2023. 20
  13. 13.Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh. Making the v in vqa matter: Elevating the role of image understanding in visual question answering. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 6904–6913, 2017. 4, 15
  14. 14.Danna Gurari, Qing Li, Abigale J Stangl, Anhong Guo, Chi Lin, Kristen Grauman, Jiebo Luo, and Jeffrey P Bigham. Vizwiz grand challenge: Answering visual questions from blind people. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 3608–3617, 2018. 14
  15. 15.Chi Han, Qifan Wang, Wenhan Xiong, Yu Chen, Heng Ji, and Sinong Wang. Lm-infinite: Simple on-the-fly length generalization for large language models. arXiv preprint arXiv:2308.16137, 2023. 20
  16. 16.Yefei He, Feng Chen, Jing Liu, Wenqi Shao, Hong Zhou, Kaipeng Zhang, and Bohan Zhuang. Zipvl: Efficient large vision-language models with dynamic token sparsification and kv cache compression. arXiv preprint arXiv:2410.08584, 2024. 12, 20
  17. 17.De-An Huang, Shijia Liao, Subhashree Radhakrishnan, Hongxu Yin, Pavlo Molchanov, Zhiding Yu, and Jan Kautz. Lita: Language instructed temporal-localization assistant. In European Conference on Computer Vision, pages 202–218. Springer, 2025. 20
  18. 18.Zhengchao Huang, Bin Xia, Zicheng Lin, Zhun Mou, and Wenming Yang. Ffaa: Multimodal large language model based explainable open-world face forgery analysis assistant. arXiv preprint arXiv:2408.10072, 2024. 20
  19. 19.Drew A Hudson and Christopher D Manning. GQA: A new dataset for real-world visual reasoning and compositional question answering. Conference on Computer Vision and Pattern Recognition (CVPR), 2019. 4, 15
  20. 20.Yunseok Jang, Yale Song, Youngjae Yu, Youngjin Kim, and Gunhee Kim. Tgif-qa: Toward spatio-temporal reasoning in visual question answering. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2758–2766, 2017. 6, 18
  21. 21.Yiren Jian, Tingkai Liu, Yunzhe Tao, Chunhui Zhang, Soroush Vosoughi, and Hongxia Yang. Expedited training of visual conditioned language generation via redundancy reduction. arXiv preprint arXiv:2310.03291, 2023. 20
  22. 22.Jacob Devlin Ming-Wei Chang Kenton and Lee Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of naacL-HLT, page 2. Minneapolis, Minnesota, 2019. 1
  23. 23.Moo Jin Kim, Karl Pertsch, Siddharth Karamcheti, Ted Xiao, Ashwin Balakrishna, Suraj Nair, Rafael Rafailov, Ethan Foster, Grace Lam, Pannag Sanketi, Quan Vuong, Thomas Kollar, Benjamin Burchfiel, Russ Tedrake, Dorsa Sadigh, Sergey Levine, Percy Liang, and Chelsea Finn. Openvla: An open-source vision-language-action model. arXiv preprint arXiv:2406.09246, 2024. 1
  24. 24.Xin Lai, Zhuotao Tian, Yukang Chen, Yanwei Li, Yuhui Yuan, Shu Liu, and Jiaya Jia. Lisa: Reasoning segmentation via large language model. arXiv preprint arXiv:2308.00692, 2023. 20
  25. 25.Xin Lai, Zhuotao Tian, Yukang Chen, Senqiao Yang, Xiangru Peng, and Jiaya Jia. Step-dpo: Step-wise preference optimization for long-chain reasoning of llms. arXiv:2406.18629, 2024. 20
  26. 26.Bohao Li, Rui Wang, Guangzhi Wang, Yuying Ge, Yixiao Ge, and Ying Shan. Seed-bench: Benchmarking multimodal llms with generative comprehension. arXiv preprint arXiv:2307.16125, 2023. 4, 13
  27. 27.Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In International conference on machine learning, 2023. 1
  28. 28.Jingyao Li, Han Shi, Xin Jiang, Zhenguo Li, Hong Xu, and Jiaya Jia. Quickllama: Query-aware inference acceleration for large language models, 2024. 20
  29. 29.Yifan Li, Yifan Du, Kun Zhou, Jinpeng Wang, Wayne Xin Zhao, and Ji-Rong Wen. Evaluating object hallucination in large vision-language models. arXiv:2305.10355, 2023. 4, 6, 15
  30. 30.Yanwei Li, Chengyao Wang, and Jiaya Jia. Llama-vid: An image is worth 2 tokens in large language models. 2024. 20
  31. 31.Yanwei Li, Yuechen Zhang, Chengyao Wang, Zhisheng Zhong, Yixin Chen, Ruihang Chu, Shaoteng Liu, and Jiaya Jia. Mini-gemini: Mining the potential of multi-modality vision language models. arXiv:2403.18814, 2024. 1, 4, 8, 13, 20
  32. 32.Bin Lin, Bin Zhu, Yang Ye, Munan Ning, Peng Jin, and Li Yuan. Video-llava: Learning united visual representation by alignment before projection. arXiv:2311.10122, 2023. 6, 8, 20
  33. 33.Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee. Improved baselines with visual instruction tuning. arXiv:2310.03744, 2023. 1, 4, 8, 13, 14, 16, 20
  34. 34.Haotian Liu, Chunyuan Li, Yuheng Li, Bo Li, Yuanhan Zhang, Sheng Shen, and Yong Jae Lee. Llava-next: Improved reasoning, ocr, and world knowledge, 2024. 1, 4, 13, 20
  35. 35.Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual instruction tuning. Advances in neural information processing systems, 2024. 20
  36. 36.Jiaming Liu, Mengzhen Liu, Zhenyu Wang, Lily Lee, Kaichen Zhou, Pengju An, Senqiao Yang, Renrui Zhang, Yandong Guo, and Shanghang Zhang. Robomamba: Multimodal state space model for efficient robot reasoning and manipulation. arXiv preprint arXiv:2406.04339, 2024. 1
  37. 37.Yuan Liu, Haodong Duan, Yuanhan Zhang, Bo Li, Songyang Zhang, Wangbo Zhao, Yike Yuan, Jiaqi Wang, Conghui He, Ziwei Liu, et al. MMBench: Is your multi-modal model an all-around player? arXiv:2307.06281, 2023. 4, 14
  38. 38.Zhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer, Trevor Darrell, and Saining Xie. A convnet for the 2020s. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 11976–11986, 2022. 6
  39. 39.Pan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu, Kai-Wei Chang, Song-Chun Zhu, Oyvind Tafjord, Peter Clark, and Ashwin Kalyan. Learn to explain: Multimodal reasoning via thought chains for science question answering. Advances in Neural Information Processing Systems, 35:2507–2521, 2022. 4, 14
  40. 40.Muhammad Maaz, Hanoona Rasheed, Salman Khan, and Fahad Shahbaz Khan. Video-chatgpt: Towards detailed video understanding via large vision and language models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (ACL 2024), 2024. 8, 20
  41. 41.Guanqiao Qu, Qiyuan Chen, Wei Wei, Zheng Lin, Xianhao Chen, and Kaibin Huang. Mobile edge intelligence for large language models: A contemporary survey. arXiv preprint arXiv:2407.18921, 2024. 1
  42. 42.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In International conference on machine learning, pages 8748–8763. PMLR, 2021. 1, 12
  43. 43.Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn. Direct preference optimization: Your language model is secretly a reward model. Advances in Neural Information Processing Systems, 36, 2024. 20
  44. 44.Yuzhang Shang, Mu Cai, Bingxin Xu, Yong Jae Lee, and Yan Yan. Llava-prumerge: Adaptive token reduction for efficient large multimodal models. arXiv preprint arXiv:2403.15388, 2024. 20
  45. 45.Tong Shao, Zhuotao Tian, Hang Zhao, and Jingyong Su. Explore the potential of clip for training-free open vocabulary semantic segmentation. In European Conference on Computer Vision, pages 139–156. Springer, 2025. 7
  46. 46.Dachuan Shi, Chaofan Tao, Ying Jin, Zhendong Yang, Chun Yuan, and Jiaqi Wang. Upop: Unified and progressive pruning for compressing vision-language transformers. In International Conference on Machine Learning, pages 31292–31311. PMLR, 2023. 20
  47. 47.Amanpreet Singh, Vivek Natarjan, Meet Shah, Yu Jiang, Xinlei Chen, Devi Parikh, and Marcus Rohrbach. Towards VQA models that can read. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 8317–8326, 2019. 4, 15
  48. 48.Enxin Song, Wenhao Chai, Guanhong Wang, Yucheng Zhang, Haoyang Zhou, Feiyang Wu, Haozhe Chi, Xun Guo, Tian Ye, Yanting Zhang, et al. Moviechat: From dense token to sparse memory for long video understanding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 18221–18232, 2024. 20
  49. 49.Yunlong Tang, Jing Bi, Siting Xu, Luchuan Song, Susan Liang, Teng Wang, Daoan Zhang, Jie An, Jingyang Lin, Rongyi Zhu, et al. Video understanding with large language models: A survey. arXiv preprint arXiv:2312.17432, 2023. 20
  50. 50.Shengbang Tong, Ellis Brown, Penghao Wu, Sanghyun Woo, Manoj Middepogu, Sai Charitha Akula, Jihan Yang, Shusheng Yang, Adithya Iyer, Xichen Pan, et al. Cambrian-1: A fully open, vision-centric exploration of multimodal llms. arXiv preprint arXiv:2406.16860, 2024. 8, 20
  51. 51.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothee Lacroix, Baptiste Roziere, Naman Goyal, Eric Hambro, Faisal Azhar, et al. Llama: Open and efficient foundation language models. arXiv:2302.13971, 2023. 1, 8, 20
  52. 52.Weihan Wang, Qingsong Lv, Wenmeng Yu, Wenyi Hong, Ji Qi, Yan Wang, Junhui Ji, Zhuoyi Yang, Lei Zhao, Xixuan Song, et al. Cogvlm: Visual expert for pretrained language models. arXiv:2311.03079, 2023. 20
  53. 53.Yuxin Wen, Qingqing Cao, Qichen Fu, Sachin Mehta, and Mahyar Najibi. Efficient vision-language models by summarizing visual tokens into compact registers. arXiv preprint arXiv:2410.14072, 2024. 20
  54. 54.Guangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han, and Mike Lewis. Efficient streaming language models with attention sinks. arXiv, 2023. 7, 20
  55. 55.Long Xing, Qidong Huang, Xiaoyi Dong, Jiajie Lu, Pan Zhang, Yuhang Zang, Yuhang Cao, Conghui He, Jiaqi Wang, Feng Wu, et al. Pyramiddrop: Accelerating your large vision-language models via pyramid visual redundancy reduction. arXiv preprint arXiv:2410.17247, 2024. 12, 20
  56. 56.Dejing Xu, Zhou Zhao, Jun Xiao, Fei Wu, Hanwang Zhang, Xiangnan He, and Yueting Zhuang. Video question answering via gradually refined attention over appearance and motion. In Proceedings of the ACM international conference on Multimedia, pages 1645–1653, 2017. 6, 18
  57. 57.Senqiao Yang, Jiaming Liu, Ray Zhang, Mingjie Pan, Zoey Guo, Xiaoqi Li, Zehui Chen, Peng Gao, Yandong Guo, and Shanghang Zhang. Lidar-llm: Exploring the potential of large language models for 3d lidar understanding. arXiv preprint arXiv:2312.14074, 2023. 1
  58. 58.Senqiao Yang, Tianyuan Qu, Xin Lai, Zhuotao Tian, Bohao Peng, Shu Liu, and Jiaya Jia. An improved baseline for reasoning segmentation with large language model. arXiv preprint arXiv:2312.17240, 2023. 20
  59. 59.Senqiao Yang, Zhuotao Tian, Li Jiang, and Jiaya Jia. Unified language-driven zero-shot domain adaptation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 23407–23415, 2024. 1
  60. 60.Yuan Yao, Tianyu Yu, Ao Zhang, Chongyi Wang, Junbo Cui, Hongji Zhu, Tianchi Cai, Haoyu Li, Weilin Zhao, Zhihui He, et al. Minicpm-v: A gpt-4v level mllm on your phone. arXiv preprint arXiv:2408.01800, 2024. 1
  61. 61.Weihao Yu, Zhengyuan Yang, Linjie Li, Jianfeng Wang, Kevin Lin, Zicheng Liu, Xinchao Wang, and Lijuan Wang. Mm-vet: Evaluating large multimodal models for integrated capabilities. In International conference on machine learning. PMLR, 2024. 4, 14
  62. 62.Zhou Yu, Dejing Xu, Jun Yu, Ting Yu, Zhou Zhao, Yueting Zhuang, and Dacheng Tao. Activitynet-qa: A dataset for understanding complex web videos via question answering. In AAAI, pages 9127–9134, 2019. 6, 18
  63. 63.Xiang Yue, Yuansheng Ni, Kai Zhang, Tianyu Zheng, Ruoqi Liu, Ge Zhang, Samuel Stevens, Dongfu Jiang, Weiming Ren, Yuxuan Sun, Cong Wei, Botao Yu, Ruibin Yuan, Renliang Sun, Ming Yin, Boyuan Zheng, Zhenzhu Yang, Yibo Liu, Wenhao Huang, Huan Sun, Yu Su, and Wenhu Chen. Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi. In Proceedings of CVPR, 2024. 4, 13
  64. 64.Xiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, and Lucas Beyer. Sigmoid loss for language image pre-training. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 11975–11986, 2023. 1
  65. 65.Kaichen Zhang, Bo Li, Peiyuan Zhang, Fanyi Pu, Joshua Adrian Cahyono, Kairui Hu, Shuai Liu, Yuanhan Zhang, Jingkang Yang, Chunyuan Li, et al. Lmms-eval: Reality check on the evaluation of large multimodal models. arXiv preprint arXiv:2407.12772, 2024. 14, 15, 16, 17, 19, 20
  66. 66.Yuan Zhang, Chun-Kai Fan, Junpeng Ma, Wenzhao Zheng, Tao Huang, Kuan Cheng, Denis Gudovskiy, Tomoyuki Okuno, Yohei Nakata, Kurt Keutzer, et al. Sparsevlm: Visual token sparsification for efficient vision-language model inference. arXiv preprint arXiv:2410.04417, 2024. 4, 6, 8, 12, 20, 21
  67. 67.Yuechen Zhang, Shengju Qian, Bohao Peng, Shu Liu, and Jiaya Jia. Prompt highlighter: Interactive control for multimodal llms. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 13215–13224, 2024. 20
  68. 68.Zhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen, Lianmin Zheng, Ruisi Cai, Zhao Song, Yuandong Tian, Christopher Re, Clark Barrett, et al. H2o: Heavy-hitter oracle for efficient generative inference of large language models. Advances in Neural Information Processing Systems, 36: 34661–34710, 2023. 20
  69. 69.Chuanyang Zheng, Yihang Gao, Han Shi, Minbin Huang, Jingyao Li, Jing Xiong, Xiaozhe Ren, Michael Ng, Xin Jiang, Zhenguo Li, and Yu Li. Dape: Data-adaptive positional encoding for length extrapolation, 2024. 20
  70. 70.Qiang Zhou, Chaohui Yu, Shaofeng Zhang, Sitong Wu, Zhibing Wang, and Fan Wang. Regionblip: A unified multimodal pre-training framework for holistic and regional comprehension. arXiv preprint arXiv:2308.02299, 2023. 20
  71. 71.Bin Zhu, Bin Lin, Munan Ning, Yang Yan, Jiaxi Cui, HongFa Wang, Yatian Pang, Wenhao Jiang, Junwu Zhang, Zongwei Li, et al. Languagebind: Extending video-language pretraining to n-modality by language-based semantic alignment. arXiv preprint arXiv:2310.01852, 2023. 12
  72. 72.Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny. Minigpt-4: Enhancing vision-language understanding with advanced large language models. arXiv:2304.10592, 2023. 1

Citation

MLA
Yang, S., et al. “VisionZip: Longer Is Better but Not Necessary in Vision Language Models”. 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2025, pp. 19792–802, https://doi.org/10.1109/CVPR52734.2025.01843.
APA
Yang, S., Chen, Y., Tian, Z., Wang, C., Li, J., Yu, B., & Jia, J. (2025). VisionZip: Longer is Better but Not Necessary in Vision Language Models. 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 19792–19802. https://doi.org/10.1109/CVPR52734.2025.01843
Chicago
Yang, S., Y. Chen, Z. Tian, et al. 2025. “VisionZip: Longer Is Better but Not Necessary in Vision Language Models”. 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 19792–802. https://doi.org/10.1109/CVPR52734.2025.01843.
Harvard
Yang, S. et al. (2025) “VisionZip: Longer is Better but Not Necessary in Vision Language Models”, 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp. 19792–19802. Available at: https://doi.org/10.1109/CVPR52734.2025.01843.
Vancouver
1. Yang S, Chen Y, Tian Z, Wang C, Li J, Yu B, Jia J (2025) VisionZip: Longer is Better but Not Necessary in Vision Language Models. In: 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp 19792–19802

BibTeX

@inproceedings{Yang_2025, title={VisionZip: Longer is Better but Not Necessary in Vision Language Models}, url={http://dx.doi.org/10.1109/CVPR52734.2025.01843}, DOI={10.1109/cvpr52734.2025.01843}, booktitle={2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)}, publisher={IEEE}, author={Yang, Senqiao and Chen, Yukang and Tian, Zhuotao and Wang, Chengyao and Li, Jingyao and Yu, Bei and Jia, Jiaya}, year={2025}, month=June, pages={19792–19802} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE