VisionZip: Longer is Better but Not Necessary in Vision Language Models
Senqiao YangYukang ChenZhuotao TianChengyao WangJingyao LiBei YuJiaya Jia
Proposes VisionZip, a visual token reduction method that selects dominant tokens and merges redundant ones from vision encoders to achieve an 8x prefilling speedup with minimal performance loss across image and video benchmarks.
Modern vision-language models achieve impressive visual and reasoning performance, but they do so at steep computational and memory costs. These architectures convert high-resolution images and video frames into thousands of visual tokens—often vastly outnumbering text tokens by a factor of twenty or more. This creates severe inference latency and resource bottlenecks that restrict deployment in latency-sensitive, resource-constrained environments such as robotics, edge devices, and autonomous systems.
The article evaluates whether all generated visual tokens are truly necessary, demonstrating that standard vision encoders suffer from extreme feature redundancy. It introduces VisionZip, a text-agnostic framework designed to aggressively reduce visual token counts while preserving model performance across diverse multimodal benchmarks.
To demonstrate this, the researchers analyzed token attention patterns within common vision encoders and evaluated VisionZip across eleven image understanding benchmarks and four video question-answering benchmarks. The approach was tested across leading open-source model families—including LLaVA-1.5, LLaVA-NeXT, and Mini-Gemini—under both training-free inference settings and an efficient fine-tuning mode that adapts only the cross-modality projector using a tiny subset of data.
The findings show that attention within vision encoders heavily concentrates into a small subset of dominant proxy tokens, leaving most visual tokens with near-zero attention. VisionZip capitalizes on this by selecting dominant tokens based on attention weights and merging remaining non-dominant tokens based on semantic similarity. Across all evaluated settings, VisionZip retained approximately 95% of original model performance using only 5% to 10% of the visual tokens. On the POPE benchmark with LLaVA-NeXT 7B, it reduced initial token prefilling time by 7.8 times and improved overall inference runtime threefold. Furthermore, because VisionZip selects tokens independently of input text prompts, it outperformed text-dependent pruning baselines by over 5% and maintained robust performance across multi-turn conversational caching without discarding context needed for follow-up questions.
These results demonstrate that increasing visual token length delivers diminishing returns while dramatically driving up operational compute costs. By extracting concentrated visual representations at the encoder stage, organizations can deploy larger, more capable language models—such as a 13-billion parameter model—at faster speeds and lower hardware costs than standard smaller models. Prior text-guided token pruning methods suffer from spatial feature misalignment, whereas prompt-agnostic token reduction provides a more reliable path for production-ready vision-language systems.
Engineering and deployment teams should consider adopting plug-and-play visual token reduction techniques like VisionZip to improve inference throughput and reduce serving infrastructure costs. For teams seeking maximal fidelity, performing a lightweight, 30-minute projector fine-tuning step is recommended to bridge the space between reduced vision tokens and language decoders. Future research and development should focus on training natively dense vision encoders that eliminate feature redundancy upstream rather than relying on oversized token sequences.
While confidence in the empirical results is high across standard vision-language benchmarks and common vision architectures, the evaluations rely on existing benchmark suites and standard self-attention encoders. Stakeholders should conduct pilot validations within their specific domain-specific tasks and production data streams to verify that aggressive token reduction maintains sufficient fine-grained visual accuracy.
- Paper: Prismatic VLMs: Investigating the Design Space of Visually-Conditioned Language Models, Siddharth Karamcheti et al. (2024). It provides a rigorous empirical analysis of visual representations and projector design spaces in vision-language models, establishing foundational architectural insights that VisionZip optimizes.
- Paper: Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution, Peng Wang et al. (2024). It establishes key dynamic-resolution vision-language model architectures and shows the explosive growth of visual tokens that VisionZip directly aims to compress.
- Paper: LLaVA-OneVision: Easy Visual Task Transfer, Bo Li et al. (2024). It introduces standard unified visual representation frameworks and token budget baselines across image and video tasks that motivate prompt-agnostic token reduction methods.
- Paper: Vision-Language Models for Vision Tasks: A Survey, Jingyi Zhang et al. (2023). It presents a comprehensive taxonomy and architectural foundation of vision-language models and cross-modality projectors essential for contextualizing token compression techniques.
- Paper: Video-LLaVA: Learning United Visual Representation by Alignment Before Projection, Bin Lin et al. (2023). It details visual-language alignment and projection architectures across image and video modalities that VisionZip evaluates and accelerates.
- Paper: Tokens-to-Token ViT: Training Vision Transformers from Scratch on ImageNet, Li Yuan et al. (2021). It introduces fundamental principles of progressive visual token merging and feature redundancy reduction inside vision transformers.
- Paper: Global Compression Commander: Plug-and-Play Inference Acceleration for High-Resolution Large Vision-Language Models, Xuyang Liu et al. (2026). It extends prompt-agnostic token reduction concepts to hierarchical high-resolution models by using global thumbnails to dynamically compress local crop tokens without retraining.
- Paper: Filter, Correlate, Compress: Training-Free Token Reduction for MLLM Acceleration, Yuhang Han et al. (2026). It advances training-free visual token reduction by combining redundancy filtering with inter-token correlation routing across both vision encoders and language decoders.
- Paper: XAttention: Block Sparse Attention with Antidiagonal Scoring, Ruyi Xu et al. (2025). It complements visual token pruning by addressing attention sparsity along long-context video and multimodal sequence processing.
- Paper: M-LLM Based Video Frame Selection for Efficient Video Understanding, Kai Hu 0010 et al. (2025). It builds on visual token reduction strategies by introducing question-aware video frame selection using compact tokenized frame proxies.
- Paper: Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation, Zhiheng Liu et al. (2026). It explores an alternative paradigm that circumvents vision encoder token redundancy entirely by feeding pixel embeddings directly into multimodal models.
