PuMer: Pruning and Merging Tokens for Efficient Vision Language Models
Qingqing CaoBhargavi ParanjapeHannaneh Hajishirzi
Introduces a token reduction framework combining text-guided pruning and modality-aware merging that doubles vision-language model inference throughput and cuts memory consumption in half with under a 1% loss in accuracy.
Modern vision-language models have achieved remarkable success across complex visual reasoning tasks, but their high computational cost and heavy memory footprint present major operational challenges. Because these architectures compute dense cross-attention between fine-grained visual image patches and text tokens across multiple deep layers, resource demands scale quadratically. This inefficiency severely restricts high-throughput cloud deployments and makes on-device execution on resource-constrained hardware virtually impractical.
The article introduces and evaluates PuMer, a lightweight token reduction framework designed to substantially accelerate vision-language model processing without sacrificing task accuracy. The framework demonstrates that progressively removing text-irrelevant visual information and merging redundant tokens directly inside cross-modal layers yields dramatic efficiency gains.
The authors designed a two-stage, non-parametric token reduction mechanism inserted at multiple layers of a vision-language model. First, text-informed image pruning removes visual tokens that show low cross-attention relevance to the input sentence, avoiding the need for extra learnable parameters. Second, modality-aware merging uses a fast bipartite matching algorithm to combine semantically similar tokens separately within each modality. To evaluate this approach, extensive computational experiments were performed across two distinct vision-language architectures—the 110-million parameter ViLT and the 330-million parameter METER—across five standard visual reasoning benchmarks, measuring real hardware throughput, peak memory usage, and task accuracy against baseline reduction methods.
The empirical findings demonstrate that PuMer effectively resolves key efficiency bottlenecks. Across all benchmark tasks, PuMer improves model inference throughput by 1.7x to 2.1x while slashing peak memory consumption by 38% to 51%. Crucially, these operational speedups incur less than a 1% drop in task accuracy compared to standard fully finetuned models. Furthermore, PuMer consistently outperforms existing vision-only token reduction baselines (such as DynamicViT and ToMe) and simple image downsampling, providing superior accuracy at equivalent throughput levels. Ablation analysis confirms that combining text-informed pruning with modality-aware merging is essential, as pruning alone causes severe information loss while merging alone leaves substantial cross-modal redundancy.
These results provide a practical path toward reducing inference hosting costs, mitigating memory risks in production environments, and enabling faster response times in interactive applications. Unlike single-modality pruning methods that discard data statically, text-guided reduction ensures that task-critical visual features are retained depending on the specific user query. Because the token reduction modules require no additional model parameters and reduce computation during forward passes, the framework also yields 15% to 20% faster training times.
Organizations deploying large-scale vision-language models should consider integrating cascaded token reduction into their serving pipelines to optimize infrastructure utilization. When implementing the framework, teams should carefully balance the reduction layer depth and compression ratios, as scattering reduction across middle layers preserves higher accuracy than aggressive early pruning. However, caution is warranted when applying this framework to architectures where the standalone image encoder accounts for the majority of the computational workload rather than the cross-modal encoder; in such settings, additional token reduction techniques within the vision backbone will be required to realize similar speedups.
- Paper: An Empirical Study of Training End-to-End Vision-and-Language Transformers, Zi-Yi Dou et al. (2022). METER is one of PuMer’s evaluated backbones, so its account of the architecture and cross-modal fusion provides direct context for the efficiency results.
- Paper: ViLT: Vision-and-Language Transformer Without Convolution or Region Supervision, Wonjae Kim et al. (2021). ViLT is PuMer’s other evaluated backbone, and understanding its patch-token vision-language transformer makes the pruning and merging interventions easier to follow.
No sufficiently relevant recommendations were found.
