PuMer: Pruning and Merging Tokens for Efficient Vision Language Models

Qingqing CaoBhargavi ParanjapeHannaneh Hajishirzi

article2023ACL60 citations

Introduces a token reduction framework combining text-guided pruning and modality-aware merging that doubles vision-language model inference throughput and cuts memory consumption in half with under a 1% loss in accuracy.

Listen

Modern vision-language models have achieved remarkable success across complex visual reasoning tasks, but their high computational cost and heavy memory footprint present major operational challenges. Because these architectures compute dense cross-attention between fine-grained visual image patches and text tokens across multiple deep layers, resource demands scale quadratically. This inefficiency severely restricts high-throughput cloud deployments and makes on-device execution on resource-constrained hardware virtually impractical.

The article introduces and evaluates PuMer, a lightweight token reduction framework designed to substantially accelerate vision-language model processing without sacrificing task accuracy. The framework demonstrates that progressively removing text-irrelevant visual information and merging redundant tokens directly inside cross-modal layers yields dramatic efficiency gains.

The authors designed a two-stage, non-parametric token reduction mechanism inserted at multiple layers of a vision-language model. First, text-informed image pruning removes visual tokens that show low cross-attention relevance to the input sentence, avoiding the need for extra learnable parameters. Second, modality-aware merging uses a fast bipartite matching algorithm to combine semantically similar tokens separately within each modality. To evaluate this approach, extensive computational experiments were performed across two distinct vision-language architectures—the 110-million parameter ViLT and the 330-million parameter METER—across five standard visual reasoning benchmarks, measuring real hardware throughput, peak memory usage, and task accuracy against baseline reduction methods.

The empirical findings demonstrate that PuMer effectively resolves key efficiency bottlenecks. Across all benchmark tasks, PuMer improves model inference throughput by 1.7x to 2.1x while slashing peak memory consumption by 38% to 51%. Crucially, these operational speedups incur less than a 1% drop in task accuracy compared to standard fully finetuned models. Furthermore, PuMer consistently outperforms existing vision-only token reduction baselines (such as DynamicViT and ToMe) and simple image downsampling, providing superior accuracy at equivalent throughput levels. Ablation analysis confirms that combining text-informed pruning with modality-aware merging is essential, as pruning alone causes severe information loss while merging alone leaves substantial cross-modal redundancy.

These results provide a practical path toward reducing inference hosting costs, mitigating memory risks in production environments, and enabling faster response times in interactive applications. Unlike single-modality pruning methods that discard data statically, text-guided reduction ensures that task-critical visual features are retained depending on the specific user query. Because the token reduction modules require no additional model parameters and reduce computation during forward passes, the framework also yields 15% to 20% faster training times.

Organizations deploying large-scale vision-language models should consider integrating cascaded token reduction into their serving pipelines to optimize infrastructure utilization. When implementing the framework, teams should carefully balance the reduction layer depth and compression ratios, as scattering reduction across middle layers preserves higher accuracy than aggressive early pruning. However, caution is warranted when applying this framework to architectures where the standalone image encoder accounts for the majority of the computational workload rather than the cross-modal encoder; in such settings, additional token reduction techniques within the vision backbone will be required to realize similar speedups.

No sufficiently relevant recommendations were found.

Cover for PuMer: Pruning and Merging Tokens for Efficient Vision Language Models

Abstract

Large-scale vision language (VL) models use Transformers to perform cross-modal interactions between the input text and image. These cross-modal interactions are computationally expensive and memory-intensive due to the quadratic complexity of processing the input image and text. We present PuMer¹: a token reduction framework that uses text-informed Pruning and modality-aware Merging strategies to progressively reduce the tokens of input image and text, improving model inference speed and reducing memory footprint. PuMer learns to keep salient image tokens related to the input text and merges similar textual and visual tokens by adding lightweight token reducer modules at several cross-modal layers in the VL model. Training PuMer is mostly the same as finetuning the original VL model but faster. Our evaluation for two vision language models on four downstream VL tasks shows PuMer increases inference throughput by up to 2x and reduces memory footprint by over 50% while incurring less than a 1% accuracy drop.²

Table of Contents

  • 1 Introduction
  • 2 Related work
  • 3 Background and Overview
  • 4 PuMer: Text-Informed Token Reduction Framework
  • 4.1 Text-Informed Image Pruning
  • 4.2 Modality-Aware Merging
  • 4.3 Training and Inference
  • 5 Evaluation Setup
  • 5.1 Backbone Vision-Language Models
  • 5.2 Evaluation Tasks
  • 5.3 Baselines
  • 5.4 Evaluation Metrics
  • 6 Experimental Results
  • 6.1 Main Results
  • 6.2 Ablation Study
  • 7 Conclusion
  • Acknowledgements
  • 8 Limitations
  • References
  • A Appendix
  • A.1 PuMer Details
  • A.2 Model Inference FLOPs Comparison

Knowls

  1. Knowl 1 — PuMer reduces tokens progressively inside cross-modal encoders

    model/method

    PuMer is a token-reduction framework for vision-language (VL) models whose cross-modal encoder processes text and image tokens together. At selected cross-modal layers, a parameter-free reducer first prunes image tokens according to their relevance to the text, then merges similar tokens separately within the image and text modalities. Unpruned and unmerged tokens continue through the encoder. Applying reducers at multiple layers progressively lowers the token count and the cost of subsequent cross-modal computation, while modality-specific merging avoids combining text and image representations with each other.

  2. Knowl 2 — Text-to-image attention scores guide image-token pruning

    equation

    At a reducer, let TT be the set of text tokens, VV the set of image tokens, ∣T∣|T| and ∣V∣|V| their counts, and HH the number of attention heads. Let AtvhA^h_{tv} be the text-to-image attention score from text token tt to image token vv in head hh. PuMer assigns each image token a text-saliency score:

    sv=1∣T∣∑t=1∣T∣∑h=1HAtvh.s_v = \frac{1}{|T|}\sum_{t=1}^{|T|}\sum_{h=1}^{H} A^h_{tv}.

    For pruning ratio kk, where kk is the fraction of image tokens to remove, PuMer retains the k′=(1−k)∣V∣k'=(1-k)|V| image tokens with the largest scores and discards the rest. Thus, tokens receiving more total attention from the text are prioritized for retention. The attention scores are already computed by the VL model's cross-modal attention.

  3. Knowl 3 — Bipartite soft matching merges tokens within each modality

    algorithm

    PuMer uses bipartite soft matching independently on the retained image tokens and on the text tokens. For a token set XX and merge ratio qq (the image merge ratio for image tokens, or the text merge ratio for text tokens), the procedure is:

    Input: token vectors X, corresponding key vectors K, merge ratio q
    Output: token vectors after merging
    1. Divide X into sets E and O according to even and odd token order.
    2. For each token in O, find its most similar token in E.
    3. Score each resulting pair by the dot product of the pair's key vectors.
    4. Select the top q|X| pairs by score.
    5. Connect the selected pairs; for each connected group, replace its token vectors
       with their average. Keep tokens not in a selected group unchanged.
    6. Return the merged and unchanged tokens.

    The similarity between tokens ii and jj is the dot product Ki⊤KjK_i^\top K_j of their key vectors, which the Transformer has already computed. The selected pairs may share tokens; tokens joined into the same connected group are averaged together. This operation reduces the number of tokens while retaining a representation of similar content.

  4. Knowl 4 — Cascaded reducer placement trades accuracy for efficiency

    model/method

    PuMer leaves an initial portion of the cross-modal encoder untouched, then applies token reducers at several later layers. Pruned image tokens are not processed by subsequent layers. Reducing tokens earlier saves computation in more layers but can cause greater accuracy loss; spreading reduction across layers limits abrupt information loss. The paper generally uses reducers at three or four locations and reduction ratios in the range 0.1–0.5. One reported ViLT configuration for SNLI-VE uses layers 2, 4, 6, and 8, with image pruning ratio 0.1, image merge ratio 0.3, and text merge ratio 0.2.

  5. Knowl 5 — PuMer is trained by fine-tuning with optional knowledge distillation

    model/method

    PuMer's reducers have no trainable parameters and can be added to an existing VL model without changing its architecture. Training otherwise follows fine-tuning of the original model; the authors add a knowledge-distillation loss using the original fine-tuned VL model as teacher to reduce the accuracy gap. In the reported setup, original ViLT and METER baselines were fine-tuned for 10 epochs, while PuMer models were trained for up to 20 epochs with early stopping after five epochs without accuracy improvement. Training used four Nvidia A100 GPUs. Because pruning and merging reduce tokens during the forward computation in training as well as inference, the authors report 15%–20% faster training in practice.

  6. Knowl 6 — PuMer retains accuracy while improving measured inference efficiency

    data/table

    The evaluation applies PuMer to ViLT (110 million parameters) and METER (330 million parameters) on Flickr30K text-to-image and image-to-text retrieval, VQAv2, SNLI-VE, and NLVR2. Models were fine-tuned on task training data, evaluated on test data, and accuracy values are averages over three runs. Throughput increase is relative to the original fine-tuned model; memory reduction is the reduction in peak inference memory. Throughput was measured for 30 seconds at the largest batch size that fit on the GPU (GTX 1080 Ti for ViLT and A40 for METER); memory comparisons used the same batch size for the original and PuMer models.

    The results show 1.74×–2.07× throughput increases and 38%–51% memory reductions, with reported accuracy decreases below one point on every listed task.

    Model Task Original accuracy PuMer accuracy Change Throughput increase Memory reduction
    METER Flickr30K text-to-image retrieval 94.7 93.8 -0.9 1.81×\times 38%
    METER Flickr30K image-to-text retrieval 82.0 81.2 -0.8 1.81×\times 38%
    METER VQAv2 77.5 76.8 -0.7 1.82×\times 38%
    METER SNLI-VE 81.1 80.3 -0.8 2.07×\times 43%
    METER NLVR2 82.7 82.2 -0.5 1.79×\times 38%
    ViLT Flickr30K text-to-image retrieval 78.2 77.6 -0.6 1.78×\times 46%
    ViLT Flickr30K image-to-text retrieval 60.2 59.6 -0.7 1.78×\times 46%
    ViLT VQAv2 69.5 68.9 -0.6 1.76×\times 45%
    ViLT SNLI-VE 76.0 75.6 -0.4 2.01×\times 51%
    ViLT NLVR2 75.5 74.9 -0.6 1.74×\times 45%

    The paper also reports inference FLOPs in GFLOPs (original to PuMer): METER uses 92 to 64.7 on VQAv2 (1.42×), 92 to 59 on SNLI-VE (1.56×), and 184 to 131 on NLVR2 (1.40×); ViLT uses 16 to 8.7 on VQAv2 (1.84×), 16 to 7.7 on SNLI-VE (2.08×), and 32 to 17.4 on NLVR2 (1.84×).

  7. Knowl 7 — PuMer compares favorably with token-reduction and image-downsampling baselines

    empirical result

    For ViLT on VQAv2, the paper compares PuMer with DynamicViT and ToMe across settings with different throughput and accuracy trade-offs. The reported comparison finds that at similar throughput increase (for example, around 1.8×), PuMer has the highest accuracy; for an accuracy-drop constraint below one point, PuMer provides a larger throughput increase.

    For METER on VQAv2, PuMer is also compared with reducing the input image resolution. Smaller images improve efficiency but produce larger accuracy drops than PuMer at comparable settings. The 320×320 baseline is 0.2 accuracy points above PuMer at 384×384, but has about 20% lower throughput. Applying PuMer to 320×320 inputs gives a further efficiency gain over the 320×320 baseline.

    Method Image resolution VQAv2 accuracy Throughput increase Memory reduction
    Resolution baseline 192×\times192 74.3 (-3.2) 4.23×\times 75%
    Resolution baseline 224×\times224 75.2 (-2.3) 3.48×\times 66%
    Resolution baseline 256×\times256 76.1 (-1.4) 2.67×\times 54%
    Resolution baseline 320×\times320 77.0 (-0.5) 1.62×\times 37%
    PuMer 320×\times320 76.3 (-1.2) 2.86×\times 59%
    PuMer 384×\times384 76.8 (-0.7) 1.82×\times 38%
    Original METER 384×\times384 77.5 1×\times 0%

    Accuracy changes in parentheses are relative to the original METER model at 384×384 resolution (77.5 accuracy). At 320×320, PuMer's 2.86× throughput relative to the original 320×320 model corresponds to 1.76× the throughput of that resolution baseline.

  8. Knowl 8 — Pruning, merging, and distillation each contribute to the ViLT result

    empirical result

    An ablation on ViLT and VQAv2 measures the effect of removing each PuMer component. PuMer with pruning, merging, and distillation reaches 68.9 accuracy (0.6 points below the original ViLT) at 1.76× throughput. Removing text-informed pruning reduces throughput to 1.52× and gives 69.2 accuracy; removing modality-aware merging reduces throughput to 1.46× and gives 69.1 accuracy. Removing distillation retains 1.76× throughput but lowers accuracy to 68.6, a 0.9-point drop from the original 69.5. The two token-reduction operations therefore provide complementary efficiency gains, while distillation recovers some of the accuracy loss.

  9. Knowl 9 — Reducer locations and ratios determine the ViLT accuracy–throughput trade-off

    data/table

    The design study evaluates ViLT on SNLI-VE; the original model has 76.0 accuracy and 1.00× throughput. Increasing pruning or merging ratios generally raises throughput but can reduce accuracy, while the layer locations change how much later computation is saved. In particular, reducing at layers 2, 3, and 4 gives 2.03× throughput but a 1.8-point accuracy drop; reducing at layers 7, 8, and 9 gives only 1.31× throughput with a 0.1-point drop. The selected four-location setting balances these outcomes at 2.01× throughput and a 0.4-point drop.

    Choice Reduction layers Prune ratio Image merge ratio Text merge ratio VE accuracy (change) Throughput increase
    Ratios 2,5,8 0.1 0.3 0.2 75.8 (-0.2) 1.77×\times
    Ratios 2,5,8 0.3 0.3 0.2 74.7 (-1.3) 2.04×\times
    Ratios 2,5,8 0.1 0.3 0.5 74.9 (-1.1) 1.89×\times
    Ratios 2,5,8 0.1 0.5 0.2 73.8 (-2.1) 2.12×\times
    Number of layers 2 0.1 0.3 0.2 75.9 (-0.15) 1.43×\times
    Number of layers 2,4 0.1 0.3 0.2 75.8 (-0.2) 1.69×\times
    Number of layers 2,4,6 0.1 0.3 0.2 75.7 (-0.3) 1.80×\times
    Locations 2,3,4 0.2 0.2 0.2 74.2 (-1.8) 2.03×\times
    Locations 7,8,9 0.2 0.2 0.2 75.9 (-0.1) 1.31×\times
    PuMer selected setting 2,4,6,8 0.1 0.3 0.2 75.6 (-0.4) 2.01×\times
    Original ViLT – – – – 76.0 1.00×\times

    The study supports distributing reduction across several layers rather than relying only on an early or late reduction: earlier locations can save more computation, but their accuracy cost can be larger.

  10. Knowl 10 — PuMer is limited when cross-modal layers are not the main computational cost

    limitation

    PuMer reduces tokens within the cross-modal encoder, so its end-to-end speed benefit is marginal for VL models in which that encoder is relatively lightweight compared with the vision encoder. The paper names ALBEF and X-VLM as examples of this architecture pattern. Reducing image tokens inside the vision encoder could address this limitation, but the paper leaves that extension unexplored.

Coverage note — No substantial contributed material was omitted; exhaustive per-task fine-tuning hyperparameters and library implementation details were left out because they are supplementary to the method and reported findings.

References

  1. 1.TPrune: Efficient Transformer Pruning for Mobile Devices: ACM Transactions on Cyber-Physical Systems: Vol 5, No 3.
  2. 2.
    1. Deepspeed.
  3. 3.
    1. huggingface/accelerate.
  4. 4.Daniel Bolya, Cheng-Yang Fu, Xiaoliang Dai, Peizhao Zhang, Christoph Feichtenhofer, and Judy Hoffman. 2022. Token Merging: Your ViT But Faster. ArXiv:2210.09461 [cs].
  5. 5.Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning. 2015. A large annotated corpus for learning natural language inference. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, pages 632–642, Lisbon, Portugal. Association for Computational Linguistics.
  6. 6.Qingqing Cao, Prerna Khanna, Nicholas D. Lane, and Aruna Balasubramanian. 2022. MobiVQA: Efficient On-Device Visual Question Answering. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies, 6(2):44:1–44:23.
  7. 7.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  8. 8.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. 2020. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.
  9. 9.Zi-Yi Dou, Yichong Xu, Zhe Gan, Jianfeng Wang, Shuohang Wang, Lijuan Wang, Chenguang Zhu, Pengchuan Zhang, Lu Yuan, Nanyun Peng, Zicheng Liu, and Michael Zeng. 2022. An empirical study of training end-to-end vision-and-language transformers. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 18166–18176.
  10. 10.Zhiyuan Fang, Jianfeng Wang, Xiaowei Hu, Lijuan Wang, Yezhou Yang, and Zicheng Liu. 2021. Compressing Visual-Linguistic Model via Knowledge Distillation. pages 1428–1438.
  11. 11.Saurabh Goyal, Anamitra Roy Choudhury, Saurabh Raje, Venkatesan Chakaravarthy, Yogish Sabharwal, and Ashish Verma. 2020. PoWER-BERT: Accelerating BERT Inference via Progressive Word-vector Elimination. In Proceedings of the 37th International Conference on Machine Learning, pages 3690–3699. PMLR. ISSN: 2640-3498.
  12. 12.Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh. 2017. Making the v in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering. pages 6904–6913.
  13. 13.Benjamin Graham, Alaaeldin El-Nouby, Hugo Touvron, Pierre Stock, Armand Joulin, Hervé Jégou, and Matthijs Douze. 2021. LeViT: A Vision Transformer in ConvNet’s Clothing for Faster Inference. pages 12259–12269.
  14. 14.Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. 2015. Distilling the Knowledge in a Neural Network. ArXiv:1503.02531 [cs, stat].
  15. 15.Wonjae Kim, Bokyung Son, and Ildoo Kim. 2021. ViLT: Vision-and-Language Transformer Without Convolution or Region Supervision. In Proceedings of the 38th International Conference on Machine Learning, pages 5583–5594. PMLR. ISSN: 2640-3498.
  16. 16.Zhenglun Kong, Peiyan Dong, Xiaolong Ma, Xin Meng, Mengshu Sun, Wei Niu, Xuan Shen, Geng Yuan, Bin Ren, Minghai Qin, Hao Tang, and Yanzhi Wang. 2022. SPViT: Enabling Faster Vision Transformers via Soft Token Pruning. ArXiv:2112.13890 [cs].
  17. 17.François Lagunas, Ella Charlaix, Victor Sanh, and Alexander Rush. 2021. Block Pruning For Faster Transformers. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 10619–10629, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  18. 18.Junnan Li, Ramprasaath Selvaraju, Akhilesh Gotmare, Shafiq Joty, Caiming Xiong, and Steven Chu Hong Hoi. 2021. Align before fuse: Vision and language representation learning with momentum distillation. In Advances in neural information processing systems, volume 34, pages 9694–9705. Curran Associates, Inc.
  19. 19.Youwei Liang, Chongjian Ge, Zhan Tong, Yibing Song, Jue Wang, and Pengtao Xie. 2021. EViT: Expediting Vision Transformers via Token Reorganizations.
  20. 20.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C. Lawrence Zitnick. 2014. Microsoft COCO: Common Objects in Context. In Computer Vision – ECCV 2014, Lecture Notes in Computer Science, pages 740–755, Cham. Springer International Publishing.
  21. 21.Weijie Liu, Peng Zhou, Zhiruo Wang, Zhe Zhao, Haotang Deng, and Qi Ju. 2020. FastBERT: a Self-distilling BERT with Adaptive Inference Time. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 6035–6044, Online. Association for Computational Linguistics.
  22. 22.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. RoBERTa: A Robustly Optimized BERT Pretraining Approach. Number: arXiv:1907.11692 arXiv:1907.11692 [cs].
  23. 23.Dmitrii Marin, Jen-Hao Rick Chang, Anurag Ranjan, Anish Prabhu, Mohammad Rastegari, and Oncel Tuzel. 2021. Token Pooling in Vision Transformers. ArXiv:2110.03860 [cs].
  24. 24.Piotr Nawrot, Jan Chorowski, Adrian Łancucki, and ´ Edoardo M. Ponti. 2022. Efficient Transformers with Dynamic Token Pooling. ArXiv:2211.09761 [cs] version: 1.
  25. 25.Michał Pietruszka, Łukasz Borchmann, and Filip Gralinski. 2020. ´ Sparsifying Transformer Models with Differentiable Representation Pooling. arXiv:2009.05169 [cs]. ArXiv: 2009.05169.
  26. 26.Bryan A. Plummer, Liwei Wang, Chris M. Cervantes, Juan C. Caicedo, Julia Hockenmaier, and Svetlana Lazebnik. 2015. Flickr30k Entities: Collecting Region-to-Phrase Correspondences for Richer Image-to-Sentence Models. In 2015 IEEE International Conference on Computer Vision (ICCV), pages 2641–2649. ISSN: 2380-7504.
  27. 27.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021. Learning Transferable Visual Models From Natural Language Supervision. In Proceedings of the 38th International Conference on Machine Learning, pages 8748–8763. PMLR. ISSN: 2640-3498.
  28. 28.Yongming Rao, Wenliang Zhao, Benlin Liu, Jiwen Lu, Jie Zhou, and Cho-Jui Hsieh. 2021. DynamicViT: Efficient Vision Transformers with Dynamic Token Sparsification. arXiv:2106.02034 [cs]. ArXiv: 2106.02034.
  29. 29.Michael S. Ryoo, A. J. Piergiovanni, Anurag Arnab, Mostafa Dehghani, and Anelia Angelova. 2021. TokenLearner: What Can 8 Learned Tokens Do for Images and Videos?
  30. 30.Roy Schwartz, Gabriel Stanovsky, Swabha Swayamdipta, Jesse Dodge, and Noah A. Smith. 2020. The Right Tool for the Job: Matching Model and Instance Complexities. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 6640–6651, Online. Association for Computational Linguistics.
  31. 31.Alane Suhr, Stephanie Zhou, Ally Zhang, Iris Zhang, Huajun Bai, and Yoav Artzi. 2019. A Corpus for Reasoning about Natural Language Grounded in Photographs. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 6418–6428, Florence, Italy. Association for Computational Linguistics.
  32. 32.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in neural information processing systems, pages 5998–6008.
  33. 33.Jianfeng Wang, Xiaowei Hu, Pengchuan Zhang, Xiujun Li, Lijuan Wang, Lei Zhang, Jianfeng Gao, and Zicheng Liu. 2020. MiniVLM: A Smaller and Faster Vision-Language Model. arXiv:2012.06946 [cs]. ArXiv: 2012.06946.
  34. 34.Peng Wang, An Yang, Rui Men, Junyang Lin, Shuai Bai, Zhikang Li, Jianxin Ma, Chang Zhou, Jingren Zhou, and Hongxia Yang. 2022. OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework. Technical Report arXiv:2202.03052, arXiv. ArXiv:2202.03052 [cs] version: 2 type: article.
  35. 35.Zirui Wang, Jiahui Yu, Adams Wei Yu, Zihang Dai, Yulia Tsvetkov, and Yuan Cao. 2021. SimVLM: Simple Visual Language Model Pretraining with Weak Supervision.
  36. 36.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander M. Rush. 2020. Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 38–45, Online. Association for Computational Linguistics.
  37. 37.Ning Xie, Farley Lai, Derek Doran, and Asim Kadav. 2019. Visual Entailment Task for Visually-Grounded Language Learning. Technical Report arXiv:1811.10582, arXiv. ArXiv:1811.10582 [cs] type: article.
  38. 38.Ji Xin, Rodrigo Nogueira, Yaoliang Yu, and Jimmy Lin. 2020. Early Exiting BERT for Efficient Document Ranking. In Proceedings of SustaiNLP: Workshop on Simple and Efficient Natural Language Processing, pages 83–88, Online. Association for Computational Linguistics.
  39. 39.Jiarui Xu, Shalini De Mello, Sifei Liu, Wonmin Byeon, Thomas Breuel, Jan Kautz, and Xiaolong Wang. 2022. GroupViT: Semantic Segmentation Emerges From Text Supervision. pages 18134–18144.
  40. 40.Hongxu Yin, Arash Vahdat, Jose M. Alvarez, Arun Mallya, Jan Kautz, and Pavlo Molchanov. 2022. A-vit: Adaptive tokens for efficient vision transformer. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 10809–10818.
  41. 41.Peter Young, Alice Lai, Micah Hodosh, and Julia Hockenmaier. 2014. From image descriptions to visual denotations: new similarity metrics for semantic inference over event descriptions. Transactions of the Association for Computational Linguistics, 2:67–78. Place: Cambridge, MA Publisher: MIT Press.
  42. 42.Fang Yu, Kun Huang, Meng Wang, Yuan Cheng, Wei Chu, and Li Cui. 2022. Width & depth pruning for vision transformers. In AAAI Conference on Artificial Intelligence (AAAI), volume 2022.
  43. 43.Hao Yu and Jianxin Wu. 2021. A Unified Pruning Framework for Vision Transformers. arXiv:2111.15127 [cs]. ArXiv: 2111.15127.
  44. 44.Yan Zeng, Xinsong Zhang, and Hang Li. 2021. Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual Concepts.
  45. 45.Pengchuan Zhang, Xiujun Li, Xiaowei Hu, Jianwei Yang, Lei Zhang, Lijuan Wang, Yejin Choi, and Jianfeng Gao. 2021. VinVL: Revisiting Visual Representations in Vision-Language Models. pages 5579–5588.
  46. 46.Wangchunshu Zhou, Canwen Xu, Tao Ge, Julian McAuley, Ke Xu, and Furu Wei. 2020. BERT loses patience: Fast and robust inference with early exit. In Advances in neural information processing systems, volume 33, pages 18330–18341. Curran Associates, Inc.

Citation

MLA
Cao, Q., et al. “PuMer: Pruning and Merging Tokens for Efficient Vision Language Models”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2023, pp. 12890–903, https://doi.org/10.18653/v1/2023.acl-long.721.
APA
Cao, Q., Paranjape, B., & Hajishirzi, H. (2023). PuMer: Pruning and Merging Tokens for Efficient Vision Language Models. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 12890–12903. https://doi.org/10.18653/v1/2023.acl-long.721
Chicago
Cao, Q., B. Paranjape, and H. Hajishirzi. 2023. “PuMer: Pruning and Merging Tokens for Efficient Vision Language Models”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 12890–903. https://doi.org/10.18653/v1/2023.acl-long.721.
Harvard
Cao, Q., Paranjape, B. and Hajishirzi, H. (2023) “PuMer: Pruning and Merging Tokens for Efficient Vision Language Models”, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 12890–12903. Available at: https://doi.org/10.18653/v1/2023.acl-long.721.
Vancouver
1. Cao Q, Paranjape B, Hajishirzi H (2023) PuMer: Pruning and Merging Tokens for Efficient Vision Language Models. In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 12890–12903

BibTeX

@inproceedings{cao-etal-2023-pumer,
    title = "{P}u{M}er: Pruning and Merging Tokens for Efficient Vision Language Models",
    author = "Cao, Qingqing  and
      Paranjape, Bhargavi  and
      Hajishirzi, Hannaneh",
    editor = "Rogers, Anna  and
      Boyd-Graber, Jordan  and
      Okazaki, Naoaki",
    booktitle = "Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = jul,
    year = "2023",
    address = "Toronto, Canada",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.acl-long.721/",
    doi = "10.18653/v1/2023.acl-long.721",
    pages = "12890--12903"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/