MADTP: Multimodal Alignment-Guided Dynamic Token Pruning for Accelerating Vision-Language Transformer

Jianjian CaoPeng YeShengze LiChong YuYansong TangJiwen LuTao Chen

article2024CVPR75 citations

Proposes a multimodal alignment-guided dynamic token pruning framework that cuts Vision-Language Transformer computation by up to 80% with minimal accuracy loss by aligning cross-modal representations to prevent false token removal and adaptively tuning layer-wise pruning ratios per input instance.

Listen

Vision-Language Transformers power many state-of-the-art multimodal artificial intelligence applications, including visual reasoning, image captioning, image-text retrieval, and visual question answering. However, processing large numbers of visual and textual tokens creates significant computational overhead, which limits real-time deployment and inflates inference costs. Existing token-reduction methods either prune tokens within single modalities independently or apply static reduction ratios. These approaches frequently discard visual details that are critical for text comprehension (or vice versa) and fail to adjust processing power based on input complexity.

The article introduces and evaluates Multimodal Alignment-Guided Dynamic Token Pruning (MADTP), a compression framework designed to accelerate multimodal models by aligning visual and language features before pruning and dynamically adjusting computational effort per input.

The framework introduces two primary mechanisms: a Multi-modality Alignment Guidance module and a Dynamic Token Pruning module. The alignment module uses shared learnable tokens to map image and text features into a common semantic space, ensuring that retained tokens are mutually relevant across modalities. The dynamic pruning module scores each token by combining class-level, self-attention, and cross-modal attention scores, and then removes less relevant tokens using an adaptive, instance-specific threshold. The authors validated MADTP on established benchmark architectures (BLIP and CLIP) across four major datasets: NLVR2 for visual reasoning, COCO and Flickr30k for retrieval and captioning, and VQA v2.0 for visual question answering.

The evaluation yielded several key findings. First, MADTP achieved aggressive computation cuts with minimal accuracy loss: on the BLIP model performing visual reasoning, it reduced floating-point operations by 80% while retaining test accuracy within 3.86% of the uncompressed model. Second, in image-text retrieval tasks at high compression ratios (around 75% computation reduction), MADTP substantially outperformed prior methods, improving image-to-text recall@1 on the COCO dataset by 10.1 percentage points over the state-of-the-art baseline. Third, on visual question answering tasks, MADTP achieved a 57% reduction in computational complexity with less than a 1% drop in accuracy. Finally, ablation experiments confirmed that cross-modal feature alignment contributed roughly a 2% improvement in performance compared to pruning without multimodal guidance.

These findings indicate that cross-modal alignment is critical for compressing multimodal systems without degrading accuracy. In practical terms, reducing computational demands by 50% to 80% directly lowers cloud inference expenses, reduces latency, and facilitates deployment of high-performing vision-language models on edge hardware and resource-constrained environments.

Organizations deploying vision-language models should consider adopting multimodal alignment-guided dynamic pruning as a standard optimization pipeline. Prior to broad deployment, engineering teams should benchmark MADTP in application-specific pilot pipelines to determine the optimal trade-off between reduction ratios and task-specific accuracy requirements.

While the results demonstrate robust performance across multiple standard benchmarks and architectures, the framework requires tuning hyperparameters such as the number and dimension of learnable tokens. Overall confidence in the reported efficiency and accuracy improvements is high for the tested tasks, though practitioners should validate performance on complex domain-specific tasks and hardware architectures.

Cover for MADTP: Multimodal Alignment-Guided Dynamic Token Pruning for Accelerating Vision-Language Transformer

Abstract

Vision-Language Transformers (VLTs) have shown great success recently, but are meanwhile accompanied by heavy computation costs, where a major reason can be attributed to the large number of visual and language tokens. Existing token pruning research for compressing VLTs mainly follows a single-modality-based scheme yet ignores the critical role of aligning different modalities for guiding the token pruning process, causing the important tokens for one modality to be falsely pruned in another modality branch. Meanwhile, existing VLT pruning works also lack the flexibility to dynamically compress each layer based on different input samples. To this end, we propose a novel framework named Multimodal Alignment-Guided Dynamic Token Pruning (MADTP) for accelerating various VLTs. Specifically, we first introduce a well-designed Multi-modality Alignment Guidance (MAG) module that can align features of the same semantic concept from different modalities, to ensure the pruned tokens are less important for all modalities. We further design a novel Dynamic Token Pruning (DTP) module, which can adaptively adjust the token compression ratio in each layer based on different input instances. Extensive experiments on various benchmarks demonstrate that MADTP significantly reduces the computational complexity of kinds of multimodal models while preserving competitive performance. Notably, when applied to the BLIP model in the NLVR2 dataset, MADTP can reduce the GFLOPs by 80% with less than 4% performance degradation. The code is available at https://github.com/double125/MADTP.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 2.1. Vision-Language Transformer
  • 2.2. Multimodal Compression
  • 2.3. Token Merging and Pruning
  • 3. Methodology
  • 3.1. Preliminaries
  • 3.2. Multi-modality Alignment Guidance
  • 3.3. Dynamic Token Pruning
  • 3.4. Objective Function
  • 4. Experiments
  • 4.1. Experimental Setup
  • 4.2. Experiments on the Visual Reasoning Task
  • 4.3. Experiments on the Retrieval Task
  • 4.4. Experiments on the Image Caption Task
  • 4.5. Experiments on the Visual QA Task
  • 5. Conclusion
  • 6. Acknowledgments
  • References

Knowls

  1. Knowl 1 — MADTP framework for multimodal dynamic token pruning

    model/method

    Multimodal Alignment-Guided Dynamic Token Pruning (MADTP) compresses Vision-Language Transformers (VLTs) by jointly using cross-modal alignment and input-adaptive token pruning. A VLT receives an image-text pair, produces visual tokens V∈RN×dvV\in\mathbb{R}^{N\times d_v} and language tokens L∈RM×dlL\in\mathbb{R}^{M\times d_l}, and processes each modality through Transformer blocks. MADTP inserts two components into this architecture:

    • Multi-modality Alignment Guidance (MAG) aligns visual and language representations through shared learnable tokens and supplies cross-modal token-attention maps.
    • Dynamic Token Pruning (DTP) is inserted between the self-attention and feed-forward layers of each Transformer block. It combines class-attention importance, within-modality self-attention importance, and MAG-derived cross-modal importance, then removes tokens using an input- and layer-dependent threshold.

    Discarded tokens are not simply lost: their information is combined into a new token using their token-importance scores, and the resulting token is passed onward with the retained tokens. MAG weights are shared across the VLT layers. The architecture diagram on page 4 depicts the two branches, the shared MAG modules between them, and DTP modules inside successive Transformer blocks.

  2. Knowl 2 — Multi-modality Alignment Guidance module

    model/method

    The MAG module maps visual and language tokens at a given VLT layer into a common dimension dd and compares both modalities through KK shared learnable tokens. For visual tokens V∈RN×dvV\in\mathbb{R}^{N\times d_v} and language tokens L∈RM×dlL\in\mathbb{R}^{M\times d_l}, the layer-specific projections are

    V′=VWv+Bv,L′=LWt+Bt,V'=VW_v+B_v,\qquad L'=LW_t+B_t,

    where Wv∈Rdv×dW_v\in\mathbb{R}^{d_v\times d}, Wt∈Rdl×dW_t\in\mathbb{R}^{d_l\times d}, Bv∈R1×dB_v\in\mathbb{R}^{1\times d}, and Bt∈R1×dB_t\in\mathbb{R}^{1\times d} are trainable parameters broadcast over tokens. Let E∈RK×dE\in\mathbb{R}^{K\times d} be the learnable-token matrix and let dkd_k denote the attention channel dimension. MAG computes visual token-attention weights row-wise as

    Atokenv=softmax⁡(EV′Tdk)∈RK×N,A^v_{\mathrm{token}}=\operatorname{softmax}\left(\frac{EV'^{\mathsf T}}{\sqrt{d_k}}\right)\in\mathbb{R}^{K\times N},

    and obtains the corresponding visual features

    Ev=AtokenvV′∈RK×d.E^v=A^v_{\mathrm{token}}V'\in\mathbb{R}^{K\times d}.

    The language attention map Atokenl∈RK×MA^l_{\mathrm{token}}\in\mathbb{R}^{K\times M} and language features El∈RK×dE^l\in\mathbb{R}^{K\times d} are computed analogously from L′L'. The same learnable-token mechanism explicitly represents semantic associations across modalities; its attention maps are subsequently used to identify tokens that are redundant in both branches. The MAG modules share their weights across the VLT layers.

  3. Knowl 3 — Multimodal token-importance score

    equation

    For each modality, MADTP assigns every input token a Token Importance Score (TIS) by averaging three normalized signals:

    TIS⁡(xi)=Scls(xi)+Sself(xi)+Stoken(xi)3,\operatorname{TIS}(x_i)=\frac{S_{\mathrm{cls}}(x_i)+S_{\mathrm{self}}(x_i)+S_{\mathrm{token}}(x_i)}{3},

    where xix_i is token ii, SclsS_{\mathrm{cls}} is the token's class-attention score, SselfS_{\mathrm{self}} measures its importance within the same modality, and StokenS_{\mathrm{token}} measures its alignment with the learnable MAG tokens. For a modality with nn tokens, self-attention produces Aself∈Rn×nA_{\mathrm{self}}\in\mathbb{R}^{n\times n} and MAG produces Atoken∈RK×nA_{\mathrm{token}}\in\mathbb{R}^{K\times n}. The two attention-derived scores are normalized across tokens as

    Sself(xi)=max⁡jAself[j,i]∑r=1nmax⁡jAself[j,r],Stoken(xi)=max⁡kAtoken[k,i]∑r=1nmax⁡kAtoken[k,r].S_{\mathrm{self}}(x_i)=\frac{\max_j A_{\mathrm{self}}[j,i]}{\sum_{r=1}^{n}\max_j A_{\mathrm{self}}[j,r]}, \qquad S_{\mathrm{token}}(x_i)=\frac{\max_{k} A_{\mathrm{token}}[k,i]}{\sum_{r=1}^{n}\max_{k} A_{\mathrm{token}}[k,r]}.

    Here jj indexes attention-source tokens, kk indexes the KK learnable MAG tokens, and i,ri,r index tokens in the current modality. The resulting normalized scores incorporate task relevance, within-modality contextual relevance, and cross-modal semantic relevance, reducing the chance that a token important to the other modality is pruned.

  4. Knowl 4 — Instance- and layer-adaptive DTP pruning rule

    algorithm

    DTP derives a different pruning threshold for every input instance and Transformer layer from the MAG alignment map. For a modality with nn tokens, let Atoken∈RK×nA_{\mathrm{token}}\in\mathbb{R}^{K\times n} be the MAG attention map, let TIS⁡∈Rn\operatorname{TIS}\in\mathbb{R}^{n} be the token-importance vector, and let T>0T>0 be the pruning temperature. The sparse attention map is computed row-wise:

    A^token=sparsemax⁡(TAtoken),\widehat A_{\mathrm{token}}=\operatorname{sparsemax}(T A_{\mathrm{token}}),

    where for a vector z∈Rnz\in\mathbb{R}^{n},

    sparsemax⁡(z)=arg⁡min⁡p∈Δn−1∥p−z∥22,Δn−1={p∈Rn:1Tp=1, p≥0}.\operatorname{sparsemax}(z)=\arg\min_{p\in\Delta^{n-1}}\|p-z\|_2^2, \qquad \Delta^{n-1}=\{p\in\mathbb{R}^{n}:\mathbf{1}^{\mathsf T}p=1,\ p\ge 0\}.

    Each of the KK rows of A^token\widehat A_{\mathrm{token}} weights the TIS values, and the smallest weighted score becomes the threshold:

    θ=min⁡k∈{1,…,K}∑i=1nA^token[k,i]TIS⁡(xi).\theta=\min_{k\in\{1,\ldots,K\}}\sum_{i=1}^{n}\widehat A_{\mathrm{token}}[k,i]\operatorname{TIS}(x_i).

    The pruning mask is

    Mp(xi)={1,TIS⁡(xi)>θ,0,TIS⁡(xi)≤θ.M_p(x_i)= \begin{cases} 1,&\operatorname{TIS}(x_i)>\theta,\\ 0,&\operatorname{TIS}(x_i)\le\theta. \end{cases}

    Tokens with mask value 11 continue individually; tokens with mask value 00 are removed from the next computation, while their TIS-weighted information is aggregated into a new token. Because θ\theta is calculated from each input's multimodal attention pattern at each layer, DTP provides both instance-wise and layer-wise dynamic compression rather than removing a fixed number of tokens everywhere. During training, the temperature TT is adjusted by epoch according to the GFLOPs of the pruned model.

  5. Knowl 5 — Alignment-aware optimization objective

    equation

    MADTP preserves the task-specific VLT objective while adding a loss that encourages the visual and language features produced by corresponding MAG learnable tokens to agree. Let Ltask\mathcal{L}_{\mathrm{task}} be the loss for the downstream multimodal task, let Eiv,Eil∈RdE_i^v,E_i^l\in\mathbb{R}^{d} be the visual and language features associated with learnable token ii, let KK be the number of learnable tokens, and let α≥0\alpha\ge 0 be the alignment-loss weight. The total training loss is

    L=Ltask+αLsim,\mathcal{L}=\mathcal{L}_{\mathrm{task}}+\alpha\mathcal{L}_{\mathrm{sim}},

    with

    Lsim=1K∑i=1K(1−cos⁡(Eiv,Eil)),\mathcal{L}_{\mathrm{sim}}=\frac{1}{K}\sum_{i=1}^{K}\left(1-\cos(E_i^v,E_i^l)\right),

    where cos⁡(⋅,⋅)\cos(\cdot,\cdot) is cosine similarity. Thus, the optimization encourages corresponding learnable MAG tokens to encode semantically related visual and language content while the task loss maintains downstream performance.

  6. Knowl 6 — Evaluation protocol and implementation configuration

    experimental setup

    MADTP was evaluated by compressing pretrained CLIP and BLIP VLTs on four multimodal datasets: NLVR2, containing 107,292 image-text pairs; COCO, containing approximately 330,000 images with five descriptions per image; Flickr30K, containing 31,783 images with descriptive titles; and VQA v2.0, an open-ended human-annotated visual question-answering dataset. Task-specific metrics were used for each benchmark, and computational complexity was measured in GFLOPs per image-text pair.

    The models were initialized from pretrained weights used by the UPop implementation. Compression training used 8 NVIDIA A100 GPUs with batch size 32 and alignment-loss coefficient α=0.1\alpha=0.1. The target reduce ratio denotes the desired fraction of original GFLOPs removed. The DTP temperature was adjusted at each epoch according to the GFLOPs of the currently pruned model. For the main NLVR2 study, the authors also implemented Static Token Pruning (STP), which removes a fixed number of tokens at every layer using the same TIS formula but without DTP's adaptive threshold.

  7. Knowl 7 — NLVR2 visual-reasoning compression results

    data/table

    The numerical comparison on page 6 evaluates BLIP compression on NLVR2. It compares fixed-count STP, the UPop baseline, and MADTP at target GFLOPs reduction ratios of 0.3–0.8. Dev Acc and Test Acc are percentages, while GFLOPs is the computation per image-text pair. MADTP preserves substantially more accuracy than UPop at the same high compression levels: at an 0.8 reduction ratio it retains 79.22 test accuracy at only 26.46 GFLOPs, versus UPop's 57.79 test accuracy at 19.08 GFLOPs.

    Could not parse LaTeX table

    Relative to the uncompressed BLIP model, MADTP reduces GFLOPs by 30%, 50%, 60%, 70%, and 80% at the five target ratios, with corresponding test-accuracy decreases of approximately 0%, 0.23%, 0.66%, 1.85%, and 3.86%. The token-mask visualizations on page 8 show that later blocks retain image patches and text words that are semantically relevant to the paired description while discarding background or weakly related tokens.

  8. Knowl 8 — Retrieval performance across CLIP and BLIP

    data/table

    The retrieval results on page 7 compare uncompressed models, UPop, and MADTP on Flickr30K and COCO for both image-to-text and text-to-image retrieval. R@1, R@5, and R@10 are recall percentages, and higher values are better. MADTP generally retains or improves retrieval quality at the same target reduction ratio; its clearest gain is CLIP on COCO at ratio 0.75, where image-to-text R@1 rises from 56.1 to 66.2 while GFLOPs fall from 105.9 to 92.4.

    Could not parse LaTeX table
  9. Knowl 9 — Captioning and visual-question-answering compression

    data/table

    The page-8 evaluation tests MADTP's generalization beyond retrieval and visual reasoning by compressing BLIP on COCO image captioning and VQA v2.0. CIDEr and SPICE measure caption quality; Test-dev and Test-std measure VQA accuracy; all task metrics are higher-is-better. At reduction ratio 0.75, MADTP obtains CIDEr 120.1 at 22.1 GFLOPs, exceeding UPop's CIDEr 117.4 at 22.2 GFLOPs. On VQA at ratio 0.5, MADTP reduces computation from 186.1 to 79.4 GFLOPs while Test-dev accuracy decreases from 77.4 to 76.8 and Test-std from 77.5 to 76.8.

    Could not parse LaTeX table

    For captioning, MADTP improves CIDEr over UPop by 2.1 points at ratio 0.5 and 2.7 points at ratio 0.75. For VQA, the reported ratio-0.5 result achieves a 57% GFLOPs reduction with less than one percentage point of accuracy degradation.

  10. Knowl 10 — Component and hyperparameter effects

    empirical result

    The ablations on BLIP/NLVR2 use a target reduction ratio of 0.5. Combining all three TIS sources is better than using only self-attention, only MAG token attention, or only class attention. Removing MAG lowers dev/test accuracy from 81.97/82.85 to 79.65/80.96, while removing DTP lowers it to 80.83/81.44. The best tested MAG configuration uses K=100K=100 learnable tokens and channel dimension dk=768d_k=768; the max-keep mini-batch operation also outperforms mean-keep.

    Could not parse LaTeX table

    The module ablation indicates that MAG contributes 2.32 dev-accuracy points and 1.89 test-accuracy points relative to the corresponding model without MAG, while DTP contributes 1.14 and 1.41 points relative to the model without DTP. Max-keep determines the number of tokens pruned in a mini-batch using the instance with the highest inferred complexity.

Coverage note — No substantial contributed material was omitted; the page-8 qualitative token-mask visualizations are summarized with the NLVR2 results, while background, references, and appendix-level implementation details are excluded.

References

  1. 1.Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C. Lawrence Zitnick, and Devi Parikh. Vqa: Visual question answering. In IEEE International Conference on Computer Vision (ICCV), 2015.
  2. 2.Sajid Anwar, Kyuyeon Hwang, and Wonyong Sung. Structured pruning of deep convolutional neural networks. ACM Journal on Emerging Technologies in Computing Systems, 13(3), 2015.
  3. 3.Zhe Bian, Zhe Wang, Wenqiang Han, and Kangping Wang. Muti-scale and token mergence: Make your vit more efficient. arXiv preprint arXiv:2306.04897, 2023.
  4. 4.Davis Blalock, Jose Javier Gonzalez Ortiz, Jonathan Frankle, and John Guttag. What is the state of neural network pruning? Proceedings of machine learning and systems, 2:129–146, 2020.
  5. 5.Daniel Bolya, Cheng-Yang Fu, Xiaoliang Dai, Peizhao Zhang, Christoph Feichtenhofer, and Judy Hoffman. Token merging: Your vit but faster. arXiv preprint arXiv:2210.09461, 2022.
  6. 6.Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin-Tat Lee, Yuanzhi Li, Scott Lundberg, Harsha Nori, Hamid Palangi, MarcoTulio Ribeiro, and Yi Zhang. Sparks of artificial general intelligence: Early experiments with gpt-4. arXiv preprint arXiv:2303.12712, 2023.
  7. 7.Arnav Chavan, Zhiqiang Shen, Zhuang Liu, Zechun Liu, Kwang-Ting Cheng, and Eric Xing. Vision transformer slimming: Multi-dimension searching in continuous optimization space. In 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022.
  8. 8.Yu Cheng, Duo Wang, Pan Zhou, and Tao Zhang. A survey of model compression and acceleration for deep neural networks. arXiv preprint arXiv:1710.09282, 2017.
  9. 9.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North, 2019.
  10. 10.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929, 2020.
  11. 11.Yunfeng Fan, Wenchao Xu, Haozhao Wang, Junxiao Wang, and Song Guo. Pmr: Prototypical modal rebalance for multimodal learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 20029–20038, 2023.
  12. 12.Zhiyuan Fang, Jianfeng Wang, Xiaowei Hu, Lijuan Wang, Yezhou Yang, and Zicheng Liu. Compressing visual-linguistic model via knowledge distillation. arXiv: Computer Vision and Pattern Recognition,arXiv: Computer Vision and Pattern Recognition, 2021.
  13. 13.Amir Gholami, Sehoon Kim, Zhen Dong, Zhewei Yao, Michael W Mahoney, and Kurt Keutzer. A survey of quantization methods for efficient neural network inference. In Low-Power Computer Vision, pages 291–326. Chapman and Hall/CRC, 2022.
  14. 14.Mitchell A Gordon, Kevin Duh, and Nicholas Andrews. Compressing bert: Studying the effects of weight pruning on transfer learning. arXiv preprint arXiv:2002.08307, 2020.
  15. 15.Jianping Gou, Baosheng Yu, Stephen J Maybank, and Dacheng Tao. Knowledge distillation: A survey. International Journal of Computer Vision, 129:1789–1819, 2021.
  16. 16.Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh. Making the v in vqa matter: Elevating the role of image understanding in visual question answering. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 6904–6913, 2017.
  17. 17.Yangyang Guo, Haoyu Zhang, Liqiang Nie, Yongkang Wong, and Mohan Kankanhalli. Elip: Efficient language-image pre-training with fewer vision tokens. arXiv preprint arXiv:2309.16738, 2023.
  18. 18.Song Han, Huizi Mao, and William J Dally. Deep compression: Compressing deep neural networks with pruning, trained quantization and huffman coding. arXiv preprint arXiv:1510.00149, 2015.
  19. 19.Yizeng Han, Gao Huang, Shiji Song, Le Yang, Honghui Wang, and Yulin Wang. Dynamic neural networks: A survey. IEEE Transactions on Pattern Analysis and Machine Intelligence, 44(11):7436–7456, 2021.
  20. 20.Andrew G Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, Weijun Wang, Tobias Weyand, Marco Andreetto, and Hartwig Adam. Mobilenets: Efficient convolutional neural networks for mobile vision applications. arXiv preprint arXiv:1704.04861, 2017.
  21. 21.Forrest N Iandola, Song Han, Matthew W Moskewicz, Khalid Ashraf, William J Dally, and Kurt Keutzer. Squeezenet: Alexnet-level accuracy with 50x fewer parameters and¡ 0.5 mb model size. arXiv preprint arXiv:1602.07360, 2016.
  22. 22.X. Jia, E. Gavves, B. Fernando, and T. Tuytelaars. Guiding the long-short term memory model for image caption generation. In 2015 IEEE International Conference on Computer Vision (ICCV), pages 2407–2415. IEEE Computer Society, 2015.
  23. 23.Justin Johnson, Bharath Hariharan, Laurens van der Maaten, Li Fei-Fei, C. Lawrence Zitnick, and Ross Girshick. Clevr: A diagnostic dataset for compositional language and elementary visual reasoning. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017.
  24. 24.Wonjae Kim, Bokyung Son, and Ildoo Kim. Vilt: Vision-and-language transformer without convolution or region supervision. In International Conference on Machine Learning, pages 5583–5594. PMLR, 2021.
  25. 25.Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi. Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation. In ICML, 2022.
  26. 26.Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. arXiv preprint arXiv:2301.12597, 2023.
  27. 27.Youwei Liang, Chongjian Ge, Zhan Tong, Yibing Song, Jue Wang, and Pengtao Xie. Not all patches are what you need: Expediting vision transformers via token reorganizations. ICLR, 2022.
  28. 28.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollar, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In Computer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part V 13, pages 740–755. Springer, 2014.
  29. 29.Xiangcheng Liu, Tianyi Wu, and Guodong Guo. Adaptive sparse vit: Towards learnable adaptive token pruning by fully exploiting self-attention. In Proceedings of the Thirty-Second International Joint Conference on Artificial Intelligence, IJCAI-23, pages 1222–1230. International Joint Conferences on Artificial Intelligence Organization, 2023. Main Track.
  30. 30.Andre Martins and Ramon Astudillo. From softmax to sparsemax: A sparse model of attention and multi-label classification. In International conference on machine learning, pages 1614–1623. PMLR, 2016.
  31. 31.Lingchen Meng, Hengduo Li, Bor-Chun Chen, Shiyi Lan, Zuxuan Wu, Yu-Gang Jiang, and Ser-Nam Lim. Adavit: Adaptive vision transformers for efficient image recognition. In 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022.
  32. 32.Xiaokang Peng, Yake Wei, Andong Deng, Dong Wang, and Di Hu. Balanced multimodal learning via on-the-fly gradient modulation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8238–8247, 2022.
  33. 33.Alec Radford, JongWook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning transferable visual models from natural language supervision. Cornell University - arXiv,Cornell University - arXiv, 2021.
  34. 34.Yongming Rao, Wenliang Zhao, Benlin Liu, Jiwen Lu, Jie Zhou, and Cho-Jui Hsieh. Dynamicvit: Efficient vision transformers with dynamic token sparsification. Neural Information Processing Systems,Neural Information Processing Systems, 2021.
  35. 35.Dachuan Shi, Chaofan Tao, Ying Jin, Zhendong Yang, Chun Yuan, and Jiaqi Wang. UPop: Unified and progressive pruning for compressing vision-language transformers. In Proceedings of the 40th International Conference on Machine Learning, pages 31292–31311. PMLR, 2023.
  36. 36.Dachuan Shi, Chaofan Tao, Anyi Rao, Zhendong Yang, Chun Yuan, and Jiaqi Wang. Crossget: Cross-guided ensemble of tokens for accelerating vision-language transformers. arXiv preprint arXiv:2305.17455, 2023.
  37. 37.Pablo Sprechmann, Alexander M Bronstein, and Guillermo Sapiro. Learning efficient sparse and low rank models. IEEE transactions on pattern analysis and machine intelligence, 37(9):1821–1833, 2015.
  38. 38.Alane Suhr, Stephanie Zhou, Ally Zhang, Iris Zhang, Huajun Bai, and Yoav Artzi. A corpus for reasoning about natural language grounded in photographs. arXiv preprint arXiv:1811.00491, 2018.
  39. 39.Shengkun Tang, Yaqing Wang, Zhenglun Kong, Tianchi Zhang, Yao Li, Caiwen Ding, Yanzhi Wang, Yi Liang, and Dongkuan Xu. You need multiple exiting: Dynamic early exiting for accelerating unified vision language model. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 10781–10791, 2023.
  40. 40.Yehui Tang, Kai Han, Yunhe Wang, Chang Xu, Jianyuan Guo, Chao Xu, and Dacheng Tao. Patch slimming for efficient vision transformers. Cornell University - arXiv,Cornell University - arXiv, 2021.
  41. 41.Wenhan Xia, Hongxu Yin, Xiaoliang Dai, and N.K. Jha. Fully dynamic inference with deep neural networks. IEEE Transactions on Emerging Topics in Computing, PP:1–1, 2021.
  42. 42.Huanrui Yang, Hongxu Yin, Maying Shen, Pavlo Molchanov, Hai Li, and Jan Kautz. Global vision transformer pruning with hessian-aware saliency. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 18547–18557, 2023.
  43. 43.Hongxu Yin, Arash Vahdat, Jose Alvarez, Arun Mallya, Jan Kautz, and Pavlo Molchanov. A-ViT: Adaptive tokens for efficient vision transformer. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022.
  44. 44.Hongxu Yin, Arash Vahdat, Jose M. Alvarez, Arun Mallya, Jan Kautz, and Pavlo Molchanov. A-vit: Adaptive tokens for efficient vision transformer. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 10809–10818, 2022.
  45. 45.Peter Young, Alice Lai, Micah Hodosh, and Julia Hockenmaier. From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions. Transactions of the Association for Computational Linguistics, 2:67–78, 2014.
  46. 46.Lu Yu and Wei Xiang. X-pruner: explainable pruning for vision transformers. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 24355–24363, 2023.
  47. 47.Xiyu Yu, Tongliang Liu, Xinchao Wang, and Dacheng Tao. On compressing deep models by low rank and sparse decomposition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 7370–7379, 2017.
  48. 48.Chuanyang Zheng, Kai Zhang, Zhi Yang, Wenming Tan, Jun Xiao, Ye Ren, Shiliang Pu, et al. Savit: Structure-aware vision transformer pruning via collaborative optimization. Advances in Neural Information Processing Systems, 35:9010–9023, 2022.

Citation

MLA
Cao, J., et al. “MADTP: Multimodal Alignment-Guided Dynamic Token Pruning for Accelerating Vision-Language Transformer”. In Proc. IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2024, 2024, http://arxiv.org/abs/2403.02991v1.
APA
Cao, J., Ye, P., Li, S., Yu, C., Tang, Y., Lu, J., & Chen, T. (2024). MADTP: Multimodal Alignment-Guided Dynamic Token Pruning for Accelerating Vision-Language Transformer. In Proc. IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2024. http://arxiv.org/abs/2403.02991v1
Chicago
Cao, J., P. Ye, S. Li, et al. 2024. “MADTP: Multimodal Alignment-Guided Dynamic Token Pruning for Accelerating Vision-Language Transformer”. In Proc. IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2024. http://arxiv.org/abs/2403.02991v1.
Harvard
Cao, J. et al. (2024) “MADTP: Multimodal Alignment-Guided Dynamic Token Pruning for Accelerating Vision-Language Transformer”, In Proc. IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2024 [Preprint]. Available at: http://arxiv.org/abs/2403.02991v1.
Vancouver
1. Cao J, Ye P, Li S, Yu C, Tang Y, Lu J, Chen T (2024) MADTP: Multimodal Alignment-Guided Dynamic Token Pruning for Accelerating Vision-Language Transformer. In Proc. IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2024

BibTeX

@article{cao2024madtp,
  title = {MADTP: Multimodal Alignment-Guided Dynamic Token Pruning for Accelerating Vision-Language Transformer},
  author = {Cao, Jianjian and Ye, Peng and Li, Shengze and Yu, Chong and Tang, Yansong and Lu, Jiwen and Chen, Tao},
  year = {2024},
  journal = {In Proc. IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2024},
  url = {http://arxiv.org/abs/2403.02991v1},
  eprint = {2403.02991}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE