MADTP: Multimodal Alignment-Guided Dynamic Token Pruning for Accelerating Vision-Language Transformer
Jianjian CaoPeng YeShengze LiChong YuYansong TangJiwen LuTao Chen
Proposes a multimodal alignment-guided dynamic token pruning framework that cuts Vision-Language Transformer computation by up to 80% with minimal accuracy loss by aligning cross-modal representations to prevent false token removal and adaptively tuning layer-wise pruning ratios per input instance.
Vision-Language Transformers power many state-of-the-art multimodal artificial intelligence applications, including visual reasoning, image captioning, image-text retrieval, and visual question answering. However, processing large numbers of visual and textual tokens creates significant computational overhead, which limits real-time deployment and inflates inference costs. Existing token-reduction methods either prune tokens within single modalities independently or apply static reduction ratios. These approaches frequently discard visual details that are critical for text comprehension (or vice versa) and fail to adjust processing power based on input complexity.
The article introduces and evaluates Multimodal Alignment-Guided Dynamic Token Pruning (MADTP), a compression framework designed to accelerate multimodal models by aligning visual and language features before pruning and dynamically adjusting computational effort per input.
The framework introduces two primary mechanisms: a Multi-modality Alignment Guidance module and a Dynamic Token Pruning module. The alignment module uses shared learnable tokens to map image and text features into a common semantic space, ensuring that retained tokens are mutually relevant across modalities. The dynamic pruning module scores each token by combining class-level, self-attention, and cross-modal attention scores, and then removes less relevant tokens using an adaptive, instance-specific threshold. The authors validated MADTP on established benchmark architectures (BLIP and CLIP) across four major datasets: NLVR2 for visual reasoning, COCO and Flickr30k for retrieval and captioning, and VQA v2.0 for visual question answering.
The evaluation yielded several key findings. First, MADTP achieved aggressive computation cuts with minimal accuracy loss: on the BLIP model performing visual reasoning, it reduced floating-point operations by 80% while retaining test accuracy within 3.86% of the uncompressed model. Second, in image-text retrieval tasks at high compression ratios (around 75% computation reduction), MADTP substantially outperformed prior methods, improving image-to-text recall@1 on the COCO dataset by 10.1 percentage points over the state-of-the-art baseline. Third, on visual question answering tasks, MADTP achieved a 57% reduction in computational complexity with less than a 1% drop in accuracy. Finally, ablation experiments confirmed that cross-modal feature alignment contributed roughly a 2% improvement in performance compared to pruning without multimodal guidance.
These findings indicate that cross-modal alignment is critical for compressing multimodal systems without degrading accuracy. In practical terms, reducing computational demands by 50% to 80% directly lowers cloud inference expenses, reduces latency, and facilitates deployment of high-performing vision-language models on edge hardware and resource-constrained environments.
Organizations deploying vision-language models should consider adopting multimodal alignment-guided dynamic pruning as a standard optimization pipeline. Prior to broad deployment, engineering teams should benchmark MADTP in application-specific pilot pipelines to determine the optimal trade-off between reduction ratios and task-specific accuracy requirements.
While the results demonstrate robust performance across multiple standard benchmarks and architectures, the framework requires tuning hyperparameters such as the number and dimension of learnable tokens. Overall confidence in the reported efficiency and accuracy improvements is high for the tested tasks, though practitioners should validate performance on complex domain-specific tasks and hardware architectures.
- Paper: BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation, Junnan Li et al. (2022). Introduces the BLIP framework and multi-modal alignment concepts that MADTP directly targets for dynamic token compression and acceleration.
- Paper: BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models, Junnan Li et al. (2023). Establishes modern vision-language interaction using Q-Former bottlenecks, providing essential background on the cross-modal token representations pruned by MADTP.
- Paper: Align before Fuse: Vision and Language Representation Learning with Momentum Distillation, Junnan Li et al. (2021). Introduces the align-before-fuse vision-language paradigm that MADTP adapts to prevent false pruning across cross-modal branches.
- Paper: AdaViT: Adaptive Vision Transformers for Efficient Image Recognition, Lingchen Meng et al. (2022). Provides the foundational dynamic token and block skipping mechanism using input-dependent gating in vision transformers.
- Paper: ViLT: Vision-and-Language Transformer Without Convolution or Region Supervision, Wonjae Kim et al. (2021). Presents a unified token-based vision-and-language transformer architecture where joint text and visual tokens create the efficiency bottlenecks tackled by MADTP.
- Paper: Dynamic Context Pruning for Efficient and Interpretable Autoregressive Transformers, Sotiris Anagnostidis et al. (2023). Formulates dynamic instance-specific context token pruning in autoregressive transformer layers.
- Paper: Tokens-to-Token ViT: Training Vision Transformers from Scratch on ImageNet, Li Yuan et al. (2021). Demonstrates early methods for progressive visual token aggregation and reduction to eliminate redundancy in transformer layers.
- Paper: VisionZip: Longer is Better but Not Necessary in Vision Language Models, Senqiao Yang et al. (2025). Extends visual token reduction techniques to modern high-resolution vision-language models via attention-guided dominant token selection and merging.
- Paper: Filter, Correlate, Compress: Training-Free Token Reduction for MLLM Acceleration, Yuhang Han et al. (2026). Builds on multimodal token reduction by introducing a training-free filter-correlate-compress framework across encoder and LLM decoder stages.
- Paper: Global Compression Commander: Plug-and-Play Inference Acceleration for High-Resolution Large Vision-Language Models, Xuyang Liu et al. (2026). Applies hierarchical and dynamic token compression to high-resolution vision-language models employing global thumbnail and local crop representations.
- Paper: InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models, Jinguo Zhu et al. (2025). Explores scaled native multimodal pre-training and variable visual context encoding in next-generation vision-language foundation models.
