keyword
vision-language transformer
A vision-language transformer is a multimodal neural network architecture based on the transformer model that processes, aligns, and integrates visual and textual information within a unified computational framework. It tokenizes visual inputs, such as image patches or video frames, alongside textual sequences, using self-attention and cross-attention mechanisms to learn joint representations and capture correlations between text and imagery. Depending on the architecture, it may employ single-stream fusion, dual-encoder alignment, or modular cross-modal attention to bridge the semantic gap between visual and linguistic features. These models serve as the backbone for numerous multimodal artificial intelligence tasks, including image captioning, visual question answering, cross-modal retrieval, and zero-shot visual reasoning.
3 items

Rethinking Multimodal Entity and Relation Extraction from a Translation Point of View
Changmeng Zheng, Junhao Feng, Yi Cai, Xiaoyong Wei, Qing Li
Why you should read this
Proposes a translation-inspired framework that resolves text-image misalignment in multimodal entity and relation extraction by combining diffusion-based back-translation with divergence estimation to outperform 14 state-of-the-art baselines.
We revisit the multimodal entity and relation extraction from a translation point of view. Special attention is paid on the misalignment issue in text-image datasets which may mislead the learning. We are motivated by the fact that the cross-modal misalignment is a similar problem of cross-lingual divergence issue in machine translation. The problem can then be transformed and existing solutions can be borrowed by treating a text and its paired image as the translation to each other. We implement a multimodal back-translation using diffusion-based generative models for pseudo-paralleled pairs and a divergence estimator by constructing a high-resource corpora as a bridge for low-resource learners. Fine-grained confidence scores are generated to indicate both types and degrees of alignments with which better representations are obtained. The method has been validated in the experiments by outperforming 14 state-of-the-art methods in both entity and relation extraction tasks. The source code is available at https://github.com/thecharm/TMR.
Added
2026-10-05

Primitive Generation and Semantic-Related Alignment for Universal Zero-Shot Segmentation
Shuting He, Henghui Ding, Wei Jiang
Why you should read this
Proposes PADing, a unified universal zero-shot segmentation framework that bridges the cross-modal domain gap by assembling learned fine-grained primitives to synthesize unseen visual features and aligning their semantic-related components with linguistic class relationships.
We study universal zero-shot segmentation in this work to achieve panoptic, instance, and semantic segmentation for novel categories without any training samples. Such zero-shot segmentation ability relies on inter-class relationships in semantic space to transfer the visual knowledge learned from seen categories to unseen ones. Thus, it is desired to well bridge semantic and visual spaces and apply the semantic relationships to visual feature learning. We introduce a generative model to synthesize features for unseen categories, which links semantic and visual spaces as well as addresses the issue of lack of unseen training data. Furthermore, to mitigate the domain gap between semantic and visual spaces, firstly, we enhance the vanilla generator with learned primitives, each of which contains fine-grained attributes related to categories, and synthesize unseen features by selectively assembling these primitives. Secondly, we propose to disentangle the visual feature into the semantic-related part and the semantic-unrelated part that contains useful visual classification clues but is less relevant to semantic representation. The inter-class relationships of semantic-related visual features are then required to be aligned with those in semantic space, thereby transferring semantic knowledge to visual feature learning. The proposed approach achieves impressively state-of-the-art performance on zero-shot panoptic segmentation, instance segmentation, and semantic segmentation.
Added
2026-09-26

MADTP: Multimodal Alignment-Guided Dynamic Token Pruning for Accelerating Vision-Language Transformer
Jianjian Cao, Peng Ye, Shengze Li, Chong Yu, Yansong Tang, Jiwen Lu, Tao Chen
Why you should read this
Proposes a multimodal alignment-guided dynamic token pruning framework that cuts Vision-Language Transformer computation by up to 80% with minimal accuracy loss by aligning cross-modal representations to prevent false token removal and adaptively tuning layer-wise pruning ratios per input instance.
Vision-Language Transformers (VLTs) have shown great success recently, but are meanwhile accompanied by heavy computation costs, where a major reason can be attributed to the large number of visual and language tokens. Existing token pruning research for compressing VLTs mainly follows a single-modality-based scheme yet ignores the critical role of aligning different modalities for guiding the token pruning process, causing the important tokens for one modality to be falsely pruned in another modality branch. Meanwhile, existing VLT pruning works also lack the flexibility to dynamically compress each layer based on different input samples. To this end, we propose a novel framework named Multimodal Alignment-Guided Dynamic Token Pruning (MADTP) for accelerating various VLTs. Specifically, we first introduce a well-designed Multi-modality Alignment Guidance (MAG) module that can align features of the same semantic concept from different modalities, to ensure the pruned tokens are less important for all modalities. We further design a novel Dynamic Token Pruning (DTP) module, which can adaptively adjust the token compression ratio in each layer based on different input instances. Extensive experiments on various benchmarks demonstrate that MADTP significantly reduces the computational complexity of kinds of multimodal models while preserving competitive performance. Notably, when applied to the BLIP model in the NLVR2 dataset, MADTP can reduce the GFLOPs by 80% with less than 4% performance degradation. The code is available at https://github.com/double125/MADTP.
Added
2026-09-26
