keyword
cross-modal attention
Cross-modal attention is a computational mechanism in multimodal machine learning models that allows information from one data type, such as text, images, or audio, to dynamically reference and integrate relevant elements from another distinct data type. Primarily implemented within neural network architectures such as transformers, it computes attention scores across different modalities by mapping queries derived from one modality against the keys and values of another. By identifying fine-grained semantic correlations and weighing complementary features across heterogeneous inputs, cross-modal attention facilitates multimodal feature fusion and joint reasoning for tasks such as visual question answering, referring image segmentation, audio-visual event localization, and cross-modal retrieval.
8 items

VLMo: Unified Vision-Language Pre-Training with Mixture-of-Modality-Experts
Hangbo Bao, Wenhui Wang, Li Dong, Qiang Liu, Owais Khan Mohammed, Kriti Aggarwal, Subhojit Som, Songhao Piao, Furu Wei
Why you should read this
Introduces a modular vision-language pre-training framework that uses modality-specific feed-forward experts and shared self-attention to unify dual-encoder retrieval speed with fusion-encoder classification accuracy.
We present a unified Vision-Language pretrained Model (VLMo) that jointly learns a dual encoder and a fusion encoder with a modular Transformer network. Specifically, we introduce Multiway Transformer, where each block contains a pool of modality-specific experts and a shared self-attention layer. Because of the modeling flexibility of Multiway Transformer, pretrained VLMo can be fine-tuned as a fusion encoder for vision-language classification tasks, or used as a dual encoder for efficient image-text retrieval. Moreover, we propose a stagewise pre-training strategy, which effectively leverages large-scale image-only and text-only data besides image-text pairs. Experimental results show that VLMo achieves state-of-the-art results on various vision-language tasks, including VQA, NLVR2 and image-text retrieval. The code and pretrained models are available at http://aka.ms/vlmo.
Added
2026-10-05

Conversation Understanding using Relational Temporal Graph Neural Networks with Auxiliary Cross-Modality Interaction
Cam-Van Thi Nguyen, Anh-Tuan Mai, The-Son Le, Hai-Dang Kieu, Duc-Trong Le
Why you should read this
Proposes a relational temporal graph neural network framework, CORECT, that jointly models utterance-level temporal dependencies and conversation-level cross-modal interactions while preserving modality-specific representations to achieve state-of-the-art multimodal emotion recognition on IEMOCAP and CMU-MOSEI.
Emotion recognition is a crucial task for human conversation understanding. It becomes more challenging with the notion of multimodal data, e.g., language, voice, and facial expressions. As a typical solution, the global and local context information are exploited to predict the emotional label for every single sentence, i.e., utterance, in the dialogue. Specifically, the global representation could be captured via modeling of cross-modal interactions at the conversation level. The local one is often inferred using the temporal information of speakers or emotional shifts, which neglects vital factors at the utterance level. Additionally, most existing approaches take fused features of multiple modalities in a unified input without leveraging modality-specific representations. Motivating from these problems, we propose the Relational Temporal Graph Neural Network with Auxiliary Cross-Modality Interaction (CORECT), an novel neural network framework that effectively captures conversation-level cross-modality interactions and utterance-level temporal dependencies with the modality-specific manner for conversation understanding. Extensive experiments demonstrate the effectiveness of CORECT via its state-of-the-art results on the IEMOCAP and CMU-MOSEI datasets for the multimodal ERC task.
Added
2026-10-03

Why is Winoground Hard? Investigating Failures in Visuolinguistic Compositionality
Anuj Diwan, Layne Berry, Eunsol Choi, David Harwath, Kyle Mahowald
Why you should read this
Reveals that vision-language model failures on the Winoground compositional benchmark stem primarily from cross-modal representation fusion and atypical visual reasoning demands rather than deficits in linguistic understanding.
Recent visuolinguistic pre-trained models show promising progress on various end tasks such as image retrieval and video captioning. Yet, they fail miserably on the recently proposed Winoground dataset (Thrush et al., 2022), which challenges models to match paired images and English captions, with items constructed to overlap lexically but differ in meaning (e.g., “there is a mug in some grass” vs. “there is some grass in a mug”). By annotating the dataset using new fine-grained tags, we show that solving the Winoground task requires not just compositional language understanding, but a host of other abilities like commonsense reasoning or locating small, out-of-focus objects in low-resolution images. In this paper, we identify the dataset’s main challenges through a suite of experiments on related tasks (probing task, image retrieval task), data augmentation, and manual inspection of the dataset. Our analysis suggests that a main challenge in visuolinguistic models may lie in fusing visual and textual representations, rather than in compositional language understanding. We release our annotation and code at https://github.com/ajd12342/why-winoground-hard.
Added
2026-10-03

Multimodal Multi-loss Fusion Network for Sentiment Analysis
Zehui Wu, Ziwei Gong, Jaywon Koo, Julia Hirschberg
Why you should read this
Proposes an end-to-end multimodal fusion framework that combines cross-modal attention, self-attention, and auxiliary single-modality training losses to achieve state-of-the-art sentiment analysis performance on CMU-MOSI, CMU-MOSEI, and CH-SIMS benchmarks.
Sentiment analysis has become increasingly important due to the rapid growth of social media data. Recent research has shown that multimodal approaches which combine text, acoustic, and visual modalities can outperform unimodal methods. In particular, fusion strategies such as early fusion, late fusion, and model-level fusion have been explored extensively. However, effectively leveraging multiple losses to guide the training process remains under-explored. To address this gap, we propose a Multimodal Multi-loss Fusion Network (MMLFN). Our approach incorporates three different loss functions—focal loss, contrastive loss, and center loss—into a unified framework. We evaluate MMLFN on several benchmark datasets, achieving state-of-the-art performance. Extensive experiments demonstrate that our method significantly improves upon existing techniques. Moreover, comprehensive ablation studies verify the effectiveness of each component in our design. With joint optimization using complementary information across modalities as well as diverse objectives via distinct supervision signals during learning phase yields robust representations leading toward superior results compared prior art hence benefit...
Added
2026-10-02

LAVT: Language-Aware Vision Transformer for Referring Image Segmentation
Zhao Yang, Jiaqi Wang, Yansong Tang, Kai Chen, Hengshuang Zhao, Philip H. S. Torr
Why you should read this
Proposes an early fusion framework that integrates linguistic cues directly into the intermediate layers of a vision Transformer encoder, replacing complex decoders with a lightweight predictor to set new performance standards on referring image segmentation benchmarks.
Referring image segmentation is a fundamental vision-language task that aims to segment out an object referred to by a natural language expression from an image. One of the key challenges behind this task is leveraging the referring expression for highlighting relevant positions in the image. A paradigm for tackling this problem is to leverage a powerful vision-language (“cross-modal”) decoder to fuse features independently extracted from a vision encoder and a language encoder. Recent methods have made remarkable advancements in this paradigm by exploiting Transformers as cross-modal decoders, concurrent to the Transformer’s overwhelming success in many other vision-language tasks. Adopting a different approach in this work, we show that significantly better cross-modal alignments can be achieved through the early fusion of linguistic and visual features in intermediate layers of a vision Transformer encoder network. By conducting cross-modal feature fusion in the visual feature encoding stage, we can leverage the well-proven correlation modeling power of a Transformer encoder for excavating helpful multi-modal context. This way, accurate segmentation results are readily harvested with a light-weight mask predictor. Without bells and whistles, our method surpasses the previous state-of-the-art methods on RefCOCO, RefCOCO+, and G-Ref by large margins.
Added
2026-09-26

Compositional Temporal Grounding with Structured Variational Cross-Graph Correspondence Learning
Juncheng Li, Junlin Xie, Long Qian, Linchao Zhu, Siliang Tang, Fei Wu, Yi Yang, Yueting Zhuang, Xin Eric Wang
Why you should read this
Introduces compositional temporal grounding benchmarks alongside a hierarchical variational cross-graph reasoning framework that aligns multi-level video and language structures to accurately localize video segments described by novel word combinations.
Temporal grounding in videos aims to localize one target video segment that semantically corresponds to a given query sentence. Thanks to the semantic diversity of natural language descriptions, temporal grounding allows activity grounding beyond pre-defined classes and has received increasing attention in recent years. The semantic diversity is rooted in the principle of compositionality in linguistics, where novel semantics can be systematically described by combining known words in novel ways (compositional generalization). However, current temporal grounding datasets do not specifically test for the compositional generalizability. To systematically measure the compositional generalizability of temporal grounding models, we introduce a new Compositional Temporal Grounding task and construct two new dataset splits, i.e., Charades-CG and ActivityNet-CG. Evaluating the state-of-the-art methods on our new dataset splits, we empirically find that they fail to generalize to queries with novel combinations of seen words. To tackle this challenge, we propose a variational cross-graph reasoning framework that explicitly decomposes video and language into multiple structured hierarchies and learns fine-grained semantic correspondence among them. Experiments illustrate the superior compositional generalizability of our approach. The repository of this work is at https://github.com/YYJMJC/Compositional-Temporal-Grounding.
Added
2026-09-26

Collecting Cross-Modal Presence-Absence Evidence for Weakly-Supervised Audio- Visual Event Perception
Junyu Gao, Mengyuan Chen, Changsheng Xu
Why you should read this
Proposes an evidential learning framework that extracts presence evidence from the primary modality and absence evidence from the complementary modality to accurately localize unsynchronized audio and visual events using only video-level labels.
With only video-level event labels, this paper targets at the task of weakly-supervised audio-visual event perception (WS-AVEP), which aims to temporally localize and categorize events belonging to each modality. Despite the recent progress, most existing approaches either ignore the unsynchronized property of audio-visual tracks or discount the complementary modality for explicit enhancement. We argue that, for an event residing in one modality, the modality itself should provide ample presence evidence of this event, while the other complementary modality is encouraged to afford the absence evidence as a reference signal. To this end, we propose to collect Cross-Modal Presence-Absence Evidence (CMPAE) in a unified framework. Specifically, by leveraging uni-modal and cross-modal representations, a presence-absence evidence collector (PAEC) is designed under Subjective Logic theory. To learn the evidence in a reliable range, we propose a joint-modal mutual learning (JML) process, which calibrates the evidence of diverse audible, visible, and audi-visible events adaptively and dynamically. Extensive experiments show that our method surpasses state-of-the-arts (e.g., absolute gains of 3.6% and 6.1% in terms of event-level visual and audio metrics). Code is available in github.com/MengyuanChen21/CVPR2023-CMPAE.
Added
2026-09-26

DF-GAN: A Simple and Effective Baseline for Text-to-Image Synthesis
Ming Tao, Hao Tang, Fei Wu, Xiaoyuan Jing, Bing-Kun Bao, Changsheng Xu
Why you should read this
Proposes a simple one-stage text-to-image GAN architecture that directly synthesizes high-resolution images using deep feature fusion blocks and a matching-aware discriminator, outperforming complex multi-stage models without relying on auxiliary networks.
Synthesizing high-quality realistic images from text descriptions is a challenging task. Existing text-to-image Generative Adversarial Networks generally employ a stacked architecture as the backbone yet still remain three flaws. First, the stacked architecture introduces the entanglements between generators of different image scales. Second, existing studies prefer to apply and fix extra networks in adversarial learning for text-image semantic consistency, which limits the supervision capability of these networks. Third, the cross-modal attention-based text-image fusion that widely adopted by previous works is limited on several special image scales because of the computational cost. To these ends, we propose a simpler but more effective Deep Fusion Generative Adversarial Networks (DF-GAN). To be specific, we propose: (i) a novel one-stage text-to-image backbone that directly synthesizes high-resolution images without entanglements between different generators, (ii) a novel Target-Aware Discriminator composed of Matching-Aware Gradient Penalty and One-Way Output, which enhances the text-image semantic consistency without introducing extra networks, (iii) a novel deep text-image fusion block, which deepens the fusion process to make a full fusion between text and visual features. Compared with current state-of-the-art methods, our proposed DF-GAN is simpler but more efficient to synthesize realistic and text-matching images and achieves better performance on widely used datasets. Code is available at https://github.com/tobran/DF-GAN.
Added
2026-09-26
