keyword
cross-modality fusion
Cross-modality fusion is the process in machine learning of combining, aligning, and synthesizing information or features extracted from multiple distinct data types into a unified representation. These heterogeneous modalities often include diverse sensory inputs and data formats such as text, audio, video, camera imagery, and spatial sensor measurements. By bridging structural and semantic differences across varied inputs, cross-modality fusion enables models to exploit complementary strengths, resolve conflicting signals, and capture rich inter-modal correlations that cannot be observed from any single modality alone. The integration can occur at different stages of computation, including early raw data aggregation, intermediate feature-level interactions via cross-attention or joint embedding spaces, and late decision-level combinations, ultimately improving model accuracy, robustness, and performance across complex multimodal tasks.
3 items

Learning Language-guided Adaptive Hyper-modality Representation for Multimodal Sentiment Analysis
Haoyu Zhang, Yu Wang, Guanghao Yin, Kejun Liu, Yuanyuan Liu, Tianshu Yu
Why you should read this
Proposes an adaptive language-guided transformer that suppresses conflicting and irrelevant visual and acoustic signals using multi-scale language cues, achieving state-of-the-art multimodal sentiment analysis performance across standard benchmarks like MOSI and MOSEI.
Though Multimodal Sentiment Analysis (MSA) proves effective by utilizing rich information from multiple sources (e.g., language, video, and audio), the potential sentiment-irrelevant and conflicting information across modalities may hinder the performance from being further improved. To alleviate this, we present Adaptive Language-guided Multimodal Transformer (ALMT), which incorporates an Adaptive Hyper-modality Learning (AHL) module to learn an irrelevance/conflict-suppressing representation from visual and audio features under the guidance of language features at different scales. With the obtained hyper-modality representation, the model can obtain a complementary and joint representation through multimodal fusion for effective MSA. In practice, ALMT achieves state-of-the-art performance on several popular datasets (e.g., MOSI, MOSEI and CH-SIMS) and an abundance of ablation demonstrates the validity and necessity of our irrelevance/conflict suppression mechanism.
Added
2026-10-01

Voxel Field Fusion for 3D Object Detection
Yanwei Li, Xiaojuan Qi, Yukang Chen, Liwei Wang, Zeming Li, Jian Sun, Jiaya Jia
Why you should read this
Proposes a cross-modality 3D object detection framework that maintains sensor consistency by projecting augmented camera features as rays into a voxel field with learnable sampling, achieving state-of-the-art results on the KITTI and nuScenes benchmarks.
In this work, we present a conceptually simple yet effective framework for cross-modality 3D object detection, named voxel field fusion. The proposed approach aims to maintain cross-modality consistency by representing and fusing augmented image features as a ray in the voxel field. To this end, the learnable sampler is first designed to sample vital features from the image plane that are projected to the voxel grid in a point-to-ray manner, which maintains the consistency in feature representation with spatial context. In addition, ray-wise fusion is conducted to fuse features with the supplemental context in the constructed voxel field. We further develop mixed augmentor to align feature-variant transformations, which bridges the modality gap in data augmentation. The proposed framework is demonstrated to achieve consistent gains in various benchmarks and outperforms previous fusion-based methods on KITTI and nuScenes datasets. Code is made available at https://github.com/dvlab-research/VFF.1
Added
2026-09-26

AVFF: Audio-Visual Feature Fusion for Video Deepfake Detection
Trevine Oorloff, Surya Koppisetti, Nicolò Bonettini, Divyaraj Solanki, Ben Colman, Yaser Yacoob, Ali Shahriyari, Gaurav Bharaj
Why you should read this
Proposes a two-stage deepfake detection framework that learns intrinsic cross-modal correspondences from real videos via complementary masking and feature fusion, achieving 98.6% accuracy on the FakeAVCeleb benchmark.
With the rapid growth in deepfake video content, we require improved and generalizable methods to detect them. Most existing detection methods either use uni-modal cues or rely on supervised training to capture the dissonance between the audio and visual modalities. While the former disregards the audio-visual correspondences entirely, the latter predominantly focuses on discerning audio-visual cues within the training corpus, thereby potentially overlooking correspondences that can help detect unseen deepfakes. We present Audio-Visual Feature Fusion (AVFF), a two-stage cross-modal learning method that explicitly captures the correspondence between the audio and visual modalities for improved deepfake detection. The first stage pursues representation learning via self-supervision on real videos to capture the intrinsic audio-visual correspondences. To extract rich cross-modal representations, we use contrastive learning and autoencoding objectives, and introduce a novel audio-visual complementary masking and feature fusion strategy. The learned representations are tuned in the second stage, where deepfake classification is pursued via supervised learning on both real and fake videos. Extensive experiments and analysis suggest that our novel representation learning paradigm is highly discriminative in nature. We report 98.6% accuracy and 99.1% AUC on the FakeAVCeleb dataset, outperforming the current audio-visual state-of-the-art by 14.9% and 9.9%, respectively.
Added
2026-09-26
