GeminiFusion: Efficient Pixel-wise Multimodal Fusion for Vision Transformer
Ding JiaJianyuan GuoKai HanHan WuChao ZhangChang XuXinghao Chen
Proposes GeminiFusion, a multimodal vision transformer framework that achieves linear computational complexity by combining intra-modal and inter-modal attention at aligned spatial positions, matching the efficiency of unimodal models while outperforming token exchange and full cross-attention across diverse segmentation, detection, and translation benchmarks.
Modern computer vision systems increasingly rely on multiple sensor inputs—such as standard color images, depth maps, light detection and ranging (LiDAR), and event data—to operate reliably in complex environments like autonomous driving and robotic navigation. While combining these different data sources enhances perception, current approaches face significant trade-offs. Cross-attention mechanisms capture rich interactions between modalities but incur heavy computational costs that grow quadratically with the number of input tokens, making them impractical for real-time deployment. Conversely, token-exchange methods save computation by replacing supposedly uninformative tokens with features from other modalities, but this pruning strategy often causes permanent data loss and underperforms.
The article introduces and evaluates GeminiFusion, an efficient multimodal fusion framework designed for vision transformers. The objective of the article is to demonstrate that restricting cross-modal attention to spatially aligned, pixel-wise token pairs preserves critical representations and achieves state-of-the-art accuracy while maintaining linear computational complexity.
The authors conducted extensive empirical evaluations across three core vision applications: multimodal semantic segmentation (using the NYUDv2, SUN RGB-D, and DeLiVER datasets), multimodal image-to-image translation (using the Taskonomy dataset), and 3D object detection (using the KITTI benchmark). The method integrates a lightweight relation discriminator to assess cross-modal disparity and injects layer-adaptive noise into the self-attention pathway to balance internal and cross-modal feature learning. Experiments evaluated standard architectures, including SegFormer and Swin Transformer encoders initialized with standard unimodal pre-training.
The evaluation produced several critical findings. First, GeminiFusion slashes the computational burden of standard cross-attention by 99.2%, reducing floating-point operations from over 17 billion to just 0.14 billion per instance by focusing strictly on spatially co-located patches. Second, in multimodal semantic segmentation, GeminiFusion consistently outperformed prior exchange-based methods, achieving performance gains across benchmarks, including a 3.4% increase in mean intersection over union when fusing four modalities (color, depth, event, and LiDAR) on the DeLiVER benchmark. Third, in image-to-image translation, the proposed approach reduced visual artifact error rates, showing a 12.6% relative improvement in the Fréchet Inception Distance metric on texture-to-RGB synthesis. Fourth, inserting the module into an existing 3D object detection framework (MVX-Net) boosted vehicle detection precision across test difficulty levels with negligible parameter overhead.
These findings indicate that multimodal vision models do not need to choose between computational efficiency and high accuracy. GeminiFusion achieves the low-latency profile of unimodal models while extracting the performance benefits of full cross-attention. This efficiency lowers memory and compute costs, making sophisticated multi-sensor perception feasible on hardware-constrained edge platforms and real-time systems. Furthermore, because the architecture works as a plug-and-play component compatible with standard pre-trained single-modality models, engineering teams can adopt it without expensive customized pre-training pipelines.
Organizations developing perception systems for autonomous driving, robotics, or spatial computing should consider piloting GeminiFusion within their existing vision transformer backbones. Development teams can also optimize latency by deploying the module selectively in later network layers, which retains competitive accuracy at even higher processing speeds. Further validation should test the architecture against noisy or degraded sensor inputs to evaluate operational safety under adverse field conditions.
The primary limitation of this approach is its requirement for spatial alignment among inputs; it is tailored for homogeneous, image-like modalities and currently cannot process unaligned heterogeneous combinations, such as paired audio and text. Within its intended scope of spatially registered visual sensors, the empirical evidence provides strong confidence in GeminiFusion's effectiveness and computational efficiency.
- Paper: Multimodal Token Fusion for Vision Transformers, Yikai Wang et al. (2022). TokenFusion establishes the token-replacement approach that GeminiFusion directly contrasts with, making its efficiency-versus-information-loss motivation and design choices clearer.
No sufficiently relevant recommendations were found.
