GeminiFusion: Efficient Pixel-wise Multimodal Fusion for Vision Transformer

Ding JiaJianyuan GuoKai HanHan WuChao ZhangChang XuXinghao Chen

article2024ICML61 citations

Proposes GeminiFusion, a multimodal vision transformer framework that achieves linear computational complexity by combining intra-modal and inter-modal attention at aligned spatial positions, matching the efficiency of unimodal models while outperforming token exchange and full cross-attention across diverse segmentation, detection, and translation benchmarks.

Listen

Modern computer vision systems increasingly rely on multiple sensor inputs—such as standard color images, depth maps, light detection and ranging (LiDAR), and event data—to operate reliably in complex environments like autonomous driving and robotic navigation. While combining these different data sources enhances perception, current approaches face significant trade-offs. Cross-attention mechanisms capture rich interactions between modalities but incur heavy computational costs that grow quadratically with the number of input tokens, making them impractical for real-time deployment. Conversely, token-exchange methods save computation by replacing supposedly uninformative tokens with features from other modalities, but this pruning strategy often causes permanent data loss and underperforms.

The article introduces and evaluates GeminiFusion, an efficient multimodal fusion framework designed for vision transformers. The objective of the article is to demonstrate that restricting cross-modal attention to spatially aligned, pixel-wise token pairs preserves critical representations and achieves state-of-the-art accuracy while maintaining linear computational complexity.

The authors conducted extensive empirical evaluations across three core vision applications: multimodal semantic segmentation (using the NYUDv2, SUN RGB-D, and DeLiVER datasets), multimodal image-to-image translation (using the Taskonomy dataset), and 3D object detection (using the KITTI benchmark). The method integrates a lightweight relation discriminator to assess cross-modal disparity and injects layer-adaptive noise into the self-attention pathway to balance internal and cross-modal feature learning. Experiments evaluated standard architectures, including SegFormer and Swin Transformer encoders initialized with standard unimodal pre-training.

The evaluation produced several critical findings. First, GeminiFusion slashes the computational burden of standard cross-attention by 99.2%, reducing floating-point operations from over 17 billion to just 0.14 billion per instance by focusing strictly on spatially co-located patches. Second, in multimodal semantic segmentation, GeminiFusion consistently outperformed prior exchange-based methods, achieving performance gains across benchmarks, including a 3.4% increase in mean intersection over union when fusing four modalities (color, depth, event, and LiDAR) on the DeLiVER benchmark. Third, in image-to-image translation, the proposed approach reduced visual artifact error rates, showing a 12.6% relative improvement in the Fréchet Inception Distance metric on texture-to-RGB synthesis. Fourth, inserting the module into an existing 3D object detection framework (MVX-Net) boosted vehicle detection precision across test difficulty levels with negligible parameter overhead.

These findings indicate that multimodal vision models do not need to choose between computational efficiency and high accuracy. GeminiFusion achieves the low-latency profile of unimodal models while extracting the performance benefits of full cross-attention. This efficiency lowers memory and compute costs, making sophisticated multi-sensor perception feasible on hardware-constrained edge platforms and real-time systems. Furthermore, because the architecture works as a plug-and-play component compatible with standard pre-trained single-modality models, engineering teams can adopt it without expensive customized pre-training pipelines.

Organizations developing perception systems for autonomous driving, robotics, or spatial computing should consider piloting GeminiFusion within their existing vision transformer backbones. Development teams can also optimize latency by deploying the module selectively in later network layers, which retains competitive accuracy at even higher processing speeds. Further validation should test the architecture against noisy or degraded sensor inputs to evaluate operational safety under adverse field conditions.

The primary limitation of this approach is its requirement for spatial alignment among inputs; it is tailored for homogeneous, image-like modalities and currently cannot process unaligned heterogeneous combinations, such as paired audio and text. Within its intended scope of spatially registered visual sensors, the empirical evidence provides strong confidence in GeminiFusion's effectiveness and computational efficiency.

Jia et al (2024).pdf
  • Paper: Multimodal Token Fusion for Vision Transformers, Yikai Wang et al. (2022). TokenFusion establishes the token-replacement approach that GeminiFusion directly contrasts with, making its efficiency-versus-information-loss motivation and design choices clearer.

No sufficiently relevant recommendations were found.

Cover for GeminiFusion: Efficient Pixel-wise Multimodal Fusion for Vision Transformer

Abstract

Cross-modal transformers have demonstrated superiority in various vision tasks by effectively integrating different modalities. This paper first critiques prior token exchange methods which replace less informative tokens with inter-modal features, and demonstrate exchange based methods underperform cross-attention mechanisms, while the computational demand of the latter inevitably restricts its use with longer sequences. To surmount the computational challenges, we propose GeminiFusion, a pixel-wise fusion approach that capitalizes on aligned cross-modal representations. GeminiFusion elegantly combines intra-modal and inter-modal attentions, dynamically integrating complementary information across modalities. We employ a layer-adaptive noise to adaptively control their interplay on a per-layer basis, thereby achieving a harmonized fusion process. Notably, GeminiFusion maintains linear complexity with respect to the number of input tokens, ensuring this multimodal framework operates with efficiency comparable to unimodal networks. Comprehensive evaluations across multimodal image-to-image translation, 3D object detection and arbitrary-modal semantic segmentation tasks, including RGB, depth, LiDAR, event data, etc. demonstrate the superior performance of our GeminiFusion against leading-edge techniques. The PyTorch code is available here.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Our Method
  • 3.1. Fusion via exchange
  • 3.2. Fusion via cross-attention
  • 3.3. GeminiFusion: pixel-wise fusion module
  • 3.4. Overall architecture
  • 4. Experiment
  • 4.1. Datasets
  • 4.2. Comparisons with TokenFusion
  • 4.3. Applying to Swin Transformer
  • 4.4. Comparisons with state-of-the-art methods
  • 4.5. Effect of each component on GeminiFusion
  • 4.6. Discussion on Inference Latency
  • 4.7. 3D Object Detection task
  • 5. Conclusion
  • Acknowledgements
  • Impact Statement
  • References
  • A. Implementation Details
  • B. More Visualization Results

Knowls

  1. Knowl 1 — Pixel-wise GeminiFusion attention

    model/method

    GeminiFusion fuses two spatially aligned modalities by attending only between the two tokens at the same spatial location, rather than attending across all locations. For location ii and layer ℓ\ell, let xm,i∈Rdx_{m,i}\in\mathbb{R}^d be the feature vector from modality m∈{1,2}m\in\{1,2\}, and let n=3−mn=3-m denote the other modality. The relation discriminator ϕ\phi produces a scalar ρm,i∈[0,1]\rho_{m,i}\in[0,1] from the two co-located features. The vectors ηℓK,ηℓV∈Rd\eta^K_\ell,\eta^V_\ell\in\mathbb{R}^d are learned, layer-specific noise parameters; WQ,WK,WV∈Rd×dW_Q,W_K,W_V\in\mathbb{R}^{d\times d} are projection matrices.

    qm,i=xm,iWQ,Km,i=[(xm,i+ηℓK)WKρm,ixm,iWK],Vm,i=[(xm,i+ηℓV)WVxn,iWV],q_{m,i}=x_{m,i}W_Q,\qquad K_{m,i}=\begin{bmatrix}(x_{m,i}+\eta^K_\ell)W_K\\\rho_{m,i}x_{m,i}W_K\end{bmatrix},\qquad V_{m,i}=\begin{bmatrix}(x_{m,i}+\eta^V_\ell)W_V\\x_{n,i}W_V\end{bmatrix}, ym,i=softmax⁡ ⁣(qm,iKm,iTd)Vm,i+xm,i.y_{m,i}=\operatorname{softmax}\!\left(\frac{q_{m,i}K_{m,i}^{\mathsf T}}{\sqrt d}\right)V_{m,i}+x_{m,i}.

    The softmax is over the two key-value entries, and the residual retains the input feature. The relation score modulates the self-derived key associated with the cross-modal value; the attention can therefore combine a modality's own representation with information from its counterpart without discarding the original feature. The paper gives the relation discriminator as a concatenation of the two features followed by a lightweight convolution and softmax. For NN aligned token locations and channel width dd, the reported complexity is O(Nd2)O(Nd^2), linear in token count, versus quadratic token-count complexity for full cross-attention. For 16,384 patches, the paper reports a reduction from over 17G to 0.14G FLOPs for one attention operation, or 99.2%.

  2. Knowl 2 — Ablations identify effective GeminiFusion components

    empirical result

    On NYUDv2 semantic segmentation with a MiT-B3 backbone and aligned training epochs, adding GeminiFusion components successively raised mIoU from 53.3 for the baseline to 55.4 with point-wise cross-attention (PWC), 56.3 after adding noised self-attention (NSA), and 56.8 after adding the attention relation discriminator (ARD). The noise ablation favored learned additive noise: on pixel accuracy / mean accuracy / mIoU, random Gaussian noise multiplied into features scored 79.2 / 69.3 / 55.5; random Gaussian noise added scored 79.2 / 68.8 / 55.3; learned noise multiplied scored 79.6 / 69.2 / 56.2; and learned noise added scored 79.9 / 69.9 / 56.8. For the relation-discriminator choices, a two-layer MLP followed by softmax scored 79.9 / 69.9 / 56.8, compared with 79.2 / 69.2 / 55.7 for a 1×1 CNN followed by softmax. These ablations support combining point-wise cross-modal exchange with self-attention, learned additive noise, and relation-based modulation.

  3. Knowl 3 — Token exchange can discard useful modality-specific information

    empirical result

    The paper's analysis of TokenFusion challenges the assumption that its learned masks reliably identify uninformative tokens. In the observed shallow layers, tokens were often all exchanged, including when the mask threshold was changed; deeper layers showed more selective exchange. In threshold experiments on NYUDv2 and SUN RGB-D segmentation, setting the threshold so that all tokens were exchanged almost invariably performed better than selective exchange. The authors interpret this as evidence that direct replacement can irreversibly remove information unique to a token or modality, and that the mask-generation process does not consistently select only negligible features. Their comparison experiments also found GeminiFusion to outperform TokenFusion under aligned training conditions.

  4. Knowl 4 — Four-stage shared-encoder architecture for aligned visual modalities

    model/method

    The semantic-segmentation implementation uses a four-stage encoder-decoder modeled on SegFormer. RGB and other image-like visual modalities, such as depth, event data, and LiDAR, are processed using shared parameters except in Layer Normalization layers. Within each stage, each modality is refined by multi-head self-attention and feed-forward blocks, and GeminiFusion integrates the aligned features. The four stages use respectively 4, 8, 16, and 32 blocks, strides of 4, 8, 16, and 32, and channel dimensions of 64, 128, 320, and 512. Features from the modalities are combined by weighted summation at each stage; an MLP-based decoder then produces segmentation predictions.

  5. Knowl 5 — Segmentation gains over TokenFusion across modality combinations

    data/table

    With training epochs aligned, GeminiFusion achieved higher semantic-segmentation scores than TokenFusion for every reported comparison. The metrics are pixel accuracy, mean accuracy, and mean IoU (all percentages); DeLiVER reports only mIoU. The results cover RGB plus depth on NYUDv2 and SUN RGB-D, and several two- and four-modality combinations on DeLiVER.

    Dataset and inputBackboneMethodPixel accuracyMean accuracymIoU
    NYUDv2, RGB+DMiT-B3TokenFusion79.066.954.2
    NYUDv2, RGB+DMiT-B3GeminiFusion79.969.956.8
    NYUDv2, RGB+DMiT-B5TokenFusion†79.167.555.1
    NYUDv2, RGB+DMiT-B5GeminiFusion80.370.457.7
    SUN RGB-D, RGB+DMiT-B3TokenFusion†82.863.651.4
    SUN RGB-D, RGB+DMiT-B3GeminiFusion83.364.652.7
    SUN RGB-D, RGB+DMiT-B5TokenFusion†83.163.951.8
    SUN RGB-D, RGB+DMiT-B5GeminiFusion83.865.353.3
    DeLiVER, RGB+DMiT-B2TokenFusion†——63.7
    DeLiVER, RGB+DMiT-B2GeminiFusion——66.4
    DeLiVER, RGB+EMiT-B2TokenFusion†——55.7
    DeLiVER, RGB+EMiT-B2GeminiFusion——58.5
    DeLiVER, RGB+LMiT-B2TokenFusion†——55.5
    DeLiVER, RGB+LMiT-B2GeminiFusion——58.6
    DeLiVER, RGB+D+E+LMiT-B2TokenFusion†——63.5
    DeLiVER, RGB+D+E+LMiT-B2GeminiFusion——66.9

    Here D denotes depth, E event data, and L LiDAR; † denotes TokenFusion results reproduced by the authors. The mIoU improvements range from 1.3 points on SUN RGB-D with MiT-B3 to 3.4 points on DeLiVER with all four modalities.

  6. Knowl 6 — Benchmark results for semantic segmentation

    empirical result

    On the NYUDv2, SUN RGB-D, and DeLiVER semantic-segmentation benchmarks, GeminiFusion with MiT-B5 (MiT-B2 on DeLiVER) reported mIoU scores of 57.7, 53.3, and 66.9, respectively, without additional strategies beyond ImageNet classification pretraining. For comparison, CMNeXt reported 56.9, 50.4, and 66.3 with MiT-B4 (MiT-B2 on DeLiVER), while CMX with MiT-B5 reported 56.9 and 52.4 on NYUDv2 and SUN RGB-D, and 62.7 on DeLiVER. A GeminiFusion model using a Swin-Large-22k encoder reported 60.2 mIoU on NYUDv2 and 54.6 on SUN RGB-D. When the SUN RGB-D-trained Swin-Large-22k model was additionally used to initialize NYUDv2 training, the NYUDv2 result was 60.9 mIoU; this latter result uses additional pretraining and is not the same training setting as the results without additional strategies.

  7. Knowl 7 — Image-to-image translation improves across four Taskonomy tasks

    data/table

    On Taskonomy image-to-image translation, GeminiFusion scored better than TokenFusion on all four reported input-to-target tasks. The evaluation uses FID/KID for RGB predictions and MAE/MSE for other predictions; lower values are better. The RGB metrics are reported at scale ×10⁻² and the other metrics at scale ×10⁻¹. Experiments followed the same sampling setup of 1,000 training and 500 testing images, with training epochs aligned.

    Inputs → targetMethodFirst metricSecond metric
    Shade+Texture → RGBTokenFusionFID 47.31KID 0.94
    Shade+Texture → RGBGeminiFusionFID 41.32KID 0.81
    Depth+Normal → RGBTokenFusionFID 103.87KID 4.24
    Depth+Normal → RGBGeminiFusionFID 96.98KID 3.71
    RGB+Shade → NormalTokenFusionMAE 0.67MSE 1.75
    RGB+Shade → NormalGeminiFusionMAE 0.65MSE 1.69
    RGB+Edge → DepthTokenFusionMAE 0.22MSE 0.55
    RGB+Edge → DepthGeminiFusionMAE 0.20MSE 0.49

    The consistent reductions across both RGB-generation and non-RGB prediction tasks show that the advantage is not limited to semantic segmentation.

  8. Knowl 8 — GeminiFusion improves KITTI vehicle detection in MVX-Net

    data/table

    GeminiFusion was inserted into the fusion layer of MVX-Net for 3D vehicle detection using image and depth inputs on the KITTI validation set. Training epochs were aligned, and the IoU threshold was 0.7. Average precision is reported for 11-point (APR11) and 40-point (APR40) evaluation at easy, medium, and hard difficulty levels.

    MethodParameters (M)APR11 easyAPR11 mediumAPR11 hardAPR40 easyAPR40 mediumAPR40 hard
    MVX-Net33.887.4977.0474.5488.4178.7774.27
    MVX-Net + GeminiFusion34.888.4977.3674.6189.4378.7674.46

    GeminiFusion improved most reported detection scores with a 1.0-million-parameter increase; APR40 medium was nearly unchanged and slightly lower, at 78.76 versus 78.77.

  9. Knowl 9 — Selective layer placement trades fusion coverage for latency

    empirical result

    On NYUDv2 and SUN RGB-D segmentation with MiT-B3, the authors measured the effect of applying GeminiFusion in only the final kk layers. Latency is the average over NYUDv2 validation samples; parameter counts and GFLOPs are also reported. Applying GeminiFusion in the last 10 layers yielded better mIoU than TokenFusion on both datasets while also reducing latency from 126 ms to 116 ms. Even placement in only the final layer achieved higher mIoU than TokenFusion at 102 ms. Applying GeminiFusion in all 28 layers gave the best reported accuracy but increased latency.

    MethodkkParameters (M)GFLOPsLatency (ms)NYUDv2 mIoUSUN RGB-D mIoU
    TokenFusion2845.910812654.251.4
    GeminiFusion2875.817415356.852.7
    GeminiFusion1062.513811656.452.2
    GeminiFusion148.811910255.151.9
    GeminiFusion, no fusion module045.91089553.351.2
  10. Knowl 10 — Current scope requires aligned homogeneous visual modalities

    limitation

    The implemented GeminiFusion framework targets homogeneous visual modalities that can be represented as image-like features of the same subject and aligned at corresponding spatial locations. The paper states that the approach does not currently handle heterogeneous combinations such as images paired with audio or text; alignment for the input pair also has to be defined. Extending fusion to these heterogeneous data types is left for future work.

Coverage note — Qualitative translation visualizations and detailed dataset inventories are omitted because they add no distinct standalone result beyond the quantified task comparisons and experimental conditions captured here.

References

  1. 1.Antol, S., Agrawal, A., Lu, J., Mitchell, M., Batra, D., Zitnick, C. L., and Parikh, D. Vqa: Visual question answering. In Proceedings of the IEEE international conference on computer vision, 2015.
  2. 2.Atrey, P. K., Hossain, M. A., El Saddik, A., and Kankanhalli, M. S. Multimodal fusion for multimedia analysis: a survey. Multimedia systems, 2010.
  3. 3.Baltrusaitis, T., Ahuja, C., and Morency, L.-P. Multimodal machine learning: A survey and taxonomy. IEEE transactions on pattern analysis and machine intelligence, 2018.
  4. 4.Ben-Younes, H., Cadene, R., Cord, M., and Thome, N. Mutan: Multimodal tucker fusion for visual question answering. In Proceedings of the IEEE international conference on computer vision, 2017.
  5. 5.Bruni, E., Tran, N.-K., and Baroni, M. Multimodal distributional semantics. Journal of artificial intelligence research, 2014.
  6. 6.Caesar, H., Bankiti, V., Lang, A. H., Vora, S., Liong, V. E., Xu, Q., Krishnan, A., Pan, Y., Baldan, G., and Beijbom, O. nuscenes: A multimodal dataset for autonomous driving. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020.
  7. 7.Cao, J., Leng, H., Lischinski, D., Cohen-Or, D., Tu, C., and Li, Y. Shapeconv: Shape-aware convolutional layer for indoor rgb-d semantic segmentation. In Proceedings of the IEEE/CVF international conference on computer vision, 2021.
  8. 8.Carion, N., Massa, F., Synnaeve, G., Usunier, N., Kirillov, A., and Zagoruyko, S. End-to-end object detection with transformers. In European conference on computer vision, 2020.
  9. 9.Chen, C., Rosa, S., Miao, Y., Lu, C. X., Wu, W., Markham, A., and Trigoni, N. Selective sensor fusion for neural visual-inertial odometry. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019.
  10. 10.De Vries, H., Strub, F., Mary, J., Larochelle, H., Pietquin, O., and Courville, A. C. Modulating early visual processing by language. Advances in Neural Information Processing Systems, 2017.
  11. 11.Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L. Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition, pp. 248–255. Ieee, 2009.
  12. 12.Dong, S., Feng, Y., Yang, Q., Huang, Y., Liu, D., and Fan, H. Efficient multimodal semantic segmentation via dual-prompt learning. arXiv preprint arXiv:2312.00360, 2023.
  13. 13.Fu, K., Fan, D.-P., Ji, G.-P., and Zhao, Q. Jl-dcf: Joint learning and densely-cooperative fusion framework for rgb-d salient object detection. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020.
  14. 14.Geiger, A., Lenz, P., and Urtasun, R. Are we ready for autonomous driving? the kitti vision benchmark suite. In Conference on Computer Vision and Pattern Recognition (CVPR), 2012.
  15. 15.Girdhar, R., Singh, M., Ravi, N., van der Maaten, L., Joulin, A., and Misra, I. Omnivore: A single model for many visual modalities. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 16102–16112, 2022.
  16. 16.Glodek, M., Tschechne, S., Layher, G., Schels, M., Brosch, T., Scherer, S., Kachele, M., Schmidt, M., Neumann, H., Palm, G., et al. Multiple classifier systems for the classification of audio-visual emotional states. In Affective Computing and Intelligent Interaction: Fourth International Conference, ACII 2011, Memphis, TN, USA, October 9–12, 2011, Proceedings, Part II, 2011.
  17. 17.Guo, J., Han, K., Wu, H., Tang, Y., Chen, X., Wang, Y., and Xu, C. Cmt: Convolutional neural networks meet vision transformers. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022a.
  18. 18.Guo, J., Tang, Y., Han, K., Chen, X., Wu, H., Xu, C., Xu, C., and Wang, Y. Hire-mlp: Vision mlp via hierarchical rearrangement. In Proceedings of the ieee/cvf conference on computer vision and pattern recognition, 2022b.
  19. 19.Gupta, S., Girshick, R., Arbelaez, P., and Malik, J. Learning rich features from rgb-d images for object detection and segmentation. In Computer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part VII 13, 2014.
  20. 20.Ha, Q., Watanabe, K., Karasawa, T., Ushiku, Y., and Harada, T. Mfnet: Towards real-time semantic segmentation for autonomous vehicles with multi-spectral scenes. In 2017 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 2017.
  21. 21.Hazirbas, C., Ma, L., Domokos, C., and Cremers, D. Fusenet: Incorporating depth into semantic segmentation via fusion-based cnn architecture. In Computer Vision–ACCV 2016: 13th Asian Conference on Computer Vision, Taipei, Taiwan, November 20-24, 2016, Revised Selected Papers, Part I 13, 2017.
  22. 22.Hori, C., Hori, T., Lee, T.-Y., Zhang, Z., Harsham, B., Hershey, J. R., Marks, T. K., and Sumi, K. Attention-based multimodal fusion for video description. In Proceedings of the IEEE international conference on computer vision, 2017.
  23. 23.Hu, J., Shen, L., and Sun, G. Squeeze-and-excitation networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, 2018.
  24. 24.Hu, X., Yang, K., Fei, L., and Wang, K. Acnet: Attention based network to exploit complementary features for rgbd semantic segmentation. In 2019 IEEE International conference on image processing (ICIP), pp. 1440–1444. IEEE, 2019.
  25. 25.Jiang, Z., Taira, H., Miyashita, N., and Okutomi, M. Self-supervised ego-motion estimation based on multi-layer fusion of rgb and inferred depth. In 2022 International Conference on Robotics and Automation (ICRA), 2022.
  26. 26.Kalra, A., Taamazyan, V., Rao, S. K., Venkataraman, K., Raskar, R., and Kadambi, A. Deep polarization cues for transparent object segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020.
  27. 27.Kazakos, E., Nagrani, A., Zisserman, A., and Damen, D. Epic-fusion: Audio-visual temporal binding for egocentric action recognition. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 2019.
  28. 28.Li, Z., Wang, W., Li, H., Xie, E., Sima, C., Lu, T., Qiao, Y., and Dai, J. Bevformer: Learning bird’s-eye-view representation from multi-camera images via spatiotemporal transformers. In European conference on computer vision, 2022.
  29. 29.Lin, D., Chen, G., Cohen-Or, D., Heng, P.-A., and Huang, H. Cascaded feature network for semantic segmentation of rgb-d images. In Proceedings of the IEEE international conference on computer vision, 2017.
  30. 30.Liu, N., Zhang, N., Wan, K., Shao, L., and Han, J. Visual saliency transformer. In Proceedings of the IEEE/CVF international conference on computer vision, 2021a.
  31. 31.Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., and Guo, B. Swin transformer: Hierarchical vision transformer using shifted windows. In Proceedings of the IEEE/CVF international conference on computer vision, 2021b.
  32. 32.Liu, Z., Wang, Y., Tu, Z., Xiao, Y., and Tang, B. Tritransnet: Rgb-d salient object detection with a triplet transformer embedding network. In Proceedings of the 29th ACM international conference on multimedia, 2021c.
  33. 33.Lu, J., Batra, D., Parikh, D., and Lee, S. Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks. Advances in neural information processing systems, 2019.
  34. 34.Nagrani, A., Yang, S., Arnab, A., Jansen, A., Schmid, C., and Sun, C. Attention bottlenecks for multimodal fusion. Advances in Neural Information Processing Systems, 2021.
  35. 35.Owens, A. and Efros, A. A. Audio-visual scene analysis with self-supervised multisensory features. In Proceedings of the European conference on computer vision (ECCV), 2018.
  36. 36.Pandeya, Y. R. and Lee, J. Deep learning-based late fusion of multimodal information for emotion classification of music video. Multimedia Tools and Applications, 2021.
  37. 37.Park, S.-J., Hong, K.-S., and Lee, S. RDFnet: Rgb-d multi-level residual feature fusion for indoor semantic segmentation. In Proceedings of the IEEE international conference on computer vision, 2017.
  38. 38.Prakash, A., Chitta, K., and Geiger, A. Multi-modal fusion transformer for end-to-end autonomous driving. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021.
  39. 39.Ramachandram, D. and Taylor, G. W. Deep multimodal learning: A survey on recent advances and trends. IEEE signal processing magazine, 2017.
  40. 40.Seichter, D., Stephan, B., Fischedick, S. B., Muller, S., Rabes, L., and Gross, H.-M. Panopticndt: Efficient and robust panoptic mapping. In 2023 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pp. 7233–7240. IEEE, 2023.
  41. 41.Shvetsova, N., Chen, B., Rouditchenko, A., Thomas, S., Kingsbury, B., Feris, R. S., Harwath, D., Glass, J., and Kuehne, H. Everything at once-multi-modal fusion transformer for video retrieval. In Proceedings of the ieee/cvf conference on computer vision and pattern recognition, 2022.
  42. 42.Silberman, N., Hoiem, D., Kohli, P., and Fergus, R. Indoor segmentation and support inference from rgbd images. In Computer Vision–ECCV 2012: 12th European Conference on Computer Vision, Florence, Italy, October 7-13, 2012, Proceedings, Part V 12, 2012.
  43. 43.Sindagi, V. A., Zhou, Y., and Tuzel, O. Mvx-net: Multimodal voxelnet for 3d object detection. In 2019 International Conference on Robotics and Automation (ICRA), pp. 7276–7282. IEEE, 2019.
  44. 44.Smith, L. and Gasser, M. The development of embodied cognition: Six lessons from babies. Artificial life, 2005.
  45. 45.Snoek, C. G., Worring, M., and Smeulders, A. W. Early versus late fusion in semantic video analysis. In Proceedings of the 13th annual ACM international conference on Multimedia, 2005.
  46. 46.Song, S., Lichtenberg, S. P., and Xiao, J. Sun rgb-d: A rgb-d scene understanding benchmark suite. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 567–576, 2015.
  47. 47.Srivastava, S. and Sharma, G. Omnivec: Learning robust representations with cross modal sharing. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pp. 1236–1248, 2024.
  48. 48.Su, W., Zhu, X., Cao, Y., Li, B., Lu, L., Wei, F., and Dai, J. Vl-bert: Pre-training of generic visual-linguistic representations. arXiv preprint arXiv:1908.08530, 2019.
  49. 49.Sun, C., Myers, A., Vondrick, C., Murphy, K., and Schmid, C. Videobert: A joint model for video and language representation learning. In Proceedings of the IEEE/CVF international conference on computer vision, 2019a.
  50. 50.Sun, Y., Zuo, W., and Liu, M. Rtfnet: Rgb-thermal fusion network for semantic segmentation of urban scenes. IEEE Robotics and Automation Letters, 2019b.
  51. 51.Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. Attention is all you need. Advances in neural information processing systems, 2017.
  52. 52.Wang, F., Pan, J., Xu, S., and Tang, J. Learning discriminative cross-modality features for rgb-d saliency detection. IEEE Transactions on Image Processing, 2022a.
  53. 53.Wang, Q., Wu, B., Zhu, P., Li, P., Zuo, W., and Hu, Q. Eca-net: Efficient channel attention for deep convolutional neural networks. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020a.
  54. 54.Wang, W., Tran, D., and Feiszli, M. What makes training multimodal classification networks hard? In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020b.
  55. 55.Wang, Y., Huang, W., Sun, F., Xu, T., Rong, Y., and Huang, J. Deep multimodal fusion by channel exchanging. Advances in neural information processing systems, 2020c.
  56. 56.Wang, Y., Chen, X., Cao, L., Huang, W., Sun, F., and Wang, Y. Multimodal token fusion for vision transformers. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022b.
  57. 57.Wei, X., Zhang, T., Li, Y., Zhang, Y., and Wu, F. Multi-modality cross attention network for image and sentence matching. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020.
  58. 58.Woo, S., Park, J., Lee, J.-Y., and Kweon, I. S. Cbam: Convolutional block attention module. In Proceedings of the European conference on computer vision (ECCV), 2018.
  59. 59.Xie, E., Wang, W., Yu, Z., Anandkumar, A., Alvarez, J. M., and Luo, P. Segformer: Simple and efficient design for semantic segmentation with transformers. Advances in Neural Information Processing Systems, 2021.
  60. 60.Xu, Y., Li, C., Li, D., Sheng, X., Jiang, F., Tian, L., and Sirasao, A. Fdvit: Improve the hierarchical architecture of vision transformer. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023.
  61. 61.Yang, X., Yuan, L., Wilber, K., Sharma, A., Gu, X., Qiao, S., Debats, S., Wang, H., Adam, H., Sirotenko, M., et al. Polymax: General dense prediction with mask transformer. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pp. 1050–1061, 2024.
  62. 62.Ye, L., Rochan, M., Liu, Z., and Wang, Y. Cross-modal self-attention network for referring image segmentation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2019.
  63. 63.Yin, B., Zhang, X., Li, Z., Liu, L., Cheng, M.-M., and Hou, Q. Dformer: Rethinking rgbd representation learning for semantic segmentation. arXiv preprint arXiv:2309.09668, 2023.
  64. 64.Zamir, A. R., Sax, A., Shen, W., Guibas, L. J., Malik, J., and Savarese, S. Taskonomy: Disentangling task transfer learning. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 3712–3722, 2018.
  65. 65.Zhang, J., Fan, D.-P., Dai, Y., Anwar, S., Saleh, F., Aliakbarian, S., and Barnes, N. Uncertainty inspired rgb-d saliency detection. IEEE transactions on pattern analysis and machine intelligence, 2021a.
  66. 66.Zhang, J., Yang, K., and Stiefelhagen, R. Exploring event-driven dynamic context for accident scene segmentation. IEEE Transactions on Intelligent Transportation Systems, 2021b.
  67. 67.Zhang, J., Liu, H., Yang, K., Hu, X., Liu, R., and Stiefelhagen, R. Cmx: Cross-modal fusion for rgb-x semantic segmentation with transformers. IEEE Transactions on Intelligent Transportation Systems, 2023a.
  68. 68.Zhang, J., Liu, R., Shi, H., Yang, K., Reiß, S., Peng, K., Fu, H., Wang, K., and Stiefelhagen, R. Delivering arbitrary-modal semantic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023b.
  69. 69.Zhang, Y. and Funkhouser, T. Deep depth completion of a single rgb-d image. In Proceedings of the IEEE conference on computer vision and pattern recognition, 2018.
  70. 70.Zhang, Y., Zhang, Q., Zhu, Z., Hou, J., and Yuan, Y. Glenet: Boosting 3d object detectors with generative label uncertainty estimation. International Journal of Computer Vision, 131(12): 3332–3352, 2023c.
  71. 71.Zhao, X., Zhang, L., Pang, Y., Lu, H., and Zhang, L. A single stream network for robust and real-time rgb-d salient object detection. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XXII 16, 2020.
  72. 72.Zhao, Y., Zhao, J., Li, J., and Chen, X. Rgb-d salient object detection with ubiquitous target awareness. IEEE Transactions on Image Processing, 2021.
  73. 73.Zheng, W., Tang, W., Jiang, L., and Fu, C.-W. Se-ssd: Self-ensembling single-stage object detector from point cloud. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 14494–14503, 2021.
  74. 74.Zhu, R., Han, C., Qian, Y., Sun, Q., Li, X., Gao, M., Cao, X., and Xian, Y. Exchanging-based multimodal fusion with transformer. arXiv preprint arXiv:2309.02190, 2023.
  75. 75.Zhuang, Z., Li, R., Jia, K., Wang, Q., Li, Y., and Tan, M. Perception-aware multi-sensor fusion for 3d lidar semantic segmentation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021a.
  76. 76.Zhuang, Z., Li, R., Jia, K., Wang, Q., Li, Y., and Tan, M. Perception-aware multi-sensor fusion for 3D LiDAR semantic segmentation. In In Proceedings of the IEEE international conference on computer vision, 2021b.

Citation

MLA
Jia, D., et al. “GeminiFusion: Efficient Pixel-wise Multimodal Fusion for Vision Transformer”. arXiv, 2024, https://doi.org/10.48550/arxiv.2406.01210.
APA
Jia, D., Guo, J., Han, K., Wu, H., Zhang, C., Xu, C., & Chen, X. (2024). GeminiFusion: Efficient Pixel-wise Multimodal Fusion for Vision Transformer. arXiv. https://doi.org/10.48550/arxiv.2406.01210
Chicago
Jia, D., J. Guo, K. Han, et al. 2024. “GeminiFusion: Efficient Pixel-wise Multimodal Fusion for Vision Transformer”. Preprint, ArXiv. https://doi.org/10.48550/arxiv.2406.01210.
Harvard
Jia, D. et al. (2024) “GeminiFusion: Efficient Pixel-wise Multimodal Fusion for Vision Transformer”. arXiv. Available at: https://doi.org/10.48550/arxiv.2406.01210.
Vancouver
1. Jia D, Guo J, Han K, Wu H, Zhang C, Xu C, Chen X (2024) GeminiFusion: Efficient Pixel-wise Multimodal Fusion for Vision Transformer. https://doi.org/10.48550/arxiv.2406.01210

BibTeX

@misc{https://doi.org/10.48550/arxiv.2406.01210,
  doi = {10.48550/ARXIV.2406.01210},
  url = {https://arxiv.org/abs/2406.01210},
  author = {Jia, Ding and Guo, Jianyuan and Han, Kai and Wu, Han and Zhang, Chao and Xu, Chang and Chen, Xinghao},
  keywords = {Computer Vision and Pattern Recognition (cs.CV), FOS: Computer and information sciences, FOS: Computer and information sciences},
  title = {GeminiFusion: Efficient Pixel-wise Multimodal Fusion for Vision Transformer},
  publisher = {arXiv},
  year = {2024},
  copyright = {Creative Commons Attribution 4.0 International}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/