keyword
intermediate fusion
Intermediate fusion is a strategy in multimodal machine learning and cooperative perception where information from multiple sensors, data modalities, or collaborating agents is combined at the level of learned feature representations rather than at the raw data or final decision stage. In this approach, each input stream is first processed independently through dedicated encoders or neural network layers to generate intermediate latent features, which are subsequently integrated using operations such as concatenation, attention mechanisms, or cross-modal layers before producing a final prediction. Positioned between early fusion, which aggregates raw inputs directly, and late fusion, which combines independent final outputs, intermediate fusion provides an effective balance between capturing rich cross-modal interactions and maintaining communication and computational efficiency.
2 items

V2V4Real: A Real-World Large-Scale Dataset for Vehicle-to-Vehicle Cooperative Perception
Runsheng Xu, Xin Xia, Jinlong Li, Hanzhao Li, Shuo Zhang, Zhengzhong Tu, Zonglin Meng, Hao Xiang, Xiaoyu Dong, Rui Song, Hongkai Yu, Bolei Zhou, Jiaqi Ma
Why you should read this
Introduces the first large-scale real-world dataset and benchmark for vehicle-to-vehicle cooperative perception, providing synchronized multimodal sensor data across 410 kilometers to evaluate cooperative 3D detection, tracking, and sim-to-real domain adaptation.
Modern perception systems of autonomous vehicles are known to be sensitive to occlusions and lack the capability of long perceiving range. It has been one of the key bottlenecks that prevents Level 5 autonomy. Recent research has demonstrated that the Vehicle-to-Vehicle (V2V) cooperative perception system has great potential to revolutionize the autonomous driving industry. However, the lack of a real-world dataset hinders the progress of this field. To facilitate the development of cooperative perception, we present V2V4Real, the first large-scale real-world multi-modal dataset for V2V perception. The data is collected by two vehicles equipped with multi-modal sensors driving together through diverse scenarios. Our V2V4Real dataset covers a driving area of 410 km, comprising 20K LiDAR frames, 40K RGB frames, 240K annotated 3D bounding boxes for 5 classes, and HDMaps that cover all the driving routes. V2V4Real introduces three perception tasks, including cooperative 3D object detection, cooperative 3D object tracking, and Sim2Real domain adaptation for cooperative perception. We provide comprehensive benchmarks of recent cooperative perception algorithms on three tasks. The V2V4Real dataset can be found at this https URL.
Added
2026-09-26

Provable Dynamic Fusion for Low-Quality Multimodal Data
Qingyang Zhang, Haitao Wu, Changqing Zhang, Qinghua Hu, Huazhu Fu, Joey Tianyi Zhou, Xi Peng
Why you should read this
Establishes theoretical generalization bounds for dynamic decision-level multimodal integration and introduces Quality-aware Multimodal Fusion to prevent performance degradation on noisy and low-quality inputs.
The inherent challenge of multimodal fusion is to precisely capture the cross-modal correlation and flexibly conduct cross-modal interaction. To fully release the value of each modality and mitigate the influence of low-quality multimodal data, dynamic multimodal fusion emerges as a promising learning paradigm. Despite its widespread use, theoretical justifications in this field are still notably lacking. Can we design a provably robust multimodal fusion method? This paper provides theoretical understandings to answer this question under a most popular multimodal fusion framework from the generalization perspective. We proceed to reveal that several uncertainty estimation solutions are naturally available to achieve robust multimodal fusion. Then a novel multimodal fusion framework termed Quality-aware Multimodal Fusion (QMF) is proposed, which can improve the performance in terms of classification accuracy and model robustness. Extensive experimental results on multiple benchmarks can support our findings.
Added
2026-09-26
