Built independently by an author, for readers. Read the story and support ChapterPal

keyword

intermediate fusion

Intermediate fusion is a strategy in multimodal machine learning and cooperative perception where information from multiple sensors, data modalities, or collaborating agents is combined at the level of learned feature representations rather than at the raw data or final decision stage. In this approach, each input stream is first processed independently through dedicated encoders or neural network layers to generate intermediate latent features, which are subsequently integrated using operations such as concatenation, attention mechanisms, or cross-modal layers before producing a final prediction. Positioned between early fusion, which aggregates raw inputs directly, and late fusion, which combines independent final outputs, intermediate fusion provides an effective balance between capturing rich cross-modal interactions and maintaining communication and computational efficiency.

2 items

V2V4Real: A Real-World Large-Scale Dataset for Vehicle-to-Vehicle Cooperative Perception

V2V4Real: A Real-World Large-Scale Dataset for Vehicle-to-Vehicle Cooperative Perception

Runsheng Xu, Xin Xia, Jinlong Li, Hanzhao Li, Shuo Zhang, Zhengzhong Tu, Zonglin Meng, Hao Xiang, Xiaoyu Dong, Rui Song, Hongkai Yu, Bolei Zhou, Jiaqi Ma

OrganizationsCleveland State UniversityFraunhofer InstituteNorthwestern UniversityTechnical University of MunichUniversity of California, Los AngelesUniversity of Texas at Austin

Why you should read this

Introduces the first large-scale real-world dataset and benchmark for vehicle-to-vehicle cooperative perception, providing synchronized multimodal sensor data across 410 kilometers to evaluate cooperative 3D detection, tracking, and sim-to-real domain adaptation.

Modern perception systems of autonomous vehicles are known to be sensitive to occlusions and lack the capability of long perceiving range. It has been one of the key bottlenecks that prevents Level 5 autonomy. Recent research has demonstrated that the Vehicle-to-Vehicle (V2V) cooperative perception system has great potential to revolutionize the autonomous driving industry. However, the lack of a real-world dataset hinders the progress of this field. To facilitate the development of cooperative perception, we present V2V4Real, the first large-scale real-world multi-modal dataset for V2V perception. The data is collected by two vehicles equipped with multi-modal sensors driving together through diverse scenarios. Our V2V4Real dataset covers a driving area of 410 km, comprising 20K LiDAR frames, 40K RGB frames, 240K annotated 3D bounding boxes for 5 classes, and HDMaps that cover all the driving routes. V2V4Real introduces three perception tasks, including cooperative 3D object detection, cooperative 3D object tracking, and Sim2Real domain adaptation for cooperative perception. We provide comprehensive benchmarks of recent cooperative perception algorithms on three tasks. The V2V4Real dataset can be found at this https URL.

Added

2026-09-26