V2V4Real: A Real-World Large-Scale Dataset for Vehicle-to-Vehicle Cooperative Perception
Runsheng XuXin XiaJinlong LiHanzhao LiShuo ZhangZhengzhong TuZonglin MengHao XiangXiaoyu DongRui Song
Introduces the first large-scale real-world dataset and benchmark for vehicle-to-vehicle cooperative perception, providing synchronized multimodal sensor data across 410 kilometers to evaluate cooperative 3D detection, tracking, and sim-to-real domain adaptation.
Autonomous driving systems rely heavily on accurate environmental perception to operate safely. However, single-vehicle perception systems face fundamental physical limitations, such as occlusions caused by surrounding traffic and limited sensor ranges. Vehicle-to-vehicle cooperative perception—where multiple connected vehicles share sensor data over wireless links—addresses these safety-critical blind spots. While theoretical concepts are well established, evaluating cooperative perception algorithms under realistic real-world road conditions and addressing performance degradation caused by data shifts remains an urgent industry challenge.
The article demonstrates and evaluates multiple cooperative perception baseline algorithms using the real-world V2V4Real dataset. It specifically benchmarks 3D object detection across various vehicle communication architectures, establishes a framework for cooperative multi-object tracking, and evaluates domain adaptation techniques designed to transfer synthetic training knowledge to real-world operational environments.
To conduct this evaluation, the analysis tested several intermediate feature fusion methods alongside early fusion, late fusion, and standalone single-vehicle baselines across urban and highway driving environments equipped with light detection and ranging sensors and cameras. The evaluation assessed both synchronized and asynchronous communication conditions, analyzed the impact of point-cloud data augmentations, and incorporated feature-level and object-level domain discriminators to adapt models trained on synthetic data to real-world scenarios.
The findings show that cooperative perception substantially improves 3D object detection accuracy compared to single-vehicle sensing, particularly in congested urban scenes where occlusions are prevalent. Among all tested architectures, advanced intermediate fusion models—specifically transformer-based methods like CoBEVT and V2X-ViT—achieved the highest overall accuracy, reaching up to 51.1% average precision under synchronized settings. Furthermore, data augmentations proved critical, as omitting rotations, scaling, and flips reduced intermediate fusion detection performance by 13.7% to 16.9%. Finally, applying domain adaptation substantially improved target detection when transferring from synthetic source data to real-world environments, though simpler feature-fusion methods exhibited higher false-positive rates in complex intersections compared to transformer-based approaches.
These results demonstrate that cooperative vehicle perception provides meaningful safety and situational awareness gains that individual vehicle sensors cannot achieve alone. For engineering and product leaders, intermediate fusion presents an optimal balance, delivering near-raw-sensor detection accuracy while maintaining low communication bandwidth requirements. Additionally, effective domain adaptation reduces the financial and operational costs associated with collecting exhaustive real-world training datasets by unlocking the value of synthetic simulation data.
Development teams should prioritize intermediate fusion architectures, particularly transformer-based designs, for connected vehicle platforms while standardizing rigorous point-cloud data augmentation pipelines. Future work should focus on testing these pipelines under broader weather, road, and high-density traffic conditions to further validate operational reliability.
- Paper: Where2comm: Communication-Efficient Collaborative Perception via Spatial Confidence Maps, Yue Hu et al. (2022). Introduces foundational communication-efficient spatial confidence mechanisms for collaborative perception across connected agents that directly motivate the intermediate fusion benchmarks in V2V4Real.
- Paper: BEVFusion: Multi-Task Multi-Sensor Fusion with Unified Bird's-Eye View Representation, Zhijian Liu et al. (2022). Establishes the unified bird's-eye view multi-sensor fusion architecture that serves as a core baseline and architectural principle evaluated across V2V4Real's vehicle-to-vehicle pipelines.
- Paper: BEVFormer: Learning Bird's-Eye-View Representation from Multi-Camera Images via Spatiotemporal Transformers, Zhiqi Li et al. (2022). Details the transformer-based bird's-eye view feature extraction principles that underpin advanced intermediate fusion models tested in the benchmark.
- Paper: Domain Adaptive Faster R-CNN for Object Detection in the Wild, Yuhua Chen et al. (2018). Presents essential adversarial domain adaptation techniques for object detection that prepare readers for V2V4Real's synthetic-to-real domain adaptation experiments.
- Paper: Center-based 3D Object Detection and Tracking, Tianwei Yin et al. (2020). Provides the anchor-free 3D detection and velocity-based tracking baseline used as the standard backbone across modern multi-object perception pipelines.
- Paper: DI-V2X: Learning Domain-Invariant Representation for Vehicle-Infrastructure Collaborative 3D Object Detection, Xiang Li et al. (2024). Extends collaborative 3D perception by tackling domain gaps caused by disparate sensor modalities between roadside infrastructure and moving vehicles.
- Paper: DeepAccident: A Motion and Accident Prediction Benchmark for V2X Autonomous Driving, Tianqi Wang et al. (2024). Builds upon V2X collaborative bird's-eye-view perception datasets to benchmark end-to-end motion forecasting and accident prediction in safety-critical collision scenarios.
- Paper: Collaboration Helps Camera Overtake LiDAR in 3D Detection, Yue Hu et al. (2023). Applies collaborative multi-agent perception principles to camera-only systems, demonstrating how sharing visual cues across agents can mitigate the absence of expensive LiDAR.
