V2V4Real: A Real-World Large-Scale Dataset for Vehicle-to-Vehicle Cooperative Perception

Runsheng XuXin XiaJinlong LiHanzhao LiShuo ZhangZhengzhong TuZonglin MengHao XiangXiaoyu DongRui Song

article2023CVPR263 citations

Introduces the first large-scale real-world dataset and benchmark for vehicle-to-vehicle cooperative perception, providing synchronized multimodal sensor data across 410 kilometers to evaluate cooperative 3D detection, tracking, and sim-to-real domain adaptation.

Listen

Autonomous driving systems rely heavily on accurate environmental perception to operate safely. However, single-vehicle perception systems face fundamental physical limitations, such as occlusions caused by surrounding traffic and limited sensor ranges. Vehicle-to-vehicle cooperative perception—where multiple connected vehicles share sensor data over wireless links—addresses these safety-critical blind spots. While theoretical concepts are well established, evaluating cooperative perception algorithms under realistic real-world road conditions and addressing performance degradation caused by data shifts remains an urgent industry challenge.

The article demonstrates and evaluates multiple cooperative perception baseline algorithms using the real-world V2V4Real dataset. It specifically benchmarks 3D object detection across various vehicle communication architectures, establishes a framework for cooperative multi-object tracking, and evaluates domain adaptation techniques designed to transfer synthetic training knowledge to real-world operational environments.

To conduct this evaluation, the analysis tested several intermediate feature fusion methods alongside early fusion, late fusion, and standalone single-vehicle baselines across urban and highway driving environments equipped with light detection and ranging sensors and cameras. The evaluation assessed both synchronized and asynchronous communication conditions, analyzed the impact of point-cloud data augmentations, and incorporated feature-level and object-level domain discriminators to adapt models trained on synthetic data to real-world scenarios.

The findings show that cooperative perception substantially improves 3D object detection accuracy compared to single-vehicle sensing, particularly in congested urban scenes where occlusions are prevalent. Among all tested architectures, advanced intermediate fusion models—specifically transformer-based methods like CoBEVT and V2X-ViT—achieved the highest overall accuracy, reaching up to 51.1% average precision under synchronized settings. Furthermore, data augmentations proved critical, as omitting rotations, scaling, and flips reduced intermediate fusion detection performance by 13.7% to 16.9%. Finally, applying domain adaptation substantially improved target detection when transferring from synthetic source data to real-world environments, though simpler feature-fusion methods exhibited higher false-positive rates in complex intersections compared to transformer-based approaches.

These results demonstrate that cooperative vehicle perception provides meaningful safety and situational awareness gains that individual vehicle sensors cannot achieve alone. For engineering and product leaders, intermediate fusion presents an optimal balance, delivering near-raw-sensor detection accuracy while maintaining low communication bandwidth requirements. Additionally, effective domain adaptation reduces the financial and operational costs associated with collecting exhaustive real-world training datasets by unlocking the value of synthetic simulation data.

Development teams should prioritize intermediate fusion architectures, particularly transformer-based designs, for connected vehicle platforms while standardizing rigorous point-cloud data augmentation pipelines. Future work should focus on testing these pipelines under broader weather, road, and high-density traffic conditions to further validate operational reliability.

arXiv: 2303.07601
Cover for V2V4Real: A Real-World Large-Scale Dataset for Vehicle-to-Vehicle Cooperative Perception

Abstract

Modern perception systems of autonomous vehicles are known to be sensitive to occlusions and lack the capability of long perceiving range. It has been one of the key bottlenecks that prevents Level 5 autonomy. Recent research has demonstrated that the Vehicle-to-Vehicle (V2V) cooperative perception system has great potential to revolutionize the autonomous driving industry. However, the lack of a real-world dataset hinders the progress of this field. To facilitate the development of cooperative perception, we present V2V4Real, the first large-scale real-world multi-modal dataset for V2V perception. The data is collected by two vehicles equipped with multi-modal sensors driving together through diverse scenarios. Our V2V4Real dataset covers a driving area of 410 km, comprising 20K LiDAR frames, 40K RGB frames, 240K annotated 3D bounding boxes for 5 classes, and HDMaps that cover all the driving routes. V2V4Real introduces three perception tasks, including cooperative 3D object detection, cooperative 3D object tracking, and Sim2Real domain adaptation for cooperative perception. We provide comprehensive benchmarks of recent cooperative perception algorithms on three tasks. The V2V4Real dataset can be found at this https URL.

Table of Contents

  • 1 Introduction
  • 2 Relaed Work
  • 2.1 Autonomous Driving Datasets.
  • 2.2 3D Detection
  • 2.3 V2V/V2X Cooperative Perception
  • 3 V2V4Real Dataset
  • 3.1 Data Acquisition
  • 3.2 Data Annotation
  • 3.3 Data Analysis
  • 4 Tasks
  • 4.1 Cooperative 3D Object Detection
  • 4.2 Object Tracking
  • 4.3 Sim2Real Domain Adaptation
  • 5 Experiments
  • 5.1 Implementation Details
  • 5.2 3D LiDAR Object Detection
  • 5.3 3D Object Tracking
  • 5.4 Sim2Real Domain Adaptation
  • 6 Conclusion
  • 7 Acknowledgement
  • References
  • A Appendix
  • B Dataset Visualization
  • C Implementation Details
  • C.1 Cooperative 3D Object Detection
  • C.2 Cooperative Tracking
  • C.3 Domain Adaption
  • D Ablation Studies
  • E Detection Results
  • F Domain Adaptation Results

Knowls

  1. Knowl 1 — Cooperative 3D Multi-Object Tracking Framework

    model/method

    The cooperative 3D tracking framework adopts a tracking-by-detection paradigm where detection bounding boxes are derived from cooperatively fused multi-agent visual perception rather than single-vehicle observations.

    At time frame tt, the cooperative detection network outputs NtN_t bounding boxes D(t)={Dit}i=1NtD(t) = \{D_i^t\}_{i=1}^{N_t}, where each detection is parameterized as Dit=(xit,yit,zit,θit,wit,hit,lit,sit)D_i^t = (x_i^t, y_i^t, z_i^t, \theta_i^t, w_i^t, h_i^t, l_i^t, s_i^t), with center coordinates (xit,yit,zit)(x_i^t, y_i^t, z_i^t), heading yaw angle θit\theta_i^t, dimensions (wit,hit,lit)(w_i^t, h_i^t, l_i^t), and detection confidence score sits_i^t.

    The tracker maintains Mt−1M_{t-1} previously confirmed trajectories T(t−1)={Tjt−1}j=1Mt−1T(t-1) = \{T_j^{t-1}\}_{j=1}^{M_{t-1}}, where each trajectory state is represented as Tjt−1=(xjt−1,yjt−1,zjt−1,θjt−1,wjt−1,hjt−1,ljt−1,sjt−1,vxjt−1,vyjt−1,vzjt−1)T_j^{t-1} = (x_j^{t-1}, y_j^{t-1}, z_j^{t-1}, \theta_j^{t-1}, w_j^{t-1}, h_j^{t-1}, l_j^{t-1}, s_j^{t-1}, vx_j^{t-1}, vy_j^{t-1}, vz_j^{t-1}), incorporating 3D linear velocities (vxjt−1,vyjt−1,vzjt−1)(vx_j^{t-1}, vy_j^{t-1}, vz_j^{t-1}).

    1. Kinematic Trajectory Prediction: A constant-velocity kinematic vehicle model within a Kalman filter predicts the spatial state for each active trajectory at frame tt: xt,predj=xt−1j+vxt−1jx_{t,\text{pred}}^j = x_{t-1}^j + vx_{t-1}^j yt,predj=yt−1j+vyt−1jy_{t,\text{pred}}^j = y_{t-1}^j + vy_{t-1}^j zt,predj=zt−1j+vzt−1jz_{t,\text{pred}}^j = z_{t-1}^j + vz_{t-1}^j Tt,predj=(xt,predj,yt,predj,zt,predj,θt−1j,wt−1j,ht−1j,lt−1j,st−1j,vxt−1j,vyt−1j,vzt−1j)T_{t,\text{pred}}^j = (x_{t,\text{pred}}^j, y_{t,\text{pred}}^j, z_{t,\text{pred}}^j, \theta_{t-1}^j, w_{t-1}^j, h_{t-1}^j, l_{t-1}^j, s_{t-1}^j, vx_{t-1}^j, vy_{t-1}^j, vz_{t-1}^j)

    2. Data Association: An affinity matrix A∈RMt−1×NtA \in \mathbb{R}^{M_{t-1} \times N_t} is constructed where entry Ai,jA_{i,j} is the 3D Intersection-over-Union (3D IoU) between predicted trajectory ii and detection bounding box jj at frame tt. Bipartite matching is solved globally via the Hungarian algorithm.

    3. State Update: For each matched pair of predicted trajectory Tt,predmT_{t,\text{pred}}^m and detection DtkD_t^k, a Kalman filter updates the trajectory state: Ttm=KF(Tt,predm,Dtk)T_t^m = \text{KF}(T_{t,\text{pred}}^m, D_t^k)

    4. Trajectory Management: When an unmatched detection enters the field of view, it is initialized as a new trajectory candidate only if it can be consistently matched across subsequent frames (preventing false positive detections from initiating tracks). When an existing active trajectory is unmatched, it is retained for several frames before being terminated as dead (preventing missed detections from prematurely ending true tracks).

  2. Knowl 2 — Dual-Level Adversarial Domain Adaptation Architecture for Cooperative Perception

    model/method

    To mitigate the simulation-to-real domain shift between synthetic cooperative datasets (such as OPV2V as the source domain) and real-world vehicle-to-vehicle datasets (such as V2V4Real as the target domain), an adversarial domain adaptation framework utilizes two distinct domain discriminators:

    1. Feature-Level Domain Discriminator: Receives the fused multi-agent feature maps output by the intermediate fusion backbone. It consists of two convolutional layers with 3×33 \times 3 kernels, where the second layer projects the channel dimension to 11 to produce a spatial domain classification map distinguishing source from target representations.

    2. Object-Level Domain Discriminator: Receives the object classification confidence score map generated by the 3D detection classification head. It comprises three linear projection layers interleaved with Rectified Linear Unit (ReLU) activation functions.

    Both discriminators are trained using binary cross-entropy loss against domain source/target labels. Gradient Reversal Layers (GRL) are placed between the feature extraction/detection heads and the discriminators to invert the gradient signs during backpropagation, enforcing domain-invariant representations across connected vehicles.

  3. Knowl 3 — PointPillars Backbone and Detection Head Configuration for Cooperative 3D Object Detection

    experimental setup

    Cooperative 3D object detection baseline models evaluated on the V2V4Real benchmark are configured as follows:

    • PointPillars Backbone: The point cloud spatial voxel resolution is set to 0.4 m×0.4 m0.4\,\text{m} \times 0.4\,\text{m} along the xx and yy axes. The maximum number of points per voxel is set to 3232, and the maximum number of voxels per scene is capped at 32,00032,000.

    • Detection Heads: Two channel-wise 1×11 \times 1 convolutional layers are placed directly on top of the fused bird's-eye-view feature maps:

      • Regression Head: Outputs 7 degrees of freedom (x,y,z,w,l,h,θ)(x, y, z, w, l, h, \theta) representing center position (x,y,z)(x, y, z), dimensions (w,l,h)(w, l, h), and yaw orientation θ\theta for predefined 3D anchor boxes. It is optimized using smoothed ℓ1\ell_1 loss.
      • Classification Head: Outputs the object versus background confidence score for each predefined anchor box. It is optimized using focal loss.
  4. Knowl 4 — Impact of Data Augmentation on Cooperative 3D Object Detection

    data/table

    Ablating point cloud data augmentations (specifically random rotation, random flipping, and random scaling) causes marked performance drops across all cooperative 3D object detection methods on the V2V4Real dataset. Intermediate fusion architectures (CoBEVT, V2X-ViT, V2VNet, AttFuse, and F-Cooper) suffer larger relative performance degradations than single-vehicle No Fusion or Early Fusion, indicating higher sensitivity to data diversity and volume in complex fusion networks.

    The quantitative detection performance (Average Precision at IoU=0.5\text{IoU}=0.5) without data augmentation across distance ranges (0–30 m0\text{--}30\,\text{m}, 30–50 m30\text{--}50\,\text{m}, 50–100 m50\text{--}100\,\text{m}) and transmission message sizes (AM in megabytes, MB) under time-synchronized (Sync) and time-asynchronous (Async) settings is:

    Method Sync (AP@IoU=0.5) Async (AP@IoU=0.5) AM
    Overall 0-30m 30-50m 50-100m Overall 0-30m 30-50m 50-100m (MB)
    No Fusion 28.7 (-11.1) 50.0 22.9 4.9 28.7 (-11.1) 50.0 22.9 4.9 0
    Late Fusion 43.0 (-12.0) 55.1 34.4 31.9 40.9 (-9.3) 54.6 33.5 30.9 0.003
    Early Fusion 48.2 (-11.5) 64.3 33.2 34.0 41.0 (-11.1) 62.5 27.3 18.1 0.96
    F-Cooper 45.6 (-15.1) 65.3 35.3 25.9 37.6 (-16.0) 62.1 28.6 13.4 0.20
    V2VNet 49.0 (-15.5) 69.2 35.0 30.6 41.5 (-14.9) 65.5 32.4 13.7 0.20
    AttFuse 47.9 (-16.8) 67.9 34.6 26.8 40.8 (-16.9) 65.8 28.6 13.9 0.20
    V2X-ViT 48.9 (-16.0) 66.0 38.1 30.0 41.6 (-14.3) 62.8 32.9 17.5 0.20
    CoBEVT 51.1 (-15.4) 69.3 40.0 32.4 44.9 (-13.7) 65.2 35.6 19.8 0.20

    Numbers in parentheses indicate the performance decrease in Average Precision percentage points compared to the identical model trained with full data augmentation.

  5. Knowl 5 — Sim2Real Domain Adaptation Trade-offs Across Cooperative Fusion Models

    empirical result

    When transferring cooperative perception models trained on simulated data (OPV2V) to real-world driving data (V2V4Real) via adversarial domain adaptation:

    1. In highway driving scenarios with higher vehicle speeds and low vehicle density, all evaluated cooperative fusion architectures benefit from domain adaptation, with AttFuse and F-Cooper showing the largest qualitative improvements in true-positive bounding box alignment.
    2. In high-density intersection scenarios characterized by heavy occlusion, intermediate fusion methods (F-Cooper, V2X-ViT, and CoBEVT) achieve superior detection accuracy. However, feature-fusion methods like F-Cooper exhibit a higher rate of false-positive bounding box predictions after adaptation, whereas vision-transformer-based architectures like V2X-ViT maintain high detection precision with fewer false positives.

Coverage note — Visual sensor point cloud renderings and camera perspective snapshots across scenarios were omitted as their qualitative and quantitative insights are fully captured in the knowls.

References

  1. 1.Qi Chen, Xu Ma, Sihai Tang, Jingda Guo, Qing Yang, and Song Fu. F-cooper: Feature based cooperative perception for autonomous vehicle edge computing system using 3d point clouds. In Proceedings of the 4th ACM/IEEE Symposium on Edge Computing, pages 88–100, 2019. 4, 7, 8
  2. 2.Qi Chen, Sihai Tang, Qing Yang, and Song Fu. Cooper: Coop- erative perception for connected autonomous vehicles based on 3d point clouds. In 2019 IEEE 39th International Confer- ence on Distributed Computing Systems (ICDCS), pages 514– 524. IEEE, 2019. 1
  3. 3.Yaroslav Ganin and Victor Lempitsky. Unsupervised domain adaptation by backpropagation. In International conference on machine learning, pages 1180–1189. PMLR, 2015. 3
  4. 4.Alex H Lang, Sourabh Vora, Holger Caesar, Lubing Zhou, Jiong Yang, and Oscar Beijbom. Pointpillars: Fast encoders for object detection from point clouds. In Proceedings of the IEEE/CVF conference on computer vision and pattern recog- nition, pages 12697–12705, 2019. 1, 4
  5. 5.Zhengzhong Tu Xin Xia Ming-Hsuan Yang Jiaqi Ma Run- sheng Xu, Hao Xiang. V2x-vit: Vehicle-to-everything coop- erative perception with vision transformer. In Proceedings of the European Conference on Computer Vision (ECCV), 2022. 1, 4, 5, 6, 7, 8
  6. 6.Tsun-Hsuan Wang, Sivabalan Manivasagam, Ming Liang, Bin Yang, Wenyuan Zeng, and Raquel Urtasun. V2vnet: Vehicle- to-vehicle communication for joint perception and prediction. In European Conference on Computer Vision, pages 605–621. Springer, 2020. 1, 4, 5, 6, 7, 8
  7. 7.Runsheng Xu, Zhengzhong Tu, Hao Xiang, Wei Shao, Bolei Zhou, and Jiaqi Ma. Cobevt: Cooperative bird’s eye view se- mantic segmentation with sparse transformers. arXiv preprint arXiv:2207.02202, 2022. 1, 4, 5, 6, 7, 8
  8. 8.Runsheng Xu, Hao Xiang, Xin Xia, Xu Han, Jinlong Li, and Jiaqi Ma. Opv2v: An open benchmark dataset and fusion pipeline for perception with vehicle-to-vehicle communica- tion. In 2022 International Conference on Robotics and Au- tomation (ICRA), pages 2583–2589. IEEE, 2022. 1, 4, 7, 8
  9. 9.Yin Zhou and Oncel Tuzel. Voxelnet: End-to-end learning for point cloud based 3d object detection. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 4490–4499, 2018. 4

Citation

MLA
Xu, R., et al. “V2V4Real: A Real-world Large-scale Dataset for Vehicle-to-Vehicle Cooperative Perception”. arXiv, 2023, http://arxiv.org/abs/2303.07601v2.
APA
Xu, R., Xia, X., Li, J., Li, H., Zhang, S., Tu, Z., Meng, Z., Xiang, H., Dong, X., Song, R., Yu, H., Zhou, B., & Ma, J. (2023). V2V4Real: A Real-world Large-scale Dataset for Vehicle-to-Vehicle Cooperative Perception. arXiv. http://arxiv.org/abs/2303.07601v2
Chicago
Xu, R., X. Xia, J. Li, et al. 2023. “V2V4Real: A Real-world Large-scale Dataset for Vehicle-to-Vehicle Cooperative Perception”. arXiv. http://arxiv.org/abs/2303.07601v2.
Harvard
Xu, R. et al. (2023) “V2V4Real: A Real-world Large-scale Dataset for Vehicle-to-Vehicle Cooperative Perception”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2303.07601v2.
Vancouver
1. Xu R, Xia X, Li J, et al (2023) V2V4Real: A Real-world Large-scale Dataset for Vehicle-to-Vehicle Cooperative Perception. arXiv

BibTeX

@article{xu2023v2v4real,
  title = {V2V4Real: A Real-world Large-scale Dataset for Vehicle-to-Vehicle Cooperative Perception},
  author = {Xu, Runsheng and Xia, Xin and Li, Jinlong and Li, Hanzhao and Zhang, Shuo and Tu, Zhengzhong and Meng, Zonglin and Xiang, Hao and Dong, Xiaoyu and Song, Rui and Yu, Hongkai and Zhou, Bolei and Ma, Jiaqi},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2303.07601v2},
  eprint = {2303.07601}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE