DI-V2X: Learning Domain-Invariant Representation for Vehicle-Infrastructure Collaborative 3D Object Detection
Xiang LiJunbo YinWei LiChengzhong XuRuigang YangJianbing Shen
Presents a teacher-student distillation framework for vehicle-infrastructure collaborative 3D object detection that overcomes sensor discrepancy domain gaps across agents through domain-mixed data augmentation, progressive distillation, and adaptive feature fusion.
Vehicle-to-Everything collaborative perception enables autonomous vehicles to share sensor data with roadside infrastructure, expanding detection range and eliminating blind spots. However, vehicles and roadside units typically use fundamentally different laser sensor setups, such as mechanical sensors on vehicles versus high-density solid-state sensors on road infrastructure. This hardware mismatch creates a significant data-level gap, causing standard collaborative systems to degrade in performance when directly combining data streams.
The article demonstrates that this sensor mismatch can be resolved using a domain-invariant framework called DI-V2X. The primary objective is to evaluate how effectively a progressive teacher-student training structure can align disparate sensor data into unified, robust representations without requiring costly high-bandwidth raw data transfers during operation.
To achieve this, the authors evaluated DI-V2X on two major benchmarks: the real-world DAIR-V2X dataset comprising 9,000 synchronized frame pairs across 100 scenes, and the synthetic V2XSet dataset containing multi-agent driving scenarios. The framework combines three main elements during training: a domain-mixing instance augmentation method that creates balanced object samples across sensors, a two-stage progressive distillation process that guides student networks using an ideal fused teacher model across overlapping and non-overlapping spatial regions, and an adaptive fusion module that corrects real-world positioning offsets while weighting domain and spatial features.
The experimental findings show clear improvements over existing collaborative perception methods. First, DI-V2X established new state-of-the-art results, achieving 66.16% average precision at strict detection thresholds on the real-world dataset, outperforming the previous leading method by 5.76 percentage points. Second, on the multi-agent synthetic benchmark, the system achieved 82.7% precision, outperforming existing attention-based baselines by 11.5 percentage points. Third, ablation analyses confirmed that all three framework components contribute positively, with adaptive fusion providing the largest individual performance gain of 8.22 percentage points. Finally, tests under degraded conditions showed that the framework retains strong detection performance even when infrastructure communication drops and the vehicle must operate independently.
These results demonstrate that explicitly resolving sensor disparities unlocks the safety and reliability benefits promised by connected infrastructure. By utilizing intermediate feature fusion rather than raw data sharing, the approach preserves low communication bandwidth while delivering superior accuracy and resilience to alignment noise. This provides an effective operational pathway for integrating heterogeneous smart-city roadside sensors with vehicle fleets.
Engineering and deployment teams should consider adopting progressive domain-distillation and calibration-aware feature fusion in collaborative perception pipelines. Prior to full-scale deployment, organizations should conduct live pilot testing to validate system performance across diverse weather conditions, communication latency variations, and differing sensor configurations not captured in the current benchmark datasets.
- Paper: Where2comm: Communication-Efficient Collaborative Perception via Spatial Confidence Maps, Yue Hu et al. (2022). Where2comm introduces bandwidth-efficient spatial confidence intermediate fusion for collaborative 3D detection on DAIR-V2X, providing the direct baseline and foundation for DI-V2X's collaborative framework.
- Paper: V2V4Real: A Real-World Large-Scale Dataset for Vehicle-to-Vehicle Cooperative Perception, Runsheng Xu et al. (2023). V2V4Real establishes real-world multi-agent cooperative perception benchmarks and baseline intermediate fusion strategies essential for understanding domain gaps across vehicular platforms.
- Paper: BEVFusion: Multi-Task Multi-Sensor Fusion with Unified Bird's-Eye View Representation, Zhijian Liu et al. (2022). BEVFusion establishes the foundational unified bird's-eye-view representation and efficient pooling pipeline that DI-V2X adapts for multi-agent feature alignment.
- Paper: BEVHeight: A Robust Framework for Vision-based Roadside 3D Object Detection, Lei Yang et al. (2023). BEVHeight addresses the specific viewpoint and sensor disparities inherent in elevated roadside unit perception, directly motivating DI-V2X's vehicle-to-infrastructure domain alignment.
- Paper: 3D-VField: Adversarial Augmentation of Point Clouds for Domain Generalization in 3D Object Detection, Alexander Lehner et al. (2022). 3D-VField develops sensor-aware geometric augmentation techniques for 3D point clouds, which directly precede DI-V2X's domain-mixing instance augmentation strategy.
- Paper: PV-RCNN: Point-Voxel Feature Set Abstraction for 3D Object Detection, Shaoshuai Shi et al. (2019). PV-RCNN supplies the point-voxel 3D object detection backbone and feature abstraction mechanisms utilized in standard LiDAR collaborative perception pipelines.
- Paper: Domain Adaptive Faster R-CNN for Object Detection in the Wild, Yuhua Chen et al. (2018). Domain Adaptive Faster R-CNN provides the foundational dual-level adversarial alignment framework for mitigating domain shifts in object detection.
- Paper: Domain Generalization: A Survey, Kaiyang Zhou et al. (2021). This survey provides a comprehensive theoretical taxonomy of domain alignment and feature-mixing augmentation methods fundamental to domain generalization.
No sufficiently relevant recommendations were found.
