DI-V2X: Learning Domain-Invariant Representation for Vehicle-Infrastructure Collaborative 3D Object Detection

Xiang LiJunbo YinWei LiChengzhong XuRuigang YangJianbing Shen

article2024AAAI56 citations

Presents a teacher-student distillation framework for vehicle-infrastructure collaborative 3D object detection that overcomes sensor discrepancy domain gaps across agents through domain-mixed data augmentation, progressive distillation, and adaptive feature fusion.

Listen

Vehicle-to-Everything collaborative perception enables autonomous vehicles to share sensor data with roadside infrastructure, expanding detection range and eliminating blind spots. However, vehicles and roadside units typically use fundamentally different laser sensor setups, such as mechanical sensors on vehicles versus high-density solid-state sensors on road infrastructure. This hardware mismatch creates a significant data-level gap, causing standard collaborative systems to degrade in performance when directly combining data streams.

The article demonstrates that this sensor mismatch can be resolved using a domain-invariant framework called DI-V2X. The primary objective is to evaluate how effectively a progressive teacher-student training structure can align disparate sensor data into unified, robust representations without requiring costly high-bandwidth raw data transfers during operation.

To achieve this, the authors evaluated DI-V2X on two major benchmarks: the real-world DAIR-V2X dataset comprising 9,000 synchronized frame pairs across 100 scenes, and the synthetic V2XSet dataset containing multi-agent driving scenarios. The framework combines three main elements during training: a domain-mixing instance augmentation method that creates balanced object samples across sensors, a two-stage progressive distillation process that guides student networks using an ideal fused teacher model across overlapping and non-overlapping spatial regions, and an adaptive fusion module that corrects real-world positioning offsets while weighting domain and spatial features.

The experimental findings show clear improvements over existing collaborative perception methods. First, DI-V2X established new state-of-the-art results, achieving 66.16% average precision at strict detection thresholds on the real-world dataset, outperforming the previous leading method by 5.76 percentage points. Second, on the multi-agent synthetic benchmark, the system achieved 82.7% precision, outperforming existing attention-based baselines by 11.5 percentage points. Third, ablation analyses confirmed that all three framework components contribute positively, with adaptive fusion providing the largest individual performance gain of 8.22 percentage points. Finally, tests under degraded conditions showed that the framework retains strong detection performance even when infrastructure communication drops and the vehicle must operate independently.

These results demonstrate that explicitly resolving sensor disparities unlocks the safety and reliability benefits promised by connected infrastructure. By utilizing intermediate feature fusion rather than raw data sharing, the approach preserves low communication bandwidth while delivering superior accuracy and resilience to alignment noise. This provides an effective operational pathway for integrating heterogeneous smart-city roadside sensors with vehicle fleets.

Engineering and deployment teams should consider adopting progressive domain-distillation and calibration-aware feature fusion in collaborative perception pipelines. Prior to full-scale deployment, organizations should conduct live pilot testing to validate system performance across diverse weather conditions, communication latency variations, and differing sensor configurations not captured in the current benchmark datasets.

arXiv: 2312.15742

No sufficiently relevant recommendations were found.

Cover for DI-V2X: Learning Domain-Invariant Representation for Vehicle-Infrastructure Collaborative 3D Object Detection

Abstract

Vehicle-to-Everything (V2X) collaborative perception has recently gained significant attention due to its capability to enhance scene understanding by integrating information from various agents, e.g., vehicles, and infrastructure. However, current works often treat the information from each agent equally, ignoring the inherent domain gap caused by the utilization of different LiDAR sensors of each agent, thus leading to suboptimal performance. In this paper, we propose DI-V2X, that aims to learn Domain-Invariant representations through a new distillation framework to mitigate the domain discrepancy in the context of V2X 3D object detection. DI-V2X comprises three essential components: a domain-mixing instance augmentation (DMA) module, a progressive domain-invariant distillation (PDD) module, and a domain-adaptive fusion (DAF) module. Specifically, DMA builds a domain-mixing 3D instance bank for the teacher and student models during training, resulting in aligned data representation. Next, PDD encourages the student models from different domains to gradually learn a domain-invariant feature representation towards the teacher, where the overlapping regions between agents are employed as guidance to facilitate the distillation process. Furthermore, DAF closes the domain gap between the students by incorporating calibration-aware domain-adaptive attention. Extensive experiments on the challenging DAIR-V2X and V2XSet benchmark datasets demonstrate DI-V2X achieves remarkable performance, outperforming all the previous V2X models. Code is available at https://github.com/Serenos/DI-V2X.

Table of Contents

  • 1 Introduction
  • 2 Related Works
  • V2X Collaborative Perception
  • Domain-Invariant Representation in Point Cloud
  • 3 The Proposed DI-V2X Framework
  • Overall Architecture
  • Domain-Mixing Instance Augmentation
  • Progressive Domain-Invariant Distillation
  • Distillation Before Fusion.
  • Distillation After Fusion.
  • Domain-Adaptive Fusion
  • 4 Experiments
  • Dataset
  • Experimental Setup
  • Main Results
  • Domain Generalization Ability of DI-V2X.
  • Ablation Study.
  • 5 Conclusion
  • 6 Acknowledgments
  • References

Knowls

  1. Knowl 1 — DI-V2X Framework for Collaborative 3D Object Detection

    model/method

    DI-V2X is a teacher-student distillation framework designed for intermediate-fusion Vehicle-to-Everything (V2X) collaborative 3D object detection to overcome data-level domain gaps between heterogeneous LiDAR sensors (e.g., mechanical vehicle-side LiDAR and solid-state roadside LiDAR).

    Let Pv∈RNv×4P_v \in \mathbb{R}^{N_v \times 4} and Pi∈RNi×4P_i \in \mathbb{R}^{N_i \times 4} denote temporally synchronized 3D point clouds from the vehicle and infrastructure agents, respectively. By transforming PiP_i into the vehicle coordinate system using extrinsic pose matrix T(i→v)∈R4×4T_{(i \to v)} \in \mathbb{R}^{4 \times 4}, an early-fused holistic point cloud PeP_e is constructed. During training, inputs are augmented via Domain-Mixing instance Augmentation (DMA) to yield {Pv′,Pi′,Pe′}\{P_v', P_i', P_e'\}.

    A 3D detection backbone (such as PointPillars with a voxel encoder) extracts bird's-eye-view (BEV) feature maps Bt∈RH×W×CB_t \in \mathbb{R}^{H \times W \times C} for the teacher from Pe′P_e', and Bv,Bi∈RH×W×CB_v, B_i \in \mathbb{R}^{H \times W \times C} for the student models from Pv′,Pi′P_v', P_i' via a weight-shared 3D encoder. A Domain-Adaptive Fusion (DAF) module aggregates student features into a unified BEV representation Bf∈RH×W×CB_f \in \mathbb{R}^{H \times W \times C}. The Progressive Domain-invariant Distillation (PDD) module aligns student representations with BtB_t both before fusion (in non-overlapping sensor regions) and after fusion (in overlapping sensor regions), as well as aligning final classification and regression predictions.

    During inference, the early-fused teacher model is discarded, and only the student intermediate fusion pipeline is deployed to preserve communication efficiency.

  2. Knowl 2 — Domain-Mixing Instance Augmentation (DMA)

    model/method

    Domain-Mixing instance Augmentation (DMA) mitigates data distribution disparities between vehicle and infrastructure LiDAR sensors by constructing an aligned, mixed ground-truth 3D instance database DmixedD_{mixed}.

    Given ground-truth 3D bounding boxes Bgt={bk}B_{gt} = \{b_k\}, points associated with each instance kk from the vehicle point cloud PvP_v and the coordinate-aligned infrastructure point cloud PiP_i are concatenated to form an early-fused instance pk=Concat(pkv,pki)∈RNk×4p_k = \text{Concat}(p_k^v, p_k^i) \in \mathbb{R}^{N_k \times 4}, where pkv∈Pvp_k^v \in P_v and pki∈Pip_k^i \in P_i. Let NkvN_k^v and NkiN_k^i be the point counts from the vehicle and infrastructure sides for instance kk. Instances are categorized into three sets using lower and upper ratio thresholds τl\tau_l and τh\tau_h:

    Di={pk  |  NkvNkv+Nki<τl}D_i = \left\{ p_k \;\middle|\; \frac{N_k^v}{N_k^v + N_k^i} < \tau_l \right\}

    Df={pk  |  τl≤NkvNkv+Nki≤τh}D_f = \left\{ p_k \;\middle|\; \tau_l \le \frac{N_k^v}{N_k^v + N_k^i} \le \tau_h \right\}

    Dv={pk  |  NkvNkv+Nki>τh}D_v = \left\{ p_k \;\middle|\; \frac{N_k^v}{N_k^v + N_k^i} > \tau_h \right\}

    The unified instance bank is Dmixed=Di∪Df∪DvD_{mixed} = D_i \cup D_f \cup D_v. During training, instances are randomly sampled from DmixedD_{mixed} with sampling probabilities of 0.80.8 for fused instances DfD_f, 0.10.1 for vehicle instances DvD_v, and 0.10.1 for infrastructure instances DiD_i, and inserted into the point cloud scenes of both teacher and student models.

  3. Knowl 3 — Progressive Domain-Invariant Distillation (PDD) and Loss Formulation

    model/method

    Progressive Domain-invariant Distillation (PDD) transfers knowledge from the early-fused teacher BEV feature Bt∈RH×W×CB_t \in \mathbb{R}^{H \times W \times C} to student representations across two stages guided by spatial overlap masks.

    Let the vehicle perception region in vehicle coordinates be Av=(0,0,2Rx,2Ry,0)A_v = (0, 0, 2R_x, 2R_y, 0), and let the infrastructure perception region transformed to vehicle coordinates be Ai=(xi,yi,2Rx,2Ry,θi)A_i = (x_i, y_i, 2R_x, 2R_y, \theta_i), where Rx,RyR_x, R_y are maximum perception ranges along the XX and YY axes. The overlapping polygon is Poverlap=Intersection(Av,Ai)P_{overlap} = \text{Intersection}(A_v, A_i). Downsampled to feature dimensions H×WH \times W, the overlapping mask M∈{0,1}H×WM \in \{0, 1\}^{H \times W} is defined by M(i,j)=1M(i, j) = 1 if coordinate (i,j)∈Poverlap(i, j) \in P_{overlap} and 00 otherwise. The non-overlapping mask is M~=1−M\tilde{M} = 1 - M.

    1. Pre-Fusion Distillation Loss: Guided by non-overlapping masks M~v,M~i\tilde{M}_v, \tilde{M}_i, each student aligns with the teacher within its independent perception area:

    Lda=1HW∑m=1H∑n=1W∣Bt(m,n)−Bv(m,n)∣⋅M~v(m,n)+1HW∑m=1H∑n=1W∣Bt(m,n)−Bi(m,n)∣⋅M~i(m,n)\mathcal{L}_{da} = \frac{1}{HW}\sum_{m=1}^H \sum_{n=1}^W |B_t(m, n) - B_v(m, n)| \cdot \tilde{M}_v(m, n) + \frac{1}{HW}\sum_{m=1}^H \sum_{n=1}^W |B_t(m, n) - B_i(m, n)| \cdot \tilde{M}_i(m, n)

    1. Post-Fusion Distillation Loss: Enforces the fused student BEV feature Bf∈RH×W×CB_f \in \mathbb{R}^{H \times W \times C} to match the teacher feature BtB_t within the overlapping area MvM_v:

    Lf=1HW∑m=1H∑n=1W∣Bt(m,n)−Bf(m,n)∣⋅Mv(m,n)\mathcal{L}_f = \frac{1}{HW}\sum_{m=1}^H \sum_{n=1}^W |B_t(m, n) - B_f(m, n)| \cdot M_v(m, n)

    1. Prediction-Level Distillation Loss: Aligns student detection outputs (CS={cks}k=1K,RS={rks}k=1KC_S = \{c_k^s\}_{k=1}^K, R_S = \{r_k^s\}_{k=1}^K) with teacher outputs (CT={ck}k=1K,RT={rk}k=1KC_T = \{c_k\}_{k=1}^K, R_T = \{r_k\}_{k=1}^K) over KK predictions:

    Lp=1K∑k=1K(∣ck−cks∣+∣rk−rks∣)\mathcal{L}_p = \frac{1}{K}\sum_{k=1}^K \left( |c_k - c_k^s| + |r_k - r_k^s| \right)

    The total training loss is:

    L=Ldetect+λkd(Lda+Lf+Lp)\mathcal{L} = \mathcal{L}_{detect} + \lambda_{kd}(\mathcal{L}_{da} + \mathcal{L}_f + \mathcal{L}_p)

    where Ldetect\mathcal{L}_{detect} is the task detection loss and λkd=1.0\lambda_{kd} = 1.0 is the distillation weighting factor.

  4. Knowl 4 — Domain-Adaptive Fusion (DAF) Module

    model/method

    The Domain-Adaptive Fusion (DAF) module corrects relative pose errors and adaptively fuses vehicle BEV features Bv∈RH×W×CB_v \in \mathbb{R}^{H \times W \times C} and infrastructure BEV features Bi∈RH×W×CB_i \in \mathbb{R}^{H \times W \times C}.

    1. Pose Calibration Offset: A convolutional network estimates a spatial offset map to compensate for relative calibration and synchronization noise:

    Δ(i→v)=Conv(Concat(Bv,Bi))∈RH×W×2\Delta_{(i \to v)} = \text{Conv}(\text{Concat}(B_v, B_i)) \in \mathbb{R}^{H \times W \times 2}

    The infrastructure feature map is warped according to the predicted offsets:

    Bi′(pk)=Bi(pk+Δ(i→v)(pk)),0≤k<HWB_i'(p_k) = B_i(p_k + \Delta_{(i \to v)}(p_k)), \quad 0 \le k < HW

    where pk∈R2p_k \in \mathbb{R}^2 represents the 2D spatial coordinate of grid location kk.

    1. Domain-Adaptive Attention: The aligned features are concatenated into Bcat=Concat(Bv,Bi′)∈RH×W×C×2B_{cat} = \text{Concat}(B_v, B_i') \in \mathbb{R}^{H \times W \times C \times 2}. Domain-specific weighting is computed via a 3×33 \times 3 convolution and two residual 1×11 \times 1 convolutions followed by softmax normalization:

    Ad=Softmax(Conv(Bcat))∈RH×W×C×2A_d = \text{Softmax}(\text{Conv}(B_{cat})) \in \mathbb{R}^{H \times W \times C \times 2}

    1. Spatially-Adaptive Attention: Saliency across the spatial grid is computed by combining multi-kernel convolutions and max-pooling over channel and domain dimensions:

    As=Conv(Bcat)+max⁡(Bcat)∈RH×WA_s = \text{Conv}(B_{cat}) + \max(B_{cat}) \in \mathbb{R}^{H \times W}

    1. Feature Fusion: The final fused feature BfB_f is obtained by applying domain and spatial modulation:

    Bf=Conv(Ad⊙As⋅Bcat)∈RH×W×CB_f = \text{Conv}(A_d \odot A_s \cdot B_{cat}) \in \mathbb{R}^{H \times W \times C}

    where ⊙\odot denotes element-wise multiplication with channel broadcasting and ⋅\cdot denotes standard element-wise multiplication.

  5. Knowl 5 — 3D Detection Benchmark Performance on DAIR-V2X

    data/table

    The detection performance of DI-V2X and baseline collaborative perception paradigms was evaluated on the real-world DAIR-V2X validation dataset. All methods used PointPillars (PP) as the base 3D detector.

    Fusion Model BEV [email protected] (%) BEV [email protected] (%)
    No Fusion PP 65.21 54.26
    Late Fusion PP-LF 63.68 45.33
    Early Fusion PP-EF 76.55 64.10
    DI-V2X (teacher) 79.81 67.08
    Intermediate Fusion PP-IF 72.85 52.67
    V2VNet 70.97 47.63
    DiscoNet 73.87 59.61
    OPV2V 73.30 55.33
    V2X-ViT 71.81 54.94
    Where2comm 74.09 59.66
    CoAlign 74.60 60.40
    DI-V2X (student) 78.82 (+4.22) 66.16 (+5.76)

    The intermediate fusion DI-V2X student outperforms all prior intermediate fusion approaches, exceeding the previous state-of-the-art model CoAlign by 4.22% in [email protected] and 5.76% in [email protected]. Furthermore, DI-V2X student surpasses standard early fusion (PP-EF) by 2.27% [email protected] and 2.06% [email protected] while avoiding raw point cloud transmission.

  6. Knowl 6 — 3D Detection Benchmark Performance on V2XSet

    data/table

    The performance of DI-V2X was benchmarked against existing collaborative perception methods on the simulated V2XSet validation dataset across multiple collaborating agents (2 to 7 agents with 36-beam LiDAR).

    Fusion Model [email protected] (%) [email protected] (%)
    No Fusion PP 60.6 40.2
    Late Fusion PP-LF 72.7 62.0
    Early Fusion PP-EF 81.9 71.9
    Intermediate Fusion PP-IF 76.9 49.2
    DiscoNet 84.4 69.5
    Where2comm 85.5 65.4
    V2X-ViT 88.2 71.2
    DI-V2X 92.7 82.7

    DI-V2X outperforms V2X-ViT by 4.5% [email protected] and 11.5% [email protected], and exceeds early fusion (PP-EF) by 10.8% in both [email protected] and [email protected].

  7. Knowl 7 — Domain Generalization Performance Under Sensor Dropout

    data/table

    To evaluate domain invariance and resilience against communication interruptions, models trained on joint vehicle and infrastructure data (V. & I.) on DAIR-V2X were evaluated under single-agent (vehicle-only or infrastructure-only) and joint-agent inference scenarios.

    Model Train Set Test Set ([email protected] %) Avg. [email protected] (%)
    Veh. Inf. V. I.
    PP Veh. 73.08 31.53 52.62 52.41
    PP Inf. 32.21 65.80 26.39 41.46
    PP-LF V. I. 63.62 65.66 45.33 58.20
    PP-EF V. I. 73.19 47.44 64.10 61.57
    PP-IF V. I. 68.55 7.45 52.67 26.22
    PP-KD V. I. 69.91 17.82 53.91 47.11
    PP-CDT V. I. 66.22 41.73 55.90 54.61
    Ours-Teacher V. I. 76.74 53.34 67.08 65.72
    Ours-Student V. I. 73.04 45.91 66.16 61.70

    When testing on vehicle data only during communication dropouts, the DI-V2X student achieves 73.04% [email protected], matching a single-domain model trained solely on vehicle data (73.08%). When testing on infrastructure data only, it achieves 45.91% [email protected], outperforming baseline intermediate fusion (PP-IF at 7.45%), knowledge distillation (PP-KD at 17.82%), and cross-domain transformer (PP-CDT at 41.73%).

  8. Knowl 8 — Ablation Analysis of DI-V2X Architectural Modules

    data/table

    The individual and cumulative contributions of Domain-Mixing instance Augmentation (DMA), Progressive Domain-invariant Distillation (PDD), and Domain-Adaptive Fusion (DAF) were ablated on the DAIR-V2X validation set using naive intermediate fusion (PP-IF) as the baseline.

    DMA PDD DAF [email protected] (%) [email protected] (%)
    - - - 72.85 52.67
    ✓ - - 73.90 54.02
    - ✓ - 74.17 55.89
    - - ✓ 75.76 60.89
    - ✓ ✓ 77.37 61.66
    ✓ ✓ ✓ 78.82 66.16
    • DMA individually improves [email protected] by 1.35% (from 52.67% to 54.02%) by enriching multi-source instance diversity.
    • PDD individually improves [email protected] by 3.22% (to 55.89%) via masked teacher-to-student feature transfer.
    • DAF individually improves [email protected] by 8.22% (to 60.89%) by correcting pose misalignments and focusing on salient domain and spatial features.
    • Combining all three components yields a net gain of +5.97% [email protected] and +13.49% [email protected] over baseline PP-IF.
  9. Knowl 9 — DI-V2X Implementation and Training Configuration

    experimental setup

    The implementation and training specifications for DI-V2X are configured as follows:

    1. Point Cloud Range and Voxelization: The detection range in vehicle coordinates is set to [−100,100][-100, 100] meters along the XX-axis, [−40,40][-40, 40] meters along the YY-axis, and [−3.5,1.5][-3.5, 1.5] meters along the ZZ-axis. The voxel size is set to [0.4,0.4,5.0][0.4, 0.4, 5.0] meters along the (X,Y,Z)(X, Y, Z) axes, resulting in a BEV feature map of size 252×100×256252 \times 100 \times 256.

    2. Hyperparameters: DMA instance classification thresholds are τl=0.2\tau_l = 0.2 and τh=0.8\tau_h = 0.8, with sampling probabilities of 0.80.8 for fused instances DfD_f, 0.10.1 for vehicle instances DvD_v, and 0.10.1 for infrastructure instances DiD_i. The distillation loss weight is set to λkd=1.0\lambda_{kd} = 1.0.

    3. Data Augmentation and Training: Scene-level augmentations include random flip, rotation, and scaling. Student models are trained on 4 NVIDIA Tesla V100 GPUs with a batch size of 4 for 40 epochs using PointPillars as the base 3D detector.

Coverage note — None was omitted. All contributed methods (DMA, PDD, DAF), mathematical loss formulations, benchmark results on DAIR-V2X and V2XSet, generalization experiments, and ablation analyses are represented.

References

  1. 1.Alonso, I.; Riazuelo, L.; Montesano, L.; and Murillo, A. C. 2021. Domain Adaptation in LiDAR Semantic Segmentation by Aligning Class Distributions. arXiv:2010.12239.
  2. 2.Chen, Q.; Ma, X.; Tang, S.; Guo, J.; Yang, Q.; and Fu, S. 2019a. F-Cooper: Feature Based Cooperative Perception for Autonomous Vehicle Edge Computing System Using 3D Point Clouds. In Proceedings of the 4th ACM/IEEE Symposium on Edge Computing.
  3. 3.Chen, Q.; Tang, S.; Yang, Q.; and Fu, S. 2019b. Cooper: Cooperative perception for connected autonomous vehicles based on 3d point clouds. In 2019 IEEE 39th International Conference on Distributed Computing Systems.
  4. 4.Chen, Y.; Hu, V. T.; Gavves, E.; Mensink, T.; and Mettes. 2020. PointMixup: Augmentation for Point Clouds. In ECCV.
  5. 5.Hu, Y.; Fang, S.; Lei, Z.; Zhong, Y.; and Chen, S. 2022. Where2comm: Communication-efficient collaborative perception via spatial confidence maps. NeurIPS.
  6. 6.Lang, A. H.; Vora, S.; Caesar, H.; Zhou, L.; Yang, J.; and Beijbom, O. 2019. PointPillars: Fast Encoders for Object Detection From Point Clouds. In CVPR.
  7. 7.Langer, F.; Milioto, A.; Haag, A.; Behley, J.; and Stachniss, C. 2020. Domain Transfer for Semantic Segmentation of LiDAR Data using Deep Neural Networks. In IROS.
  8. 8.Lee, D.; Lee, J.; Lee, J.; Lee, H.; Lee, M.; Woo, S.; and Lee, S. 2021. Regularization Strategy for Point Cloud via Rigidly Mixed Sample. In CVPR.
  9. 9.Li, Y.; Ren, S.; Wu, P.; Chen, S.; Feng, C.; and Zhang, W. 2021. Learning Distilled Collaboration Graph for Multi-Agent Perception. In NeurIPS.
  10. 10.Liu, Y.-C.; Tian, J.; Glaser, N.; and Kira, Z. 2020a. When2com: Multi-Agent Perception via Communication Graph Grouping. In CVPR.
  11. 11.Liu, Y.-C.; Tian, J.; Ma, C.-Y.; Glaser, N.; Kuo, C.-W.; and Kira, Z. 2020b. Who2com: Collaborative Perception via Learnable Handshake Communication. In ICRA.
  12. 12.Lu, Y.; Li, Q.; Liu, B.; Dianati, M.; Feng, C.; Chen, S.; and Wang, Y. 2023. Robust collaborative 3d object detection in presence of pose errors. In ICRA.
  13. 13.Nekrasov, A.; Schult, J.; Litany, O.; Leibe, B.; and Engelmann, F. 2021. Mix3D: Out-of-Context Data Augmentation for 3D Scenes. In 2021 International Conference on 3D Vision (3DV).
  14. 14.Ryu, K.; Hwang, S.; and Park, J. 2023. Instant Domain Augmentation for LiDAR Semantic Segmentation. In CVPR.
  15. 15.Triess, L. T.; Dreissig, M.; Rist, C. B.; and Marius Zöllner, J. 2021. A Survey on Deep Domain Adaptation for LiDAR Perception. In 2021 IEEE Intelligent Vehicles Symposium Workshops (IV Workshops).
  16. 16.Wang, T.-H.; Manivasagam, S.; Liang, M.; Yang, B.; Zeng, W.; and Urtasun, R. 2020. V2vnet: Vehicle-to-vehicle communication for joint perception and prediction. In ECCV.
  17. 17.Xu, R.; Chen, W.; Xiang, H.; Xia, X.; Liu, L.; and Ma, J. 2023a. Model-agnostic multi-agent perception framework. In ICRA.
  18. 18.Xu, R.; Li, J.; Dong, X.; Yu, H.; and Ma, J. 2023b. Bridging the domain gap for multi-agent perception. In ICRA.
  19. 19.Xu, R.; Xiang, H.; Tu, Z.; Xia, X.; Yang, M.-H.; and Ma, J. 2022a. V2x-vit: Vehicle-to-everything cooperative perception with vision transformer. In ECCV.
  20. 20.Xu, R.; Xiang, H.; Xia, X.; Han, X.; Li, J.; and Ma, J. 2022b. Opv2v: An open benchmark dataset and fusion pipeline for perception with vehicle-to-vehicle communication. In ICRA.
  21. 21.Yan, Y.; Mao, Y.; and Li, B. 2018. SECOND: Sparsely Embedded Convolutional Detection. Sensors, 18(10): 3337.
  22. 22.Yu, H.; Luo, Y.; Shu, M.; Huo, Y.; Yang, Z.; Shi, Y.; Guo, Z.; Li, H.; Hu, X.; Yuan, J.; et al. 2022. Dair-v2x: A large-scale dataset for vehicle-infrastructure cooperative 3d object detection. In CVPR.
  23. 23.Yu, H.; Yang, W.; Ruan, H.; Yang, Z.; Tang, Y.; Gao, X.; Hao, X.; Shi, Y.; Pan, Y.; Sun, N.; Song, J.; Yuan, J.; Luo, P.; and Nie, Z. 2023. V2X-Seq: A large-scale sequential dataset for vehicle-infrastructure cooperative perception and forecasting. In CVPR.
  24. 24.Zhang, H.; Cissé, M.; Dauphin, Y. N.; and Lopez-Paz, D. 2018. mixup: Beyond Empirical Risk Minimization. In ICLR.

Citation

MLA
Xiang, L., et al. “DI-V2X: Learning Domain-Invariant Representation for Vehicle-Infrastructure Collaborative 3D Object Detection”. arXiv, 2023, http://arxiv.org/abs/2312.15742v1.
APA
Xiang, L., Yin, J., Li, W., Xu, C.-Z., Yang, R., & Shen, J. (2023). DI-V2X: Learning Domain-Invariant Representation for Vehicle-Infrastructure Collaborative 3D Object Detection. arXiv. http://arxiv.org/abs/2312.15742v1
Chicago
Xiang, L., J. Yin, W. Li, C.-Z. Xu, R. Yang, and J. Shen. 2023. “DI-V2X: Learning Domain-Invariant Representation for Vehicle-Infrastructure Collaborative 3D Object Detection”. arXiv. http://arxiv.org/abs/2312.15742v1.
Harvard
Xiang, L. et al. (2023) “DI-V2X: Learning Domain-Invariant Representation for Vehicle-Infrastructure Collaborative 3D Object Detection”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2312.15742v1.
Vancouver
1. Xiang L, Yin J, Li W, Xu C-Z, Yang R, Shen J (2023) DI-V2X: Learning Domain-Invariant Representation for Vehicle-Infrastructure Collaborative 3D Object Detection. arXiv

BibTeX

@article{xiang2023v2x,
  title = {DI-V2X: Learning Domain-Invariant Representation for Vehicle-Infrastructure Collaborative 3D Object Detection},
  author = {Xiang, Li and Yin, Junbo and Li, Wei and Xu, Cheng-Zhong and Yang, Ruigang and Shen, Jianbing},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2312.15742v1},
  eprint = {2312.15742}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF