RADIANT: Radar-Image Association Network for 3D Object Detection

Yunfei LongAbhinav KumarDaniel D. MorrisXiaoming LiuMarcos CastroPunarjay Chakravarty

article2023AAAI60 citations

Proposes a radar-camera fusion framework that predicts 3D spatial offsets between radar returns and object centers, effectively resolving radar-to-object association errors and correcting depth inaccuracies in monocular 3D detection without retraining the base camera model.

Listen

Reliable three-dimensional object detection is essential for advanced automotive safety systems, robotics, and collision-avoidance maneuvers. While standard single-camera systems achieve high accuracy in two-dimensional image recognition, their ability to pinpoint three-dimensional locations degrades sharply due to severe depth estimation errors. Direct depth sensors like light detection and ranging (LiDAR) offer high precision but remain prohibitively expensive for mass deployment. Automotive radar provides an inexpensive, robust, and widely installed alternative; however, radar data is inherently sparse, lacks height information, and exhibits measurement noise. Furthermore, radar return points do not usually align with actual object centers, creating a challenging association problem that previously undermined the benefits of fusing radar with camera systems.

The article demonstrates a sensor-fusion framework called RADIANT (Radar-Image Association Network) that improves three-dimensional object detection by pairing existing single-camera detection pipelines with automotive radar. The primary objective is to resolve the spatial misalignment between radar points and physical object centers, enabling radar-derived depth to correct image-based distance estimates without requiring changes to underlying camera architectures.

To accomplish this, the authors engineered a dual-branch neural network evaluated on the large-scale, real-world nuScenes autonomous driving dataset, which spans over 40,000 traffic scenes. Rather than attempting to retrain the image network, the system freezes a pre-trained camera model and operates a parallel radar branch. This branch projects radar points into the image plane, borrows contextual features from the camera, and predicts the exact geometric offsets between radar reflections and three-dimensional object centers. A dedicated depth-weighting neural network then evaluates whether the camera or radar depth is more reliable for each detected object, dynamically calculating a fused depth estimate.

The evaluation produced several decisive findings. First, predicting geometric offsets between radar points and object centers drastically reduces depth error: the radar branch reduced overall depth prediction error to 0.79 meters compared to 3.42 meters for the camera-only baseline, with the greatest advantage occurring at distances exceeding 30 meters. Second, when integrated with leading monocular vision models, RADIANT substantially outperformed stand-alone camera detection, decreasing translation error and lifting overall precision across standard automotive benchmarks. Third, RADIANT surpassed CenterFusion—the previous leading radar-camera fusion method—yielding gains exceeding 18% in precision for cars and 20% for pedestrians. Finally, ablation studies showed that intelligent confidence-based weighting via the depth network significantly outperformed simple mathematical averaging of sensor depths.

These findings indicate that automakers and autonomy developers can achieve significant gains in vehicle perception accuracy and safety without adopting costly LiDAR hardware. By effectively leveraging low-cost radar already present on modern vehicles, RADIANT offers an economically viable pathway to enhance collision avoidance and driver assistance systems. Crucially, because RADIANT attaches as a modular add-on without altering or retraining existing camera models, engineering teams can upgrade current visual detection stacks with minimal architectural disruption and reduced development timelines.

Moving forward, development teams should adopt offset-prediction strategies when combining radar with visual data rather than attempting naive point-to-box associations. Technical leadership should consider pilot testing RADIANT-style fusion modules across existing vision stacks to validate real-time computational performance on automotive-grade microprocessors. Additional research is recommended to expand radar fusion beyond depth correction, specifically targeting object velocity, orientation, and physical dimensions.

Confidence in these findings is supported by thorough benchmark testing against top-tier baselines on the standard nuScenes dataset. However, stakeholders should note specific limitations: the evaluated method focuses exclusively on refining depth, leaving object boundaries, orientations, and velocities dependent on the underlying vision model. Additionally, performance remains bounded by the quality of the baseline camera network and the presence of detectable radar returns on target objects.

Cover for RADIANT: Radar-Image Association Network for 3D Object Detection

Abstract

As a direct depth sensor, radar holds promise as a tool to improve monocular 3D object detection, which suffers from depth errors, due in part to the depth-scale ambiguity. On the other hand, leveraging radar depths is hampered by difficulties in precisely associating radar returns with 3D estimates from monocular methods, effectively erasing its benefits. This paper proposes a fusion network that addresses this radar-camera association challenge. We train our network to predict the 3D offsets between radar returns and object centers, enabling radar depths to enhance the accuracy of 3D monocular detection. By using parallel radar and camera backbones, our network fuses information at both the feature level and detection level, while at the same time leveraging a state-of-the-art monocular detection technique without retraining it. Experimental results show significant improvement in mean average precision and translation error on the nuScenes dataset over monocular counterparts. Our source code is available at https://github.com/longyunf/radiant.

Table of Contents

  • Introduction
  • Related Work
  • Background and Definitions
  • RADIANT
  • Radar Branch
  • Radar Heads
  • Depth Fusion Module
  • Experiments
  • Depth Errors for Camera and Radar Heads
  • nuScenes Quantitative Results
  • nuScenes Qualitative Results
  • Ablation Studies
  • Conclusions
  • Acknowledgments
  • References

Knowls

  1. Knowl 1 — Dual-Branch Parallel Architecture of RADIANT

    model/method

    RADIANT (RADar-Image Association NeTwork) is a multimodal 3D object detection architecture that fuses monocular camera images with sparse radar point clouds in image space. The system employs two parallel branches:

    1. Camera Branch: A standard monocular 3D detection network (such as FCOS3D or PGD) that is pre-trained and kept frozen during radar branch training. It outputs 3D bounding box proposals, including projected 2D centers, full depth estimates, bounding box dimensions, and orientations.
    2. Radar Branch: An independent network consisting of a ResNet-18 backbone, a Feature Pyramid Network (FPN) neck producing representations across five resolution levels (P3,P4,P5,P6,P7P_3, P_4, P_5, P_6, P_7), and a shared radar detection head.

    The radar input is structured as a multi-channel image-plane map evaluated at the pixel projections of radar returns. The radar channels contain measured radar depth, bird's-eye view (BEV) coordinates, Doppler radial velocity, and a binary occupancy mask indicating radar-point pixels. Mid-level feature fusion is achieved by extracting image backbone feature maps at three resolution stages (C3,C4,C5C_3, C_4, C_5) and passing them through fusion convolutions to concatenate with corresponding radar backbone feature maps prior to the radar neck.

  2. Knowl 2 — Radar Residual Depth Offset and Positive Sample Definition

    model/method

    Radar returns typically do not correspond directly to the geometric 3D center of an object; rather, they reflect off exterior surfaces, internal parts, or occluded boundaries, and suffer from sensor measurement errors. To correct this discrepancy, RADIANT defines the residual radar depth offset Δz^r\Delta \hat{z}^r as:

    Δz^r=z−zr\Delta \hat{z}^r = z - z^r

    where z∈R+z \in \mathbb{R}^+ denotes the true 3D object center depth and zr∈R+z^r \in \mathbb{R}^+ represents the direct depth measurement of the radar return point.

    Because automotive datasets such as nuScenes lack point-wise object annotations for radar returns, positive radar pixels are assigned during training using a spatial-velocity consistency rule: a radar point projected onto the image is treated as a positive object candidate if and only if:

    1. The Euclidean distance between the radar point in 3D and the ground-truth 3D bounding box is below a specified spatial distance threshold; and
    2. The projection of the ground-truth object's 3D velocity vector along the sensor's radial direction matches the radar point's measured Doppler velocity within a velocity tolerance threshold.
  3. Knowl 3 — Radar Head for Point-Wise 3D Center Offset Prediction

    model/method

    The radar branch processes fused radar-image feature maps across five feature levels (P3,P4,P5,P6,P7P_3, P_4, P_5, P_6, P_7) using a shared detection head. For each positive radar pixel (ujr,vjr)(u^r_j, v^r_j), the radar head predicts:

    • A classification vector predicting the object class Y^jr\hat{Y}^r_j and classification confidence score σ^jr\hat{\sigma}^r_j.
    • An image-plane 2D pixel offset vector (Δu^jr,Δv^jr)(\Delta \hat{u}^r_j, \Delta \hat{v}^r_j) pointing to the projected 3D object center, yielding the predicted projected center (u^jr,v^jr)=(ujr+Δu^jr,vjr+Δv^jr)(\hat{u}^r_j, \hat{v}^r_j) = (u^r_j + \Delta \hat{u}^r_j, v^r_j + \Delta \hat{v}^r_j).
    • A 1D residual depth offset Δz^jr\Delta \hat{z}^r_j, yielding the radar-derived 3D object center depth z^jr=zjr+Δz^jr\hat{z}^r_j = z^r_j + \Delta \hat{z}^r_j.

    Because sparse radar returns cannot convey 3D object shape or bounding box orientation, the radar head predicts only position offsets and class scores, explicitly omitting 3D bounding box dimensions and orientations.

  4. Knowl 4 — Radar-Camera Detection-Level Association Procedure

    algorithm

    Detection-level association pairs camera box proposals with radar detection candidates by matching predicted categories, projected 2D centers, and estimated 3D depths.

    Input: Camera proposals Bc={bic}B^c = \{b^c_i\}, Radar candidates Br={bjr}B^r = \{b^r_j\}, Score thresholds Tc,TrT_c, T_r, Pixel distance threshold TpT_p, Depth difference threshold TdT_d
    Output: Matched radar detections set(i)\text{set}(i) for each camera proposal bicb^c_i
    Filter BcB^c by keeping proposals with camera detection score σ^ic>Tc\hat{\sigma}^c_i > T_c
    Filter BrB^r by keeping radar candidates with radar detection score σ^jr>Tr\hat{\sigma}^r_j > T_r
    for each radar proposal bjr∈Brb^r_j \in B^r do
        Compute predicted image center (u^jr,v^jr)=(ujr+Δu^jr,vjr+Δv^jr)(\hat{u}^r_j, \hat{v}^r_j) = (u^r_j + \Delta \hat{u}^r_j, v^r_j + \Delta \hat{v}^r_j)
        Compute predicted center depth z^jr=zjr+Δz^jr\hat{z}^r_j = z^r_j + \Delta \hat{z}^r_j
    for each camera proposal bic∈Bcb^c_i \in B^c do
        Initialize set(i)←∅\text{set}(i) \leftarrow \emptyset
        for each radar proposal bjr∈Brb^r_j \in B^r do
            if Y^ic==Y^jr\hat{Y}^c_i == \hat{Y}^r_j and ∥(u^ic,v^ic)−(u^jr,v^jr)∥2<Tp\|(\hat{u}^c_i, \hat{v}^c_i) - (\hat{u}^r_j, \hat{v}^r_j)\|_2 < T_p and ∣z^ic−z^jr∣<Td|\hat{z}^c_i - \hat{z}^r_j| < T_d then
                set(i)←set(i)∪{j}\text{set}(i) \leftarrow \text{set}(i) \cup \{j\}
    return {set(i)}\{\text{set}(i)\}

    The matching procedure evaluates pair-wise compatibility with an algorithmic complexity of O(MN)\mathcal{O}(MN), where MM is the number of camera proposals and NN is the number of radar candidates after threshold filtering.

  5. Knowl 5 — Depth Weighting Network (DWN)

    model/method

    The Depth Weighting Network (DWN) is a 4-layer multi-layer perceptron (MLP) trained to determine the relative confidence between a radar-estimated depth and a monocular camera-estimated depth for an associated detection pair (bic,bjr)(b^c_i, b^r_j).

    The DWN receives an input feature vector containing the output feature representations from both heads, the raw sensor depths, the distance between candidate centers, and the consistency between the Doppler velocity and the predicted camera motion. It outputs a scalar confidence score αj∈[0,1]\alpha_j \in [0, 1], where αj→1\alpha_j \to 1 indicates higher confidence in the radar depth and αj→0\alpha_j \to 0 indicates higher confidence in the monocular camera depth.

    The network is trained with binary cross-entropy loss against ground-truth supervisory labels defined by:

    α={1,if ∣z^jr−z∣<∣z^ic−z∣0,otherwise\alpha = \begin{cases} 1, & \text{if } |\hat{z}^r_j - z| < |\hat{z}^c_i - z| \\ 0, & \text{otherwise} \end{cases}

    where zz is the true 3D object center depth, z^jr\hat{z}^r_j is the radar-corrected depth, and z^ic\hat{z}^c_i is the monocular camera depth prediction.

  6. Knowl 6 — Confidence-Weighted Fused Depth Calculation

    equation

    For a given monocular 3D camera proposal bicb^c_i associated with a set of matched radar candidate detections j∈set(i)j \in \text{set}(i), each having an estimated radar depth z^jr\hat{z}^r_j and a DWN confidence weight αj\alpha_j, the final fused depth z^fuse\hat{z}_{\text{fuse}} is computed as:

    z^fuse={∑j∈set(i)αjz^jr∑j∈set(i)αj,if ∃j∈set(i) such that αj>Tαz^ic,if ∀j∈set(i),αj≤Tα\hat{z}_{\text{fuse}} = \begin{cases} \frac{\sum_{j \in \text{set}(i)} \alpha_j \hat{z}^r_j}{\sum_{j \in \text{set}(i)} \alpha_j}, & \text{if } \exists j \in \text{set}(i) \text{ such that } \alpha_j > T_\alpha \\ \hat{z}^c_i, & \text{if } \forall j \in \text{set}(i), \alpha_j \le T_\alpha \end{cases}

    where z^ic\hat{z}^c_i is the original monocular camera depth estimate and TαT_\alpha is a predefined depth weighting confidence threshold.

  7. Knowl 7 — Depth Prediction Error Comparison on nuScenes Validation Subset

    data/table

    The depth estimation accuracy of the proposed radar head, the baseline monocular head (FCOS3D), and unadjusted raw radar depth was evaluated on 900 images from the nuScenes validation set. Mean absolute depth error (in meters) was computed for close (≤10\le 10,m), medium (10−3010-30,m), and far (≥30\ge 30,m) objects having both positive camera and radar annotations:

    Method ≤10\le 10m 10−3010-30m ≥30\ge 30m All
    Monocular Heads 0.563 1.442 6.042 3.415
    Radar Heads 0.413 0.649 1.017 0.791
    Raw Radar Depth 1.056 1.082 1.361 1.204

    The radar head with learned residual depth offsets achieves the lowest error across all distance ranges (0.791,m overall error), substantially outperforming both monocular estimation (which degrades severely at long range to 6.042,m) and raw radar measurements without offset compensation (1.204,m).

  8. Knowl 8 — 3D Detection Performance Comparison on the nuScenes Test Set

    data/table

    Performance of RADIANT compared against monocular baselines (MonoDIS-M, CenterNet, FCOS3D, PGD) and the state-of-the-art radar-camera fusion method (CenterFusion) on the official nuScenes 3D object detection test benchmark:

    Method mATE ↓\downarrow Mean AP ↑\uparrow Car Truck Bus Trailer CV Ped. Motor. Bicycle TC Barrier
    MonoDIS-M 0.738 0.304 0.478 0.220 0.188 0.176 0.074 0.370 0.290 0.245 0.487 0.511
    CenterNet 0.658 0.338 0.536 0.270 0.248 0.251 0.086 0.375 0.291 0.207 0.583 0.533
    FCOS3D 0.690 0.358 0.524 0.270 0.277 0.255 0.117 0.397 0.345 0.298 0.557 0.538
    PGD 0.646 0.360 0.547 0.268 0.253 0.243 0.087 0.422 0.379 0.300 0.584 0.525
    CenterFusion 0.631 0.326 0.509 0.258 0.234 0.235 0.077 0.370 0.314 0.201 0.575 0.484
    FCOS3D + RADIANT 0.622 0.374 0.582 0.301 0.257 0.248 0.145 0.439 0.386 0.302 0.579 0.500
    PGD + RADIANT 0.609 0.380 0.602 0.302 0.267 0.242 0.107 0.444 0.416 0.312 0.604 0.503

    Applying RADIANT reduces mean Average Translation Error (mATE) and improves mean Average Precision (mAP) for both FCOS3D and PGD, while outperforming CenterFusion in both overall mAP (0.380 vs 0.326) and mATE (0.609 vs 0.631).

  9. Knowl 9 — Ablation of Depth Fusion Strategies on nuScenes Validation Set

    data/table

    An ablation study comparing different depth fusion strategies on the nuScenes validation dataset using FCOS3D and PGD base models. The strategies evaluated are:

    1. None: Baseline monocular detector without radar fusion.
    2. Average: Simple unweighted averaging of camera box proposal depth and neighboring radar proposals.
    3. DWN: RADIANT's confidence-weighted fusion using the Depth Weighting Network.
    Monocular Method Fusion Strategy mATE ↓\downarrow mAP ↑\uparrow
    FCOS3D None 0.739 0.326
    FCOS3D Average 0.711 0.342
    FCOS3D DWN 0.653 0.363
    PGD None 0.658 0.368
    PGD Average 0.647 0.371
    PGD DWN 0.617 0.384

    While simple averaging provides modest improvements over pure monocular baselines, the DWN fusion substantially lowers translation error (mATE) and improves mean Average Precision (mAP) across both architectures, demonstrating the benefit of adaptive confidence-weighted depth merging.

  10. Knowl 10 — Scope and Attribute Limitations of RADIANT Depth Fusion

    limitation

    RADIANT specifically updates and refines the 3D center depth (and therefore the 3D translation) of monocular bounding box candidates. Due to the high spatial sparsity of automotive radar point clouds, the network does not estimate or refine other 3D bounding box attributes, including 3D bounding box dimensions (size), heading orientation, or full 3D object velocities, which remain dependent solely on the underlying monocular camera detector.

Coverage note — Omitted fine-grained channel dimensions of internal convolutional layers and specific supplementary details on input vector concatenations to the Depth Weighting Network.

References

  1. 1.Barnes, D.; Gadd, M.; Murcutt, P.; Newman, P.; and Posner, I. 2020. The Oxford Radar RobotCar Dataset: A radar extension to the Oxford RobotCar Dataset. In ICRA.
  2. 2.Brazil, G.; and Liu, X. 2019a. M3D-RPN: Monocular 3D region proposal network for object detection. In ICCV.
  3. 3.Brazil, G.; and Liu, X. 2019b. Pedestrian detection with autoregressive network phases. In CVPR.
  4. 4.Brazil, G.; Pons-Moll, G.; Liu, X.; and Schiele, B. 2020. Kinematic 3D object detection in monocular video. In ECCV.
  5. 5.Caesar, H.; Bankiti, V.; Lang, A.; Vora, S.; Liong, V.; Xu, Q.; Krishnan, A.; Pan, Y.; Baldan, G.; and Beijbom, O. 2020. nuScenes: A multimodal dataset for autonomous driving. In CVPR.
  6. 6.Chabot, F.; Chaouch, M.; Rabarisoa, J.; Teuliere, C.; and Chateau, T. 2017. Deep MANTA: A coarse-to-fine many-task network for joint 2D and 3D vehicle analysis from monocular image. In CVPR.
  7. 7.Ding, M.; Huo, Y.; Yi, H.; Wang, Z.; Shi, J.; Lu, Z.; and Luo, P. 2020. Learning depth-guided convolutions for monocular 3D object detection. In CVPR Workshops.
  8. 8.Dong, X.; Wang, P.; Zhang, P.; and Liu, L. 2020. Probabilistic oriented object detection in automotive radar. In CVPR Workshops.
  9. 9.Geiger, A.; Lenz, P.; and Urtasun, R. 2012. Are we ready for autonomous driving? the KITTI vision benchmark suite. In CVPR.
  10. 10.He, K.; Zhang, X.; Ren, S.; and Sun, J. 2016. Deep residual learning for image recognition. In CVPR.
  11. 11.Kumar, A.; Brazil, G.; Corona, E.; Parchami, A.; and Liu, X. 2022. DEVIANT: Depth EquiVarIAnt NeTwork for Monocular 3D Object Detection. In ECCV.
  12. 12.Kumar, A.; Brazil, G.; and Liu, X. 2021. GrooMeD-NMS: Grouped mathematically differentiable NMS for monocular 3D object detection. In CVPR.
  13. 13.Lee, W.; Jovanov, L.; and Philips, W. 2021. Semantic-guided radar-vision fusion for depth estimation and object detection. In ICCV.
  14. 14.Li, L.; and Xie, Y. 2020. A feature pyramid fusion detection algorithm based on radar and camera sensor. In ICSP.
  15. 15.Lien, J.; Gillian, N.; Karagozler, E.; Amihood, P.; Schwesig, C.; Olson, E.; Raja, H.; and Poupyrev, I. 2016. Soli: Ubiquitous gesture sensing with millimeter wave radar. TOG.
  16. 16.Lin, T.-Y.; Dollar, P.; Girshick, R.; He, K.; Hariharan, B.;´ and Belongie, S. 2017. Feature pyramid networks for object detection. In CVPR.
  17. 17.Liu, F.; and Liu, X. 2021. Voxel-based 3D detection and reconstruction of multiple objects from a single image. In NeurIPS.
  18. 18.Liu, L.; Lu, J.; Xu, C.; Tian, Q.; and Zhou, J. 2019. Deep fitting degree scoring network for monocular 3D object detection. In CVPR.
  19. 19.Liu, Z.; Zhou, D.; Lu, F.; Fang, J.; and Zhang, L. 2021. AutoShape: Real-time shape-aware monocular 3D object detection. In ICCV.
  20. 20.Long, Y.; Morris, D.; Liu, X.; Castro, M.; Chakravarty, P.; and Narayanan, P. 2021a. Full-velocity radar returns by radarcamera fusion. In ICCV.
  21. 21.Long, Y.; Morris, D.; Liu, X.; Castro, M.; Chakravarty, P.; and Narayanan, P. 2021b. Radar-camera pixel depth association for depth completion. In CVPR.
  22. 22.Lu, Y.; Ma, X.; Yang, L.; Zhang, T.; Liu, Y.; Chu, Q.; Yan, J.; and Ouyang, W. 2021. Geometry uncertainty projection network for monocular 3D object detection. In ICCV.
  23. 23.Ma, X.; Liu, S.; Xia, Z.; Zhang, H.; Zeng, X.; and Ouyang, W. 2020. Rethinking Pseudo-LiDAR representation. In ECCV.
  24. 24.Ma, X.; Wang, Z.; Li, H.; Zhang, P.; Ouyang, W.; and Fan, X. 2019. Accurate monocular 3D object detection via colorembedded 3D reconstruction for autonomous driving. In ICCV.
  25. 25.Ma, X.; Zhang, Y.; Xu, D.; Zhou, D.; Yi, S.; Li, H.; and Ouyang, W. 2021. Delving into localization errors for monocular 3D object detection. In CVPR.
  26. 26.Major, B.; Fontijne, D.; Ansari, A.; Teja Sukhavasi, R.; Gowaikar, R.; Hamilton, M.; Lee, S.; Grzechnik, S.; and Subramanian, S. 2019. Vehicle detection with automotive radar using deep learning on range-azimuth-Doppler tensors. In ICCV Workshops.
  27. 27.Nabati, R.; Harris, L.; and Qi, H. 2021. CFTrack: Centerbased radar and camera fusion for 3D multi-object tracking. In Intelligent Vehicles Symposium Workshops.
  28. 28.Nabati, R.; and Qi, H. 2019. RRPN: Radar region proposal network for object detection in autonomous vehicles. In ICIP.
  29. 29.Nabati, R.; and Qi, H. 2021. CenterFusion: Center-based radar and camera fusion for 3D object detection. In WACV.
  30. 30.Ouaknine, A.; Newson, A.; Rebut, J.; Tupin, F.; and Perez, P. 2021. CARRADA dataset: Camera and automotive radar with range-angle-Doppler annotations. In ICPR.
  31. 31.Park, D.; Ambrus, R.; Guizilini, V.; Li, J.; and Gaidon, A. 2021. Is Pseudo-LiDAR needed for monocular 3D object detection? In ICCV.
  32. 32.Rematas, K.; Kemelmacher-Shlizerman, I.; Curless, B.; and Seitz, S. 2018. Soccer on your tabletop. In CVPR.
  33. 33.Saxena, A.; Driemeyer, J.; and Ng, A. 2008. Robotic grasping of novel objects using vision. IJRR.
  34. 34.Shah, M.; Huang, Z.; Laddha, A.; Langford, M.; Barber, B.; Zhang, S.; Vallespi-Gonzalez, C.; and Urtasun, R. 2020. LiRaNet: End-to-end trajectory prediction using spatiotemporal radar fusion. CoRL.
  35. 35.Shi, S.; Wang, X.; and Li, H. 2019. PointRCNN: 3D object proposal generation and detection from point cloud. In CVPR.
  36. 36.Shuai, X.; Shen, Y.; Tang, Y.; Shi, S.; Ji, L.; and Xing, G. 2021. milliEye: A lightweight mmWave radar and camera fusion system for robust object detection. In International Conference on Internet-of-Things Design and Implementation.
  37. 37.Simonelli, A.; Bulo, S.; Porzi, L.; Antequera, M.; and` Kontschieder, P. 2020. Disentangling monocular 3D object detection: From single to multi-class recognition. TPAMI.
  38. 38.Simonelli, A.; Bulo, S.; Porzi, L.; Kontschieder, P.; and Ricci,` E. 2021. Are we missing confidence in Pseudo-LiDAR methods for monocular 3D object detection? In ICCV.
  39. 39.Tang, Y.; Dorn, S.; and Savani, C. 2020. Center3D: Centerbased monocular 3D object detection with joint depth understanding. In GCPR.
  40. 40.Wang, L.; Zhang, L.; Zhu, Y.; Zhang, Z.; He, T.; Li, M.; and Xue, X. 2021a. Progressive coordinate transforms for monocular 3D object detection. In NeurIPS.
  41. 41.Wang, T.; Zhu, X.; Pang, J.; and Lin, D. 2021b. FCOS3D: Fully convolutional one-stage monocular 3D object detection. In ICCV Workshops.
  42. 42.Wang, T.; Zhu, X.; Pang, J.; and Lin, D. 2021c. Probabilistic and geometric depth: Detecting objects in perspective. In CoRL.
  43. 43.Wang, Y.; Chao, W.-L.; Garg, D.; Hariharan, B.; Campbell, M.; and Weinberger, K. 2019. Pseudo-LiDAR from visual depth estimation: Bridging the gap in 3D object detection for autonomous driving. In CVPR.
  44. 44.Wang, Y.; Guizilini, V.; Zhang, T.; Wang, Y.; Zhao, H.; and Solomon, J. 2021d. DETR3D: 3D object detection from multi-view images via 3D-to-2D queries. In CoRL.
  45. 45.Wang, Y.; Jiang, Z.; Li, Y.; Hwang, J.-N.; Xing, G.; and Liu, H. 2021e. RODNet: A real-time radar object detection network cross-supervised by camera-radar fused object 3D localization. IEEE Journal of Selected Topics in Signal Processing.
  46. 46.Wang, Y.; Mao, Q.; Zhu, H.; Zhang, Y.; Ji, J.; and Zhang, Y. 2021f. Multi-modal 3D object detection in autonomous driving: a survey. arXiv preprint arXiv:2106.12735.
  47. 47.Xu, Q.; Zhou, Y.; Wang, W.; Qi, C.; and Anguelov, D. 2021. SPG: Unsupervised domain adaptation for 3D object detection via semantic point generation. In ICCV.
  48. 48.Yadav, R.; Vierling, A.; and Berns, K. 2020. Radar+RGB fusion for robust object detection in autonomous vehicle. In ICIP.
  49. 49.Yang, B.; Guo, R.; Liang, M.; Casas, S.; and Urtasun, R. 2020. RadarNet: Exploiting Radar for robust perception of dynamic objects. In ECCV.
  50. 50.Zhou, X.; Wang, D.; and Krahenb ¨ uhl, P. 2019. Objects as ¨ points. arXiv preprint arXiv:1904.07850.
  51. 51.Zhou, Y.; He, Y.; Zhu, H.; Wang, C.; Li, H.; and Jiang, Q. 2021. MonoEF: Extrinsic parameter free monocular 3D object detection. TPAMI.

Citation

MLA
Long, Y., et al. “RADIANT: Radar-Image Association Network for 3D Object Detection”. Proceedings of the AAAI Conference on Artificial Intelligence, vol. 37, no. 2, 2023, pp. 1808–16, https://doi.org/10.1609/AAAI.V37I2.25270.
APA
Long, Y., Kumar, A., Morris, D., Liu, X., Castro, M., & Chakravarty, P. (2023). RADIANT: Radar-Image Association Network for 3D Object Detection. Proceedings of the AAAI Conference on Artificial Intelligence, 37(2), 1808–1816. https://doi.org/10.1609/AAAI.V37I2.25270
Chicago
Long, Y., A. Kumar, D. Morris, X. Liu, M. Castro, and P. Chakravarty. 2023. “RADIANT: Radar-Image Association Network for 3D Object Detection”. Proceedings of the AAAI Conference on Artificial Intelligence 37 (2): 1808–16. https://doi.org/10.1609/AAAI.V37I2.25270.
Harvard
Long, Y. et al. (2023) “RADIANT: Radar-Image Association Network for 3D Object Detection”, Proceedings of the AAAI Conference on Artificial Intelligence, 37(2), pp. 1808–1816. Available at: https://doi.org/10.1609/AAAI.V37I2.25270.
Vancouver
1. Long Y, Kumar A, Morris D, Liu X, Castro M, Chakravarty P (2023) RADIANT: Radar-Image Association Network for 3D Object Detection. Proceedings of the AAAI Conference on Artificial Intelligence 37:1808–1816

BibTeX

@article{Long_2023, title={RADIANT: Radar-Image Association Network for 3D Object Detection}, volume={37}, ISSN={2159-5399}, url={http://dx.doi.org/10.1609/AAAI.V37I2.25270}, DOI={10.1609/aaai.v37i2.25270}, number={2}, journal={Proceedings of the AAAI Conference on Artificial Intelligence}, publisher={Association for the Advancement of Artificial Intelligence (AAAI)}, author={Long, Yunfei and Kumar, Abhinav and Morris, Daniel and Liu, Xiaoming and Castro, Marcos and Chakravarty, Punarjay}, year={2023}, month=June, pages={1808–1816} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF