Raw High-Definition Radar for Multi-Task Learning

Julien RebutArthur OuaknineWaqas MalikPatrick Pérez

article2022CVPR132 citations

Presents FFT-RadNet, a computationally efficient architecture for simultaneous vehicle detection and free-driving space segmentation directly from raw high-definition radar data, supported by the RADIAl dataset featuring synchronized raw radar, camera, and LiDAR recordings.

Listen

Automated driving systems require robust, all-weather perception sensors to ensure passenger safety. While high-definition radar provides critical resilience in adverse conditions and measures vehicle speed, its integration into production vehicles is severely constrained by hardware demands. Traditional radar processing requires computing expensive range, angle, and Doppler maps, which consumes significant onboard memory and processing power.

The article demonstrates an efficient deep-learning architecture, called FFT-RadNet, that performs vehicle detection and free-driving-space segmentation directly from raw range-Doppler radar data without traditional, resource-heavy angle calculations.

The researchers trained FFT-RadNet on RADIal, a newly recorded multimodal dataset consisting of two hours of driving data across city streets, highways, and rural roads, capturing synchronized raw high-definition radar, laser scanner, and camera signals. They benchmarked FFT-RadNet against standard perception architectures that rely on pre-processed radar point clouds and range-azimuth representations.

The findings show that FFT-RadNet performs on par with or better than computationally intensive baselines while reducing compute and memory overhead. For vehicle detection, the model achieves a 96.8% average precision, outperforming point-cloud-based methods and closely matching pre-calculated range-azimuth models. For free-driving-space segmentation, it outperforms existing models with an average Intersection-over-Union score of 74.0%, representing a 13.4 percentage-point improvement. Crucially, the architecture completely eliminates the dedicated angle-of-arrival processing stage, which typically adds 45 to nearly 500 Giga Floating Point Operations per Second on high-definition systems, while using under four million parameters.

These results demonstrate that deep learning models can directly learn angular positions from compressed sensor signals on embedded vehicle hardware. This reduces compute costs and hardware strain, enhances functional safety through reliable radar-based road segmentation, and eliminates expensive end-of-line sensor calibration during manufacturing.

Stakeholders and engineering teams should adopt direct range-Doppler processing pipelines when deploying high-definition automotive radars to reduce onboard compute budgets. Further work should explore multi-sensor fusion algorithms using raw radar streams and expand the dataset to refine automated labeling quality, as pitch variations during vehicle motion can introduce minor projection errors in segmentation annotations.

Cover for Raw High-Definition Radar for Multi-Task Learning

Abstract

With their robustness to adverse weather conditions and ability to measure speeds, radar sensors have been part of the automotive landscape for more than two decades. Recent progress toward High Definition (HD) Imaging radar has driven the angular resolution below the degree, thus approaching laser scanning performance. However, the amount of data a HD radar delivers and the computational cost to estimate the angular positions remain a challenge. In this paper, we propose a novel HD radar sensing model, FFT-RadNet, that eliminates the overhead of computing the range-azimuth-Doppler 3D tensor, learning instead to recover angles from a range-Doppler spectrum. FFT-RadNet is trained both to detect vehicles and to segment free driving space. On both tasks, it competes with the most recent radar-based models while requiring less compute and memory. Also, we collected and annotated 2-hour worth of raw data from synchronized automotive-grade sensors (camera, laser, HD radar) in various environments (city street, highway, countryside road). This unique dataset, nick-named RAdial for “Radar, LiDAR et al.”, is available at https://github.com/valeoai/RAdial.

Table of Contents

  • 1. Introduction
  • 2. Radar background
  • 3. Related work
  • 4. FFT-RadNet architecture
  • 4.1. MIMO pre-encoder
  • 4.2. FPN encoder
  • 4.3. Range-angle decoder
  • 4.4. Multi-task learning
  • 5. RADIal dataset
  • 6. Experiments
  • 6.1. Training details
  • 6.2. Baselines
  • 6.3. Evaluation metric
  • 6.4. Performance analysis
  • 6.5. Complexity analysis
  • 7. Conclusion
  • References

Knowls

  1. Knowl 1 — FFT-RadNet Architecture for High-Definition Radar Multi-Task Perception

    model/method

    FFT-RadNet is an end-to-end multi-task deep neural network that performs vehicle detection and free-driving space segmentation directly from raw high-definition (HD) Frequency-Modulated Continuous-Wave (FMCW) radar Range-Doppler (RD) spectra, bypassing the computationally expensive Angle-of-Arrival (AoA) estimation and 3D Range-Azimuth-Doppler (RAD) tensor computation.

    The architecture consists of five interconnected modules:

    1. MIMO Pre-Encoder: Ingests the complex Range-Doppler spectrum of dimension (BR,BD,2NRx)(B_R, B_D, 2N_{\text{Rx}}) (where BRB_R is the number of range bins, BDB_D is the number of Doppler bins, and real/imaginary parts of NRxN_{\text{Rx}} receiver channels are concatenated). It applies a 1D atrous convolution along the Doppler axis matched to the transmitter shift followed by a 3×33\times 3 standard convolution to de-interleave and compress multi-input multi-output (MIMO) antenna signatures.
    2. Feature Pyramid Network (FPN) Encoder: A backbone composed of 4 residual blocks containing 3, 6, 6, and 3 residual layers respectively. Each block performs 2×22\times 2 spatial downsampling, resulting in an overall spatial reduction factor of 16 in height and width with 3×33\times 3 convolutional kernels to prevent signal overlaps between transmitters while extracting multi-scale features.
    3. Range-Angle Decoder: Converts multi-scale range-Doppler feature maps into a latent Range-Azimuth (RA) representation. It applies a 1×11\times 1 convolution to match azimuth channel width, swaps the Doppler and azimuth dimensions, upscales along the range dimension via deconvolution layers concatenated with corresponding encoder skip connections, and processes the output through two Conv-BatchNorm-ReLU layers.
    4. Detection Head: A single-stage detection branch inspired by PIXOR. It passes the range-azimuth latent representation through four Conv-BatchNorm layers (with 144, 96, 96, and 96 filters) and splits into: (i) a classification sub-branch outputting a probability map of size (BR/4,BA/8)(B_R/4, B_A/8) with a coarse resolution of 0.8 m0.8\,\text{m} in range and 0.8∘0.8^\circ in azimuth; and (ii) a regression sub-branch utilizing a 3×33\times 3 convolution to predict continuous range and azimuth offset corrections for positive detections.
    5. Free-Space Segmentation Head: Predicts a binary drivable space probability grid of resolution 0.4 m0.4\,\text{m} in range and 0.2∘0.2^\circ in azimuth over a field of view [−45∘,45∘][-45^\circ, 45^\circ] and range [0,50 m][0, 50\,\text{m}]. It processes the range-azimuth latent features through two successive blocks of two Conv-BatchNorm-ReLU layers (128 and 64 filters) followed by a 1×11\times 1 convolution and a sigmoid activation.
  2. Knowl 2 — MIMO Pre-Encoder for Interleaved Range-Doppler Radar Spectra

    model/method

    In Time-Division Multiplexed or Phase-Shifted Multiple-Input Multiple-Output (MIMO) automotive FMCW radars with NTxN_{\text{Tx}} transmitters and NRxN_{\text{Rx}} receivers, a transmitting phase shift Δϕ\Delta \phi is introduced between consecutive antennas to avoid signal interference. Consequently, an object at radial distance RR with relative radial velocity DD produces NTxN_{\text{Tx}} distinct interleaved signatures in the Range-Doppler (RD) spectrum of each receiver antenna, located at:

    (R,(D+kΔ) mod Dmax⁡)k=1,…,NTx(R, (D + k\Delta) \bmod D_{\max})_{k=1,\dots,N_{\text{Tx}}}

    where Δ\Delta is the Doppler shift induced by Δϕ\Delta \phi, and Dmax⁡D_{\max} is the maximum unambiguous measurable Doppler velocity.

    The trainable MIMO pre-encoder resolves this interleaving directly without calculating classical Fourier beamforming along the channel axis:

    1. The real and imaginary components of the complex RD spectrum across all NRxN_{\text{Rx}} receivers are stacked along the channel dimension, yielding an input tensor of size BR×BD×2NRxB_R \times B_D \times 2N_{\text{Rx}}, where BRB_R is the number of range discretization bins and BDB_D is the number of Doppler bins.
    2. An Atrous (dilated) convolution layer is applied along the Doppler axis with a 1×NTx1 \times N_{\text{Tx}} kernel, NRxN_{\text{Rx}} input channels, and a dilation rate of δ=ΔBDDmax⁡\delta = \frac{\Delta B_D}{D_{\max}}, which aligns with the exact Doppler bin spacing between the repeated signatures.
    3. A second 3×33 \times 3 convolutional layer combines and compresses these aligned channels into a compact latent representation prior to feature pyramid encoding.
  3. Knowl 3 — FFT-RadNet Multi-Task Loss Objective

    equation

    FFT-RadNet is trained end-to-end to jointly minimize vehicle detection loss Ldet\mathcal{L}_{\text{det}} and free-space segmentation loss Lfree\mathcal{L}_{\text{free}} across training examples x\mathbf{x}:

    LMTL=∑xLdet(x,yclas,yreg)+λLfree(x,yseg)\mathcal{L}_{\text{MTL}} = \sum_{\mathbf{x}} \mathcal{L}_{\text{det}}(\mathbf{x}, \mathbf{y}_{\text{clas}}, \mathbf{y}_{\text{reg}}) + \lambda \mathcal{L}_{\text{free}}(\mathbf{x}, \mathbf{y}_{\text{seg}})

    where λ>0\lambda > 0 is a multi-task balancing weight (set empirically to λ=100\lambda = 100).

    The detection loss Ldet\mathcal{L}_{\text{det}} is defined as:

    Ldet(x,yclas,yreg)=focal(yclas,y^clas)+β smooth-L1(yreg−y^reg)\mathcal{L}_{\text{det}}(\mathbf{x}, \mathbf{y}_{\text{clas}}, \mathbf{y}_{\text{reg}}) = \text{focal}(\mathbf{y}_{\text{clas}}, \hat{\mathbf{y}}_{\text{clas}}) + \beta\,\text{smooth-L1}(\mathbf{y}_{\text{reg}} - \hat{\mathbf{y}}_{\text{reg}})

    where yclas∈{0,1}BR/4×BA/8\mathbf{y}_{\text{clas}} \in \{0, 1\}^{B_R/4 \times B_A/8} and y^clas∈[0,1]BR/4×BA/8\hat{\mathbf{y}}_{\text{clas}} \in [0, 1]^{B_R/4 \times B_A/8} denote the ground-truth and predicted vehicle occupancy maps at coarse range-azimuth resolution (BRB_R range bins, BAB_A azimuth bins); yreg,y^reg∈R2×BR/4×BA/8\mathbf{y}_{\text{reg}}, \hat{\mathbf{y}}_{\text{reg}} \in \mathbb{R}^{2 \times B_R/4 \times B_A/8} are ground-truth and predicted continuous (r,a)(r, a) offset coordinates evaluated only on positive classification locations; β\beta is a regression weighting hyperparameter (set to β=100\beta = 100); and the focal loss uses focusing parameter γ=2\gamma = 2.

    The free-driving-space segmentation loss Lfree\mathcal{L}_{\text{free}} is the binary cross-entropy (BCE) over the high-resolution drivable grid Ω=[ ⁣[1,BR/2] ⁣]×[ ⁣[1,BA/4] ⁣]\Omega = [\![1, B_R/2]\!] \times [\![1, B_A/4]\!]:

    Lfree(x,yseg)=∑(r,a)∈ΩBCE(yseg(r,a),y^seg(r,a))\mathcal{L}_{\text{free}}(\mathbf{x}, \mathbf{y}_{\text{seg}}) = \sum_{(r, a) \in \Omega} \text{BCE}\left(\mathbf{y}_{\text{seg}}(r,a), \hat{\mathbf{y}}_{\text{seg}}(r,a)\right)

    where yseg∈{0,1}BR/2×BA/4\mathbf{y}_{\text{seg}} \in \{0, 1\}^{B_R/2 \times B_A/4} is the binary drivable space ground truth and y^seg∈[0,1]BR/2×BA/4\hat{\mathbf{y}}_{\text{seg}} \in [0, 1]^{B_R/2 \times B_A/4} is the predicted probability.

  4. Knowl 4 — Range-Angle Decoder Latent Feature Conversion

    model/method

    The Range-Angle Decoder converts the output of the FPN encoder—which operates on range, Doppler, and multi-channel features—into a Range-Azimuth (RA) latent spatial grid suitable for bird's-eye-view vehicle localization and drivable space segmentation.

    The conversion procedure is as follows:

    1. A 1×11 \times 1 convolution is applied to the feature maps from each FPN level to project the channel dimension directly to the target azimuth discretization dimension.
    2. The Doppler axis and the azimuth channel axis are swapped so that the feature map dimensions correspond directly to Range ×\times Azimuth.
    3. Because the range axis has been decimated by a factor of 16 across the 4 FPN blocks while the azimuth dimension is preserved, deconvolution layers are applied solely along the range axis to incrementally upscale the range resolution.
    4. Feature maps at each stage are concatenated with the corresponding higher-resolution skip connections from earlier pyramid levels.
    5. The combined representation passes through a final block of two Conv-BatchNorm-ReLU layers to produce the unified Range-Azimuth latent tensor.
  5. Knowl 5 — RADIal Multimodal High-Definition Radar Dataset

    experimental setup

    RADIal is an open-source multimodal automotive driving dataset providing raw high-definition radar data along with synchronized camera, LiDAR, GPS, and vehicle CAN bus odometry traces collected across city streets, highways, and countryside roads.

    Key specifications and annotations include:

    • Sensor Suite: 16-receiver, 12-transmitter HD FMCW radar delivering raw Analog-to-Digital Converter (ADC) data (allowing extraction of raw ADC, Range-Doppler spectrum, Range-Azimuth map, Range-Azimuth-Doppler tensor, and 3D point cloud), a monocular RGB camera, a 16-beam automotive laser scanner (LiDAR), and vehicle CAN odometry/GPS.
    • Dataset Scale: 91 sequences (1 to 4 minutes each), totaling 2 hours and roughly 25,000 synchronized frames. The standard split allocates approximately 70% to training, 15% to validation, and 15% to testing by sequence.
    • Vehicle Annotations: 8,252 frames contain 9,550 annotated vehicles. Each vehicle instance is labeled with a 2D bounding box in the image plane, metric radial distance to the radar, and relative radial velocity (Doppler value). Initial labels are extracted via RetinaNet on camera images, filtered by consensus between radar and LiDAR point clouds, and manually audited.
    • Free-Space Annotations: Drivable space binary masks are generated via DeepLabV3+ (pre-trained on Cityscapes and fine-tuned on manually annotated driving frames), projected into radar Cartesian coordinates ([−45∘,45∘][-45^\circ, 45^\circ] azimuth FoV, up to 50 m50\,\text{m} range), with dynamic vehicle bounding boxes subtracted.
  6. Knowl 6 — Object Detection Performance on RADIal Dataset

    data/table

    Vehicle detection performance evaluated on the RADIal test set compares FFT-RadNet (taking raw Range-Doppler spectra as input) against PIXOR baselines trained on 3D voxelized radar point clouds (Pixor-PC) and pre-processed 2D Range-Azimuth maps (Pixor-RA). Metrics include Average Precision (AP, in %) and Average Recall (AR, in %) at an Intersection-over-Union threshold of 50%, as well as mean Range error (RR, in meters) and Azimuth angle error (AA, in degrees).

    Model Radar Overall Easy Hard
    Input AP(%) ↑\uparrow AR(%) ↑\uparrow R(m) ↓\downarrow A(∘^\circ) ↓\downarrow AP(%) ↑\uparrow AR(%) ↑\uparrow R(m) ↓\downarrow A(∘^\circ) ↓\downarrow AP(%) ↑\uparrow AR(%) ↑\uparrow R(m) ↓\downarrow A(∘^\circ) ↓\downarrow
    Pixor PC 96.46 32.32 0.17 0.25 99.02 28.83 0.15 0.19 93.28 38.69 0.19 0.33
    Pixor RA 96.56 81.68 0.10 0.20 96.86 88.02 0.09 0.16 95.88 70.10 0.12 0.27
    FFT-RadNet RD 96.84 82.18 0.11 0.17 98.49 91.69 0.10 0.13 92.93 64.82 0.13 0.26

    FFT-RadNet achieves 96.84% overall AP and 82.18% AR, matching or exceeding Pixor-RA (96.56% AP, 81.68% AR) and vastly outperforming the point-cloud baseline Pixor-PC (32.32% AR) which suffers from point sparsity at longer ranges. FFT-RadNet also achieves a lower angular error (0.17∘0.17^\circ vs. 0.20∘0.20^\circ for Pixor-RA), demonstrating that angular recovery can be learned directly from RD spectra without classical digital beamforming.

  7. Knowl 7 — Free-Driving-Space Segmentation Performance on RADIal Dataset

    data/table

    Free-driving-space segmentation performance evaluated on the RADIal test set compares FFT-RadNet (using RD spectra as input) against PolarNet (using pre-processed Range-Azimuth maps as input). The evaluation metric is mean Intersection-over-Union (mIoU, in %) on binary road classification (free vs. occupied) within a [0,50 m][0, 50\,\text{m}] radial range.

    Model Radar input mIoU (%) ↑\uparrow
    Overall Easy Hard
    PolarNet RA 60.6 61.9 57.4
    FFT-RadNet RD 74.0 74.6 72.3

    FFT-RadNet achieves 74.0% overall mIoU, outperforming PolarNet by +13.4 percentage points. This improvement holds across both easy scenes (74.6% vs. 61.9%) and hard scenes (72.3% vs. 57.4%), which is attributed to FFT-RadNet retaining phase and elevation information present in the complex RD tensor that is discarded when compressing raw radar signals into 2D RA maps.

  8. Knowl 8 — Computational and Memory Footprint of Radar Signal Pipelines

    data/table

    A comparison of data volume, parameter count, and compute requirements between conventional signal-processing pipelines (Point Cloud Pixor and Range-Azimuth Pixor) and the raw Range-Doppler pipeline (FFT-RadNet).

    Method Input size (MB) ↓\downarrow # Params (10610^6) ↓\downarrow Complexity (GFLOPS) ↓\downarrow
    AoA proc. Model
    PCL Pixor 98.30 6.93 8 741
    RA Pixor 1.75 6.92 45* 761
    FFT-RadNet 16.00 3.79 0 584
    • AoA Pre-processing Overhead: Traditional Range-Azimuth generation requires correlation with a calibration matrix costing O(NTxNRxBABE)\mathcal{O}(N_{\text{Tx}} N_{\text{Rx}} B_A B_E) operations per RD point. For 1 elevation bin, AoA pre-processing takes 45 GFLOPS; for all BE=11B_E = 11 elevations, it requires up to 496 GFLOPS before neural inference begins. FFT-RadNet incurs 0 GFLOPS of AoA pre-processing.
    • Model Compute and Size: FFT-RadNet uses 3.79M parameters and 584 GFLOPS for combined detection and segmentation, compared to 6.92M parameters and 761 GFLOPS for RA Pixor (detection only) and 6.93M parameters and 741 GFLOPS for PCL Pixor.
  9. Knowl 9 — FFT-RadNet Performance Trade-offs in Disturbed Radar Environments

    limitation

    While FFT-RadNet achieves superior recall and accuracy on clean and easy radar samples, its detection performance degrades on 'Hard' test cases (samples exhibiting severe multi-path reflections on metallic surfaces, prominent antenna side-lobes, or external radar-to-radar interference) compared to models trained on traditionally processed Range-Azimuth (RA) maps. On the RADIal Hard test set, Pixor-RA achieves 70.10% Average Recall (AR) versus 64.82% AR for FFT-RadNet. This discrepancy arises because classical signal-processing chains incorporate deterministic filtering and thresholding heuristics that explicitly mitigate specific physical interference patterns before feature extraction.

Coverage note — None was omitted; all primary architectural components, signal formulations, dataset details, empirical results, and stated limitations are included.

References

  1. 1.Dan Barnes, Matthew Gadd, Paul Murcutt, Paul Newman, and Ingmar Posner. The Oxford Radar RobotCar dataset: a radar extension to the Oxford RobotCar dataset. In ICRA, 2020.
  2. 2.Daniel Brodeski, Igal Bilik, and Raja Giryes. Deep radar detector. In RadarConf, 2019.
  3. 3.Graham M Brooker. Understanding millimetre wave fmcw radars. In ICST, 2005.
  4. 4.Holger Caesar, Varun Bankiti, Alex H. Lang, Sourabh Vora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Giancarlo Baldan, and Oscar Beijbom. nuScenes: a multimodal dataset for autonomous driving. In CVPR, 2020.
  5. 5.Liang-Chieh Chen, Yukun Zhu, George Papandreou, Florian Schroff, and Hartwig Adam. Encoder-decoder with atrous separable convolution for semantic image segmentation. In ECCV, 2018.
  6. 6.Marius Cordts, Mohamed Omran, Sebastian Ramos, Timo Rehfeld, Markus Enzweiler, Rodrigo Benenson, Uwe Franke, Stefan Roth, and Bernt Schiele. The cityscapes dataset for semantic urban scene understanding. In CVPR, 2016.
  7. 7.Andreas Danzer, Thomas Griebel, Martin Bach, and Klaus Dietmayer. 2D car detection in radar data with PointNets. In ITSC, 2019.
  8. 8.Xu Dong, Pengluo Wang, Pengyue Zhang, and Langechuan Liu. Probabilistic oriented object detection in automotive radar. In CVPR Worshops, 2020.
  9. 9.B J Donnet and I D Longstaff. Mimo radar, techniques and opportunities. In EuRAD, 2006.
  10. 10.S. Franceschini, M. Ambrosanio, S. Vitale, F. Baselice, A. Gifuni, G. Grassini, and V. Pascazio. Hand gesture recognition via radar sensors and convolutional neural networks. In RadarConf, 2020.
  11. 11.Xiangyu Gao, Guanbin Xing, Sumit Roy, and Hui Liu. RAMP-CNN: a novel neural network for enhanced automotive radar object recognition. In Sensors, 2020.
  12. 12.Andreas Geiger, Philip Lenz, Christoph Stiller, and Raquel Urtasun. Vision meets robotics: The kitti dataset. IJRR, 2013.
  13. 13.Antoine Ghaleb. Micro-Doppler analysis of non-stationary moving targets in radar imaging. PhD thesis, Telecom Paris, 2009.
  14. 14.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In CVPR, 2015.
  15. 15.Xinrui Jiang, Ye Zhang, Qi Yang, Bin Deng, and Hongqiang Wang. Millimeter-wave array radar-based human gait recognition using multi-channel three-dimensional convolutional neural network. Sensors, 2020.
  16. 16.Prannay Kaul, Daniele De Martini, Matthew Gadd, and Paul Newman. RSS-Net: weakly-supervised multi-class semantic segmentation with FMCW radar. In IV, 2020.
  17. 17.Giseop Kim, Yeong Sang Park, Younghun Cho, Jinyong Jeong, and Ayoung Kim. Mulran: Multimodal range dataset for urban place recognition. In ICRA, 2020.
  18. 18.Diederik P. Kingma and Jimmy Ba. Adam: A method for stochastic optimization. In ICLR, 2015.
  19. 19.Alex H. Lang, Sourabh Vora, Holger Caesar, Lubing Zhou, Jiong Yang, and Oscar Beijbom. Pointpillars: Fast encoders for object detection from point clouds. In CVPR, 2019.
  20. 20.Tsung-Yi Lin, Piotr Dollár, Ross B. Girshick, Kaiming He, Bharath Hariharan, and Serge J. Belongie. Feature pyramid networks for object detection. In CVPR, 2017.
  21. 21.Tsung-Yi Lin, Priya Goyal, Ross B. Girshick, Kaiming He, and Piotr Dollár. Focal loss for dense object detection. In ICCV, 2017.
  22. 22.Jakob Lombacher, Kilian Laudt, Markus Hahn, Jurgen Dickmann, and Christian Wohler. Semantic radar grids. In IV, 2017.
  23. 23.Bence Major, Daniel Fontijne, Amin Ansari, Ravi Teja Sukhavasi, Radhika Gowaikar, Michael Hamilton, Sean Lee, Slawomir Grzechnik, and Sundar Subramanian. Vehicle detection with automotive radar using deep learning on range-azimuth-doppler tensors. In ICCV Workshops, 2019.
  24. 24.Michael Meyer and Georg Kuschk. Automotive radar dataset for deep learning based 3d object detection. In EuRAD, 2019.
  25. 25.Michael Meyer and Georg Kuschk. Deep learning based 3d object detection for automotive radar and camera. In EuRAD, 2019.
  26. 26.Casian Miron, Alexandru Pasarica, and Radu Timofte. Efficient cnn architecture for multi-modal aerial view object classification. In CVPR Workshops, 2021.
  27. 27.Mohammadreza Mostajabi, Ching Ming Wang, Darsh Ranjan, and Gilbert Hsyu. High resolution radar dataset for semi-supervised learning of dynamic objects. In CVPR Workshop, 2020.
  28. 28.Weichong Ng, Guohua Wang, Siddhartha, Zhiping Lin, and Bhaskar Jyoti Dutta. Range-Doppler detection in automotive radar with deep learning. In IJCNN, 2020.
  29. 29.Farzan Erlik Nowruzi, Dhanvin Kolhatkar, Prince Kapoor, Elnaz Jahani Heravi, Fahed Al Hassanat, Robert Laganiere, Julien Rebut, and Waqas Malik. Polarnet: Accelerated deep open space segmentation using automotive radar in polar domain. In VEHITS, 2021.
  30. 30.Arthur Ouaknine, Alasdair Newson, Patrick Pérez, Florence Tupin, and Julien Rebut. Multi-view radar semantic segmentation. In ICCV, 2021.
  31. 31.Arthur Ouaknine, Alasdair Newson, Julien Rebut, Florence Tupin, and Patrick Pérez. Carrada dataset: Camera and automotive radar with range- angle- doppler annotations. In ICPR, 2021.
  32. 32.Andras Palffy, Jiaao Dong, Julian F. P. Kooij, and Dariu M. Gavrila. CNN based road user detection using the 3D radar cube. In RAL, 2020.
  33. 33.Robert Prophet, Anastasios Deligiannis, Juan-Carlos Fuentes-Michel, Ingo Weber, and Martin Vossiek. Semantic segmentation on 3D occupancy grids for automotive radar. In IEEE Access, 2020.
  34. 34.Robert Prophet, Gang Li, Christian Sturm, and Martin Vossiek. Semantic segmentation on automotive radar maps. In IV, 2019.
  35. 35.Nicolas Scheiner, Florian Kraus, Fangyin Wei, Buu Phan, Fahim Mannan, Nils Appenrodt, Werner Ritter, Jürgen Dickmann, Klaus Dietmayer, Bernhard Sick, and Felix Heide. Seeing around street corners: non-line-of-sight detection and tracking in-the-wild using Doppler radar. In CVPR, 2020.
  36. 36.Ole Schumann, Markus Hahn, Nicolas Scheiner, Fabio Weishaupt, Julius F. Tilly, Jürgen Dickmann, and Christian Wöhler. Radarscenes: A real-world radar point cloud data set for automotive applications. In ArXiv, 2021.
  37. 37.Ole Schumann, Jakob Lombacher, Markus Hahn, Christian Wohler, and Jurgen Dickmann. Scene understanding with automotive radar. In IV, 2020.
  38. 38.Marcel Sheeny, Emanuele De Pellegrin, Saptarshi Mukherjee, Alireza Ahrabian, Sen Wang, and Andrew Wallace. RADIATE: a radar dataset for automotive perception. In ArXiv, 2020.
  39. 39.Liat Sless, Gilad Cohen, Bat El Shlomo, and Shaul Oron. Road scene understanding by occupancy grid learning from sparse radar clusters using semantic segmentation. In ICCV Workshops, 2019.
  40. 40.Yizhou Wang, Zhongyu Jiang, Yudong Li, Jenq-Neng Hwang, Guanbin Xing, and Hui Liu. RODNet: a real-time radar object detection network cross-supervised by camera-radar fused object 3D localization. In SP, 2021.
  41. 41.Yizhou Wang, Gaoang Wang, Hung-Min Hsu, Hui Liu, and Jenq-Neng Hwang. Rethinking of radar’s role: A camera-radar dataset and systematic annotator via coordinate alignment. In CVPR Workshop, 2021.
  42. 42.Bin Yang, Wenjie Luo, and Raquel Urtasun. PIXOR: real-time 3d object detection from point clouds. In CVPR, 2018.
  43. 43.Ao Zhang, Farzan Erlik Nowruzi, and Robert Laganière. Raddet: Range-azimuth-doppler based radar object detection for dynamic road users. In CVR, 2021.
  44. 44.Guoqiang Zhang, Haopeng Li, and Fabian Wenger. Object detection and 3d estimation via an FMCW radar using a fully convolutional network. In ICASSP, 2020.
  45. 45.Hengshuang Zhao, Jianping Shi, Xiaojuan Qi, Xiaogang Wang, and Jiaya Jia. Pyramid scene parsing network. In CVPR, 2017.

Citation

MLA
Rebut, J., et al. “Raw High-Definition Radar for Multi-Task Learning”. CVPR2022, 2021, http://arxiv.org/abs/2112.10646v3.
APA
Rebut, J., Ouaknine, A., Malik, W., & Pérez, P. (2021). Raw High-Definition Radar for Multi-Task Learning. CVPR2022. http://arxiv.org/abs/2112.10646v3
Chicago
Rebut, J., A. Ouaknine, W. Malik, and P. Pérez. 2021. “Raw High-Definition Radar for Multi-Task Learning”. CVPR2022. http://arxiv.org/abs/2112.10646v3.
Harvard
Rebut, J. et al. (2021) “Raw High-Definition Radar for Multi-Task Learning”, CVPR2022 [Preprint]. Available at: http://arxiv.org/abs/2112.10646v3.
Vancouver
1. Rebut J, Ouaknine A, Malik W, Pérez P (2021) Raw High-Definition Radar for Multi-Task Learning. CVPR2022

BibTeX

@article{rebut2021raw,
  title = {Raw High-Definition Radar for Multi-Task Learning},
  author = {Rebut, Julien and Ouaknine, Arthur and Malik, Waqas and Pérez, Patrick},
  year = {2021},
  journal = {CVPR2022},
  url = {http://arxiv.org/abs/2112.10646v3},
  eprint = {2112.10646}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE