Raw High-Definition Radar for Multi-Task Learning
Julien RebutArthur OuaknineWaqas MalikPatrick Pérez
Presents FFT-RadNet, a computationally efficient architecture for simultaneous vehicle detection and free-driving space segmentation directly from raw high-definition radar data, supported by the RADIAl dataset featuring synchronized raw radar, camera, and LiDAR recordings.
Automated driving systems require robust, all-weather perception sensors to ensure passenger safety. While high-definition radar provides critical resilience in adverse conditions and measures vehicle speed, its integration into production vehicles is severely constrained by hardware demands. Traditional radar processing requires computing expensive range, angle, and Doppler maps, which consumes significant onboard memory and processing power.
The article demonstrates an efficient deep-learning architecture, called FFT-RadNet, that performs vehicle detection and free-driving-space segmentation directly from raw range-Doppler radar data without traditional, resource-heavy angle calculations.
The researchers trained FFT-RadNet on RADIal, a newly recorded multimodal dataset consisting of two hours of driving data across city streets, highways, and rural roads, capturing synchronized raw high-definition radar, laser scanner, and camera signals. They benchmarked FFT-RadNet against standard perception architectures that rely on pre-processed radar point clouds and range-azimuth representations.
The findings show that FFT-RadNet performs on par with or better than computationally intensive baselines while reducing compute and memory overhead. For vehicle detection, the model achieves a 96.8% average precision, outperforming point-cloud-based methods and closely matching pre-calculated range-azimuth models. For free-driving-space segmentation, it outperforms existing models with an average Intersection-over-Union score of 74.0%, representing a 13.4 percentage-point improvement. Crucially, the architecture completely eliminates the dedicated angle-of-arrival processing stage, which typically adds 45 to nearly 500 Giga Floating Point Operations per Second on high-definition systems, while using under four million parameters.
These results demonstrate that deep learning models can directly learn angular positions from compressed sensor signals on embedded vehicle hardware. This reduces compute costs and hardware strain, enhances functional safety through reliable radar-based road segmentation, and eliminates expensive end-of-line sensor calibration during manufacturing.
Stakeholders and engineering teams should adopt direct range-Doppler processing pipelines when deploying high-definition automotive radars to reduce onboard compute budgets. Further work should explore multi-sensor fusion algorithms using raw radar streams and expand the dataset to refine automated labeling quality, as pitch variations during vehicle motion can introduce minor projection errors in segmentation annotations.
- Paper: Deep Multi-Modal Object Detection and Semantic Segmentation for Autonomous Driving: Datasets, Methods, and Challenges, Di Feng et al. (2019). Provides a comprehensive survey of multi-modal perception and sensor representations in autonomous driving, establishing the foundational trade-offs between point cloud, image, and raw-signal pipelines.
- Paper: Lift, Splat, Shoot: Encoding Images From Arbitrary Camera Rigs by Implicitly Unprojecting to 3D, Jonah Philion et al. (2020). Introduces foundational concepts for lifting low-level sensor representations into bird's-eye-view feature maps for multi-task driving perception without relying on heavy intermediate stages.
- Paper: Center-based 3D Object Detection and Tracking, Tianwei Yin et al. (2020). Establishes anchor-free, center-based object detection heads that motivate the multi-task detection architectures utilized in raw automotive radar processing.
- Paper: Multi-view 3D Object Detection Network for Autonomous Driving, Xiaozhi Chen et al. (2017). Introduces the paradigm of multi-view proposal generation and bird's-eye-view representations that traditional radar and LiDAR networks adapt for 3D perception.
- Paper: VoxelNet: End-to-End Learning for Point Cloud Based 3D Object Detection, Yin Zhou et al. (2017). Presents end-to-end learning directly from sparse raw spatial sensors, serving as a key benchmark against point-cloud-based perception methods.
- Paper: LinkNet: Exploiting encoder representations for efficient semantic segmentation, Abhishek Chaurasia et al. (2017). Demonstrates lightweight encoder-decoder architectures with direct bypass connections that inform efficient real-time segmentation and multi-task network design on embedded hardware.
- Paper: RADIANT: Radar-Image Association Network for 3D Object Detection, Yunfei Long et al. (2023). Explores how raw or processed radar signals can be directly associated with monocular visual features to overcome depth ambiguity in multi-sensor 3D object detection.
- Paper: BEVFusion: Multi-Task Multi-Sensor Fusion with Unified Bird's-Eye View Representation, Zhijian Liu et al. (2022). Extends multi-task spatial perception by creating a unified bird's-eye-view representation that fuses complementary sensor streams like camera and LiDAR efficiently.
- Paper: BEVFormer: Learning Bird's-Eye-View Representation from Multi-Camera Images via Spatiotemporal Transformers, Zhiqi Li et al. (2022). Advances top-down bird's-eye-view feature learning by incorporating spatiotemporal transformer attention across multi-sensor driving sequences.
- Paper: SparseOcc: Rethinking Sparse Latent Representation for Vision-Based Semantic Occupancy Prediction, Pin Tang et al. (2024). Extends efficient 3D scene understanding and drivable space estimation to complete 3D semantic occupancy prediction using sparse latent representations.
