BEVHeight: A Robust Framework for Vision-based Roadside 3D Object Detection
Lei YangKaicheng YuTao TangJun LiKun YuanLi WangXinyu ZhangPeng Chen
Proposes a bird's-eye-view framework that predicts per-pixel height to the ground rather than traditional depth, achieving distance-agnostic 3D object detection that maintains high accuracy despite camera extrinsic variations on roadside perception benchmarks.
Autonomous driving systems increasingly rely on roadside infrastructure, such as elevated cameras mounted on poles, to overcome the visual blind spots and limited perception range of vehicle-mounted sensors. However, existing camera-only 3D object detection frameworks perform poorly when adapted to roadside units. These methods typically estimate the distance (depth) of objects relative to the camera center. As objects move farther away from elevated roadside cameras, the depth differences between the ground and objects rapidly vanish. Furthermore, roadside cameras often shift due to wind, vibration, or maintenance, causing depth-based models to fail severely in real-world deployment.
The article demonstrates a novel vision-based framework, named BEVHeight, designed to accurately detect 3D objects from roadside cameras by predicting an object's height above the ground rather than its depth from the camera lens.
To evaluate this approach, the authors developed a specialized height-based projection method that maps 2D image features into a unified 3D bird's-eye-view representation. They tested the framework across two large-scale roadside datasets (DAIR-V2X-I, containing approximately 10,000 images, and Rope3D, containing over 500,000 images) against established monocular and bird's-eye-view detectors. They also simulated real-world camera orientation disturbances (rotational roll and pitch noise) to assess system stability under maintenance and weather-induced shifts.
The findings show that BEVHeight achieves state-of-the-art accuracy, outperforming leading camera-only methods on clean datasets by about 2% to 6% across vehicle, pedestrian, and cyclist categories. Crucially, in simulated noisy environments with perturbed camera angles, traditional depth-based models suffered catastrophic performance collapses (dropping from approximately 61% accuracy down to under 10% on vehicle detection). In contrast, BEVHeight maintained 51.77% accuracy under the same severe disturbances, delivering an absolute performance advantage of over 26% to 42% over previous baselines. Analysis also confirmed that height estimation significantly reduces localization errors at mid-to-long distances.
These results demonstrate that estimating ground height provides a robust, distance-consistent geometric foundation for elevated camera perception. Transitioning to height-based modeling mitigates the risk of roadside sensor misalignment, reducing the need for costly frequent physical recalibrations while enhancing traffic monitoring safety and reliability. However, testing also revealed a boundary condition: the framework underperforms depth-based methods when mounted close to the ground on passenger cars, though it retains its superiority when installed on taller commercial vehicles such as trucks.
Organizations developing intelligent transportation infrastructure or cooperative vehicle-to-infrastructure systems should prioritize height-based geometric projection over standard depth-based pipelines for roadside and elevated sensors. Before broad rollout, practitioners should conduct live pilot testing on physical roadside intersections to validate real-time computational overhead and evaluate the model under diverse weather conditions and varied pole heights.
- Paper: Lift, Splat, Shoot: Encoding Images From Arbitrary Camera Rigs by Implicitly Unprojecting to 3D, Jonah Philion et al. (2020). This paper establishes the foundational lift-splat projection pipeline for lifting 2D features into bird's-eye-view representations, which BEVHeight directly re-architects around ground height rather than depth.
- Paper: BEVFormer: Learning Bird's-Eye-View Representation from Multi-Camera Images via Spatiotemporal Transformers, Zhiqi Li et al. (2022). This work introduces state-of-the-art camera-only bird's-eye-view perception frameworks, providing the architectural and methodological baseline that BEVHeight adapts for roadside infrastructure.
- Paper: Learning Rich Features from RGB-D Images for Object Detection and Segmentation, Saurabh Gupta et al. (2014). This study introduces geocentric feature representations based on height above ground, establishing the core geometric rationale underlying BEVHeight's elevation-based mapping.
No sufficiently relevant recommendations were found.
