BEVHeight: A Robust Framework for Vision-based Roadside 3D Object Detection

Lei YangKaicheng YuTao TangJun LiKun YuanLi WangXinyu ZhangPeng Chen

article2023CVPR140 citations

Proposes a bird's-eye-view framework that predicts per-pixel height to the ground rather than traditional depth, achieving distance-agnostic 3D object detection that maintains high accuracy despite camera extrinsic variations on roadside perception benchmarks.

Listen

Autonomous driving systems increasingly rely on roadside infrastructure, such as elevated cameras mounted on poles, to overcome the visual blind spots and limited perception range of vehicle-mounted sensors. However, existing camera-only 3D object detection frameworks perform poorly when adapted to roadside units. These methods typically estimate the distance (depth) of objects relative to the camera center. As objects move farther away from elevated roadside cameras, the depth differences between the ground and objects rapidly vanish. Furthermore, roadside cameras often shift due to wind, vibration, or maintenance, causing depth-based models to fail severely in real-world deployment.

The article demonstrates a novel vision-based framework, named BEVHeight, designed to accurately detect 3D objects from roadside cameras by predicting an object's height above the ground rather than its depth from the camera lens.

To evaluate this approach, the authors developed a specialized height-based projection method that maps 2D image features into a unified 3D bird's-eye-view representation. They tested the framework across two large-scale roadside datasets (DAIR-V2X-I, containing approximately 10,000 images, and Rope3D, containing over 500,000 images) against established monocular and bird's-eye-view detectors. They also simulated real-world camera orientation disturbances (rotational roll and pitch noise) to assess system stability under maintenance and weather-induced shifts.

The findings show that BEVHeight achieves state-of-the-art accuracy, outperforming leading camera-only methods on clean datasets by about 2% to 6% across vehicle, pedestrian, and cyclist categories. Crucially, in simulated noisy environments with perturbed camera angles, traditional depth-based models suffered catastrophic performance collapses (dropping from approximately 61% accuracy down to under 10% on vehicle detection). In contrast, BEVHeight maintained 51.77% accuracy under the same severe disturbances, delivering an absolute performance advantage of over 26% to 42% over previous baselines. Analysis also confirmed that height estimation significantly reduces localization errors at mid-to-long distances.

These results demonstrate that estimating ground height provides a robust, distance-consistent geometric foundation for elevated camera perception. Transitioning to height-based modeling mitigates the risk of roadside sensor misalignment, reducing the need for costly frequent physical recalibrations while enhancing traffic monitoring safety and reliability. However, testing also revealed a boundary condition: the framework underperforms depth-based methods when mounted close to the ground on passenger cars, though it retains its superiority when installed on taller commercial vehicles such as trucks.

Organizations developing intelligent transportation infrastructure or cooperative vehicle-to-infrastructure systems should prioritize height-based geometric projection over standard depth-based pipelines for roadside and elevated sensors. Before broad rollout, practitioners should conduct live pilot testing on physical roadside intersections to validate real-time computational overhead and evaluate the model under diverse weather conditions and varied pole heights.

arXiv: 2303.08498

No sufficiently relevant recommendations were found.

Cover for BEVHeight: A Robust Framework for Vision-based Roadside 3D Object Detection

Abstract

While most recent autonomous driving system focuses on developing perception methods on ego-vehicle sensors, people tend to overlook an alternative approach to leverage intelligent roadside cameras to extend the perception ability beyond the visual range. We discover that the state-of-the-art vision-centric bird’s eye view detection methods have inferior performances on roadside cameras. This is because these methods mainly focus on recovering the depth regarding the camera center, where the depth difference between the car and the ground quickly shrinks while the distance increases. In this paper, we propose a simple yet effective approach, dubbed BEVHeight, to address this issue. In essence, instead of predicting the pixel-wise depth, we regress the height to the ground to achieve a distance-agnostic formulation to ease the optimization process of camera-only perception methods. On popular 3D detection benchmarks of roadside cameras, our method surpasses all previous vision-centric methods by a significant margin. The code is available at https://github.com/ADLab-AutoDrive/BEVHeight.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Method
  • 3.1. Problem Definition
  • 3.2. Comparing the depth and height
  • 3.3. BEVHeight
  • 4. Experiments
  • 4.1. Datasets
  • 4.2. Experimental Settings
  • 4.3. Comparing with state-of-the-art
  • 4.4. Ablation Study
  • 5. Conclusion
  • Acknowledgments
  • References

Knowls

  1. Knowl 1 — BEVHeight Architecture for Vision-Based Roadside 3D Object Detection

    model/method

    BEVHeight is a camera-only 3D object detection framework designed for roadside perception systems. Unlike conventional vision-centric Bird's-Eye-View (BEV) detectors that predict depth from the camera center—which suffers from shrinking foreground-background depth contrast at long distances and severe vulnerability to extrinsic parameter perturbations—BEVHeight predicts the per-pixel height relative to the ground plane.

    The framework consists of five sequential stages:

    1. Image-View Feature Encoder: Given a roadside image I∈R3×H×WI \in \mathbb{R}^{3 \times H \times W}, a 2D backbone with a Feature Pyramid Network (FPN) extracts multi-scale 2D features F2d∈RCF×H16×W16F^{2d} \in \mathbb{R}^{C_F \times \frac{H}{16} \times \frac{W}{16}}, where CFC_F is the feature channel count, and H,WH, W are the image dimensions.
    2. HeightNet: From F2dF^{2d}, Squeeze-and-Excitation layers generate context features Fcontext∈RCc×H16×W16F^{context} \in \mathbb{R}^{C_c \times \frac{H}{16} \times \frac{W}{16}}, and deformable convolution layers predict a categorical height distribution Hpred∈RCH×H16×W16H^{pred} \in \mathbb{R}^{C_H \times \frac{H}{16} \times \frac{W}{16}} across CHC_H discretized ground-height bins.
    3. Context and Height Fusion: Context features and predicted height distributions are combined via an outer product to form fused features Ffused∈RCc×CH×H16×W16F^{fused} \in \mathbb{R}^{C_c \times C_H \times \frac{H}{16} \times \frac{W}{16}}.
    4. Height-Based 2D-to-3D Projection: A geometric transformation module unprojects FfusedF^{fused} into a 3D wedge-shaped volume feature tensor Fwedge∈RX×Y×Z×CcF^{wedge} \in \mathbb{R}^{X \times Y \times Z \times C_c} in the ego coordinate system using a virtual coordinate system and a unit-distance reference plane.
    5. Voxel Pooling and Detection Head: The wedge-shaped volume features FwedgeF^{wedge} are pooled along the vertical (height) dimension into a unified BEV feature map FbevF^{bev}. A convolutional detection head processes FbevF^{bev} to regress 7-DoF bounding box parameters (x,y,z,l,w,h,θ)(x, y, z, l, w, h, \theta) for detected objects, representing center position (x,y,z)(x, y, z), 3D dimensions (l,w,h)(l, w, h), and yaw orientation θ\theta.
  2. Knowl 2 — Height-Based 2D-to-3D Projection Formulation

    equation

    To unproject 2D image coordinates to 3D ego coordinates using predicted height from the ground rather than camera depth, BEVHeight constructs a virtual coordinate system whose origin coincides with the camera center OcamO_{cam} and whose YY-axis is perpendicular to the ground plane. A reference plane planerefplane_{ref} is defined parallel to the image plane at unit depth dref=1d_{ref} = 1.

    For an image pixel pimage=(u,v)p_{image} = (u, v) with camera intrinsic matrix K∈R3×3K \in \mathbb{R}^{3 \times 3}, the unprojected point on the reference plane in the camera coordinate frame is:

    Prefcam=K−1dref[u,v,1]T=K−1[u,v,1]TP_{ref}^{cam} = K^{-1} d_{ref} [u, v, 1]^T = K^{-1} [u, v, 1]^T

    Using the camera-to-virtual rotation matrix Tcamvirt.T_{cam}^{virt.}, the point in the virtual coordinate system is:

    Prefvirt.=[xrefvirt.,yrefvirt.,zrefvirt.]T=Tcamvirt.PrefcamP_{ref}^{virt.} = [x_{ref}^{virt.}, y_{ref}^{virt.}, z_{ref}^{virt.}]^T = T_{cam}^{virt.} P_{ref}^{cam}

    Let HH be the camera height from the ground (the vertical distance from the virtual origin to the ground plane) and hih_i be the height above ground corresponding to the ii-th height bin. By similar triangles along the ground-perpendicular axis, the ii-th 3D point in the virtual coordinate system is:

    Pivirt.=H−hiyrefvirt.Prefvirt.P_i^{virt.} = \frac{H - h_i}{y_{ref}^{virt.}} P_{ref}^{virt.}

    Transforming Pivirt.P_i^{virt.} into the ego coordinate frame via virtual-to-ego transformation matrix Tvirt.egoT_{virt.}^{ego} yields the unified projection equation:

    Piego=Tvirt.egoH−hiyrefvirt.Tcamvirt.K−1[u,v,1]TP_i^{ego} = T_{virt.}^{ego} \frac{H - h_i}{y_{ref}^{virt.}} T_{cam}^{virt.} K^{-1} [u, v, 1]^T

  3. Knowl 3 — Height-Based 2D to 3D Projection Algorithm

    algorithm

    The height-based 2D to 3D projector places fused image-context and height features into a wedge-shaped 3D volume grid by mapping each feature pixel along its height-bin-derived spatial coordinates.

    Input: Fused feature map Ffused={f1fused,…,fMfused}F^{fused} = \{f_1^{fused}, \dots, f_{M}^{fused}\} where M=H16×W16M = \frac{H}{16} \times \frac{W}{16} and each slice fmfused∈RCH×Ccf_m^{fused} \in \mathbb{R}^{C_H \times C_c}, roadside camera ground height HH, intrinsic matrix KK, coordinate transformations Tcamvirt.T_{cam}^{virt.} and Tvirt.egoT_{virt.}^{ego}, and height bin values {hi}i=0CH−1\{h_i\}_{i=0}^{C_H-1}.
    Output: 3D wedge-shaped volume feature tensor Fwedge∈RX×Y×Z×CcF^{wedge} \in \mathbb{R}^{X \times Y \times Z \times C_c}.
    Fwedge←{}F^{wedge} \leftarrow \{\}
    for each spatial feature index mm corresponding to pixel (u,v)(u, v) do
        Prefcam←K−1[u,v,1]TP_{ref}^{cam} \leftarrow K^{-1} [u, v, 1]^T
        Prefvirt.←Tcamvirt.PrefcamP_{ref}^{virt.} \leftarrow T_{cam}^{virt.} P_{ref}^{cam}
        Let Prefvirt.=(xrefvirt.,yrefvirt.,zrefvirt.)P_{ref}^{virt.} = (x_{ref}^{virt.}, y_{ref}^{virt.}, z_{ref}^{virt.})
        for i←0i \leftarrow 0 to CH−1C_H - 1 do
            Pivirt.←H−hiyrefvirt.Prefvirt.P_i^{virt.} \leftarrow \frac{H - h_i}{y_{ref}^{virt.}} P_{ref}^{virt.}
            Piego←Tvirt.egoPivirt.P_i^{ego} \leftarrow T_{virt.}^{ego} P_i^{virt.}
            Fwedge←Fwedge∪associate(Piego,fmfused[i])F^{wedge} \leftarrow F^{wedge} \cup \text{associate}(P_i^{ego}, f_m^{fused}[i])
        end for
    end for
    return FwedgeF^{wedge}

    The operation associate(Piego,fmfused[i])\text{associate}(P_i^{ego}, f_m^{fused}[i]) maps the CcC_c-dimensional context feature associated with the ii-th height bin to the corresponding discrete voxel cell in FwedgeF^{wedge} determined by the ego coordinate position PiegoP_i^{ego}.

  4. Knowl 4 — Dynamic Height Discretization

    equation

    To cast continuous ground-height estimation into a discrete classification problem over height bins, BEVHeight uses a dynamic non-linear discretization strategy:

    hi=⌊N×h−hminhmax−hminα⌋h_i = \left\lfloor N \times \sqrt[\alpha]{\frac{h - h_{min}}{h_{max} - h_{min}}} \right\rfloor

    where hh is the continuous height of a point relative to the ground plane, hminh_{min} and hmaxh_{max} denote the lower and upper boundaries of the modeled height range (e.g., −1 m-1\text{ m} to 2 m2\text{ m}), NN is the total number of categorical height bins (CHC_H), hi∈{0,1,…,N−1}h_i \in \{0, 1, \dots, N-1\} is the resulting height bin index, and α\alpha is an adjustable hyperparameter controlling the concentration of height bins across the elevation spectrum.

  5. Knowl 5 — Context and Height Feature Fusion

    equation

    Within the HeightNet module, the 2D feature map F2d∈RCF×H16×W16F^{2d} \in \mathbb{R}^{C_F \times \frac{H}{16} \times \frac{W}{16}} extracted by the image backbone and FPN is processed into context features Fcontext∈RCc×H16×W16F^{context} \in \mathbb{R}^{C_c \times \frac{H}{16} \times \frac{W}{16}} and categorical height distribution logits Hpred∈RCH×H16×W16H^{pred} \in \mathbb{R}^{C_H \times \frac{H}{16} \times \frac{W}{16}}.

    The fused feature tensor FfusedF^{fused} is generated via the outer product of the context features and the height probability distributions at each spatial pixel location:

    Ffused=Fcontext⊗HpredF^{fused} = F^{context} \otimes H^{pred}

    Ffused∈RCc×CH×H16×W16F^{fused} \in \mathbb{R}^{C_c \times C_H \times \frac{H}{16} \times \frac{W}{16}}

    where CcC_c is the channel dimension of context representations, CHC_H is the number of discrete height bins, and H16×W16\frac{H}{16} \times \frac{W}{16} is the spatial resolution of the downsampled feature map.

  6. Knowl 6 — 3D Object Detection Performance on DAIR-V2X-I Dataset

    data/table

    The DAIR-V2X-I roadside perception validation set is evaluated using 3D Average Precision (AP3DAP_{3D}) under Easy, Mid (Moderate), and Hard difficulty criteria following the KITTI evaluation standard. The IoU thresholds are set to 0.5 for Vehicle and 0.25 for Pedestrian and Cyclist.

    Method Modality Vehicle (IoU=0.5) Pedestrian (IoU=0.25) Cyclist (IoU=0.25)
    Easy Mid Hard Easy Mid Hard Easy Mid Hard
    PointPillars LiDAR 63.07 54.00 54.01 38.53 37.20 37.28 38.46 22.60 22.49
    SECOND LiDAR 71.47 53.99 54.00 55.16 52.49 52.52 54.68 31.05 31.19
    MVXNet LiDAR+Cam 71.04 53.71 53.76 55.83 54.45 54.40 54.05 30.79 31.06
    ImvoxelNet Camera 44.78 37.58 37.55 6.81 6.746 6.73 21.06 13.57 13.17
    BEVFormer Camera 61.37 50.73 50.73 16.89 15.82 15.95 22.16 22.13 22.06
    BEVDepth Camera 75.50 63.58 63.67 34.95 33.42 33.27 55.67 55.47 55.34
    BEVHeight Camera 77.78 65.77 65.85 41.22 39.29 39.46 60.23 60.08 60.54

    BEVHeight establishes state-of-the-art results among vision-centric detectors, outperforming BEVDepth by 2.19% on Vehicle (Mid), 5.87% on Pedestrian (Mid), and 4.61% on Cyclist (Mid), while also surpassing LiDAR-only baselines (PointPillars, SECOND) across all three classes.

  7. Knowl 7 — 3D Object Detection Performance on Rope3D Dataset

    data/table

    Evaluation results on the validation set of the Rope3D roadside benchmark for Car and Big Vehicle categories using 40-recall-point 3D Average Precision (AP3D∣R40AP_{3D|R40}, abbreviated as AP) and the consolidated RopescoreRope_{score} (abbreviated as Rope), evaluated at IoU thresholds of 0.5 and 0.7 under the homologous dataset split.

    Method IoU = 0.5 IoU = 0.7
    Car Big Vehicle Car Big Vehicle
    AP Rope AP Rope AP Rope AP Rope
    M3D-RPN 54.19 62.65 33.05 44.94 16.75 32.90 6.86 24.19
    Kinematic3D 50.57 58.86 37.60 48.08 17.74 32.90 6.10 22.88
    MonoDLE 51.70 60.36 40.34 50.07 13.58 29.46 9.63 25.80
    MonoFlex 60.33 66.86 37.33 47.96 33.78 46.12 10.08 26.16
    BEVFormer 50.62 58.78 34.58 45.16 24.64 38.71 10.05 25.56
    BEVDepth 69.63 74.70 45.02 54.64 42.56 53.05 21.47 35.82
    BEVHeight 74.60 78.72 48.93 57.70 45.73 55.62 23.07 37.04

    BEVHeight outperforms prior monocular and BEV methods across all metrics and classes, surpassing BEVDepth by 4.97% AP / 4.02% Rope on Car (IoU=0.5) and 3.91% AP / 3.06% Rope on Big Vehicle (IoU=0.5).

  8. Knowl 8 — Robustness Against Camera Extrinsic Perturbations

    data/table

    Roadside cameras undergo real-world pose shifts due to wind vibrations and maintenance. Extrinsic perturbations are simulated on the DAIR-V2X-I dataset by injecting zero-mean Gaussian rotational noise N(0,1.67∘)\mathcal{N}(0, 1.67^\circ) into the camera roll and pitch angles.

    Model Disturbed Vehicle (IoU=0.5) Pedestrian (IoU=0.25) Cyclist (IoU=0.25)
    roll pitch Easy Mid Hard Easy Mid Hard Easy Mid Hard
    BEVFormer 61.37 50.73 50.73 16.89 15.82 15.95 22.16 22.13 22.00
    ✓ 50.65 42.90 42.95 10.16 9.41 9.47 13.62 13.71 13.08
    ✓ 46.40 38.26 38.37 9.12 8.44 8.55 8.99 8.43 8.42
    ✓ ✓ 19.24 16.35 16.47 3.93 3.43 3.52 4.93 4.98 4.98
    BEVDepth 71.56 60.75 60.85 21.55 20.51 20.75 40.83 40.66 40.26
    ✓ 34.82 28.32 28.35 4.49 4.36 4.39 10.48 9.51 9.73
    ✓ 14.04 11.41 11.49 3.01 2.67 2.75 6.43 6.23 6.83
    ✓ ✓ 11.84 9.48 9.54 2.16 1.84 1.89 4.31 4.14 4.26
    BEVHeight 75.58 63.49 63.59 26.93 25.47 25.78 47.97 47.45 48.12
    ✓ 66.06 54.99 55.14 18.66 17.63 17.78 34.45 26.93 27.68
    ✓ 68.49 56.98 57.11 17.94 16.87 17.09 34.48 27.82 28.67
    ✓ ✓ 62.64 51.77 51.90 14.38 14.01 14.09 31.28 25.24 26.02

    Under combined roll and pitch perturbations, depth-based BEVDepth degrades severely from 60.75% to 9.48% Mid AP on Vehicle (-51.27%), whereas BEVHeight maintains 51.77% Mid AP (surpassing BEVDepth by 42.29% under disturbance). This demonstrates that height estimation is markedly more resilient to extrinsic camera rotations than depth estimation.

  9. Knowl 9 — Error Sensitivity Dependency on Camera Mounting Height

    empirical result

    The 3D ground localization error resulting from an error in height estimation Δh\Delta h is inversely proportional to the camera mounting height HH relative to the ground. For elevated cameras (such as roadside poles or tall trucks), the steep viewing angle ensures that errors in predicted height translate into small horizontal distance errors along the ground plane.

    To demonstrate this effect independently of roadside fixed-infrastructure settings, BEVHeight and BEVDepth were evaluated on approximately 13,000 sequences collected from a moving truck platform with a high camera mounting height of H=3.14 mH = 3.14\text{ m}, annotated with 3D bounding boxes following nuScenes protocols:

    Method Car (IoU=0.5) Big Vehicle (IoU=0.5)
    Easy Mod. Hard Easy Mod. Hard
    BEVDepth 50.05 36.82 36.82 30.15 24.74 24.74
    BEVHeight 51.77 40.96 40.96 34.65 29.01 29.01

    On the elevated truck platform, BEVHeight outperforms BEVDepth by 4.14% Moderate AP on Car and 4.27% Moderate AP on Big Vehicle, confirming that the advantage of height-based projection is governed by camera elevation geometry rather than the assumption of a static ground plane.

  10. Knowl 10 — Performance Limitation on Low-Mounted Ego-Vehicle Setups

    limitation

    When evaluated on standard ego-vehicle autonomous driving benchmarks where cameras are mounted at passenger car height (approximately 1.5 m1.5\text{ m} above the ground, such as on the nuScenes dataset), BEVHeight achieves lower performance than depth-based methods.

    Evaluation on the nuScenes validation split (input resolution 256×704256 \times 704, trained for 24 epochs without CBGS, EMA, or multi-frame fusion):

    Method mAP ↑\uparrow NDS ↑\uparrow mATE ↓\downarrow mASE ↓\downarrow mAOE ↓\downarrow mAVE ↓\downarrow mAAE ↓\downarrow
    BEVDepth 0.315 0.367 0.702 0.271 0.621 1.042 0.315
    BEVDepth* 0.313 0.354 0.713 0.280 0.655 1.230 0.377
    BEVHeight 0.291 0.342 0.722 0.278 0.674 1.230 0.361
    • indicates reproduced results.

    BEVHeight trails BEVDepth by approximately 0.02 mAP and 0.012 NDS. Because the camera on a passenger vehicle has a small height HH relative to the ground, the shallow ray angle magnifies any height prediction error into a large horizontal distance error, making height-based projection sub-optimal for low-mounted vehicle cameras.

Coverage note — No substantial contributed material was omitted. All architectural components, projection algorithms, dynamic discretization equations, primary benchmark comparisons on DAIR-V2X-I and Rope3D, robustness tests, error sensitivity analyses, and limitation studies are represented.

References

  1. 1.Garrick Brazil and Xiaoming Liu. M3d-rpn: Monocular 3d region proposal network for object detection. In Proceedings of the IEEE International Conference on Computer Vision, pages 9287–9296, 2019.
  2. 2.Garrick Brazil, Gerard Pons-Moll, Xiaoming Liu, and Bernt Schiele. Kinematic 3d object detection in monocular video. In European Conference on Computer Vision, pages 135–152. Springer, 2020.
  3. 3.Holger Caesar, Varun Bankiti, Alex H Lang, Sourabh Vora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Giancarlo Baldan, and Oscar Beijbom. nuscenes: A multi-modal dataset for autonomous driving. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 11621–11631, 2020.
  4. 4.Shaoyu Chen, Xinggang Wang, Tianheng Cheng, Q. Zhang, Chang Huang, and Wenyu Liu. Polar parametrization for vision-based surround-view 3d detection. ArXiv, abs/2206.10965, 2022.
  5. 5.Jiaxun Cui, Hang Qiu, Dian Chen, Peter Stone, and Yuke Zhu. Coopernaut: End-to-end driving with cooperative perception for networked vehicles. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 17252–17262, 2022.
  6. 6.Huan Fu, Mingming Gong, Chaohui Wang, Kayhan Batmanghelich, and Dacheng Tao. Deep ordinal regression network for monocular depth estimation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2018.
  7. 7.Andreas Geiger, Philip Lenz, and Raquel Urtasun. Are we ready for autonomous driving? the kitti vision benchmark suite. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3354–3361, 2012.
  8. 8.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 770–778, 2016.
  9. 9.Junjie Huang and Guan Huang. Bevdet4d: Exploit temporal cues in multi-camera 3d object detection. arXiv preprint arXiv:2203.17054, 2022.
  10. 10.Junjie Huang, Guan Huang, Zheng Zhu, and Dalong Du. Bevdet: High-performance multi-camera 3d object detection in bird-eye-view. arXiv preprint arXiv:2112.11790, 2021.
  11. 11.Lei Huang and Wenzhun Huang. Rd-yolo: An effective and efficient object detector for roadside perception system. Sensors, 22(21):8097, 2022.
  12. 12.Yanqin Jiang, Li Zhang, Zhenwei Miao, Xiatian Zhu, Jin Gao, Weiming Hu, and Yu-Gang Jiang. Polarformer: Multi-camera 3d object detection with polar transformers. arXiv preprint arXiv:2206.15398, 2022.
  13. 13.Alex H Lang, Sourabh Vora, Holger Caesar, Lubing Zhou, Jiong Yang, and Oscar Beijbom. Pointpillars: Fast encoders for object detection from point clouds. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 12697–12705, 2019.
  14. 14.Yinhao Li, Han Bao, Zheng Ge, Jinrong Yang, Jianjian Sun, and Zeming Li. Bevstereo: Enhancing depth estimation in multi-view 3d object detection with dynamic temporal stereo. arXiv preprint arXiv:2209.10248, 2022.
  15. 15.Yinhao Li, Zheng Ge, Guanyi Yu, Jinrong Yang, Zengran Wang, Yukang Shi, Jianjian Sun, and Zeming Li. Bevdepth: Acquisition of reliable depth for multi-view 3d object detection. arXiv preprint arXiv:2206.10092, 2022.
  16. 16.Zhiqi Li, Wenhai Wang, Hongyang Li, Enze Xie, Chonghao Sima, Tong Lu, Qiao Yu, and Jifeng Dai. Bevformer: Learning bird’s-eye-view representation from multi-camera images via spatiotemporal transformers. arXiv preprint arXiv:2203.17270, 2022.
  17. 17.Yingfei Liu, Tiancai Wang, Xiangyu Zhang, and Jian Sun. Petr: Position embedding transformation for multi-view 3d object detection. arXiv preprint arXiv:2203.05625, 2022.
  18. 18.Yingfei Liu, Junjie Yan, Fan Jia, Shuailin Li, Qi Gao, Tiancai Wang, Xiangyu Zhang, and Jian Sun. Petrv2: A unified framework for 3d perception from multi-camera images. arXiv preprint arXiv:2206.01256, 2022.
  19. 19.Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101, 2017.
  20. 20.Xinzhu Ma, Yinmin Zhang, Dan Xu, Dongzhan Zhou, Shuai Yi, Haojie Li, and Wanli Ouyang. Delving into localization errors for monocular 3d object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4721–4730, 2021.
  21. 21.Jinhyung Park, Chenfeng Xu, Shijia Yang, Kurt Keutzer, Kris Kitani, Masayoshi Tomizuka, and Wei Zhan. Time will tell: New outlooks and a baseline for temporal multi-view 3d object detection. arXiv preprint arXiv:2210.02443, 2022.
  22. 22.Jonah Philion and Sanja Fidler. Lift, splat, shoot: Encoding images from arbitrary camera rigs by implicitly unprojecting to 3d. In European Conference on Computer Vision, pages 194–210. Springer, 2020.
  23. 23.Cody Reading, Ali Harakeh, Julia Chae, and Steven L Waslander. Categorical depth distribution network for monocular 3d object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8555–8564, 2021.
  24. 24.Danila Rukhovich, Anna Vorontsova, and Anton Konushin. Imvoxelnet: Image to voxels projection for monocular and multi-view general-purpose 3d object detection. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 2397–2406, 2022.
  25. 25.Avishkar Saha, Oscar Alejandro Mendez Maldonado, Chris Russell, and R. Bowden. Translating images into maps. 2022 International Conference on Robotics and Automation (ICRA), pages 9200–9206, 2022.
  26. 26.Andrea Simonelli, Samuel Rota Bulò, Lorenzo Porzi, Manuel Lopez-Antequera, and Peter Kontschieder. Disentangling monocular 3d object detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 1991–1999, 2019.
  27. 27.Vishwanath A Sindagi, Yin Zhou, and Oncel Tuzel. Mvxnet: Multimodal voxelnet for 3d object detection. In 2019 International Conference on Robotics and Automation (ICRA), pages 7276–7282. IEEE, 2019.
  28. 28.Zhiying Song, Fuxi Wen, Hailiang Zhang, and Jun Li. An efficient and robust object-level cooperative perception framework for connected and automated driving. arXiv preprint arXiv:2210.06289, 2022.
  29. 29.Pei Sun, Henrik Kretzschmar, Xerxes Dotiwalla, Aurelien Chouard, Vijaysai Patnaik, Paul Tsui, James Guo, Yin Zhou, Yuning Chai, Benjamin Caine, et al. Scalability in perception for autonomous driving: Waymo open dataset. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 2446–2454, 2020.
  30. 30.Yunlei Tang, Sebastian Dorn, and Chiragkumar Savani. Center3d: Center-based monocular 3d object detection with joint depth understanding. arXiv: Computer Vision and Pattern Recognition, 2020.
  31. 31.Ching-Yu Tseng, Yi-Rong Chen, Hsin-Ying Lee, Tsung-Han Wu, Wen-Chin Chen, and Winston Hsu. Crossdtr: Crossview and depth-guided transformers for 3d object detection. arXiv preprint arXiv:2209.13507, 2022.
  32. 32.Yue Wang, Vitor Campagnolo Guizilini, Tianyuan Zhang, Yilun Wang, Hang Zhao, and Justin Solomon. Detr3d: 3d object detection from multi-view images via 3d-to-2d queries. In Conference on Robot Learning, pages 180–191. PMLR, 2022.
  33. 33.Zengran Wang, Chen Min, Zheng Ge, Yinhao Li, Zeming Li, Hongyu Yang, and Di Huang. Sts: Surround-view temporal stereo for multi-view 3d detection. arXiv preprint arXiv:2208.10145, 2022.
  34. 34.Runsheng Xu, Hao Xiang, Zhengzhong Tu, Xin Xia, Ming-Hsuan Yang, and Jiaqi Ma. V2x-vit: Vehicle-to-everything cooperative perception with vision transformer. arXiv preprint arXiv:2203.10638, 2022.
  35. 35.Yan Yan, Yuxing Mao, and Bo Li. Second: Sparsely embedded convolutional detection. Sensors, 18(10):3337, 2018.
  36. 36.Xiaoqing Ye, Mao Shu, Hanyu Li, Yifeng Shi, Yingying Li, Guangjie Wang, Xiao Tan, and Errui Ding. Rope3d: The roadside perception dataset for autonomous driving and monocular 3d object detection task. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 21341–21350, 2022.
  37. 37.Haibao Yu, Yizhen Luo, Mao Shu, Yiyi Huo, Zebang Yang, Yifeng Shi, Zhenglong Guo, Hanyu Li, Xing Hu, Jirui Yuan, et al. Dair-v2x: A large-scale dataset for vehicle-infrastructure cooperative 3d object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 21361–21370, 2022.
  38. 38.Kaicheng Yu, Tang Tao, Hongwei Xie, Zhiwei Lin, Zhongwei Wu, Zhongyu Xia, Tingting Liang, Haiyang Sun, Jiong Deng, Dayang Hao, et al. Benchmarking the robustness of lidar-camera fusion for 3d object detection. arXiv preprint arXiv:2205.14951, 2022.
  39. 39.Yunpeng Zhang, Jiwen Lu, and Jie Zhou. Objects are different: Flexible monocular 3d object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3289–3298, 2021.
  40. 40.Benjin Zhu, Zhengkai Jiang, Xiangxin Zhou, Zeming Li, and Gang Yu. Class-balanced grouping and sampling for point cloud 3d object detection. arXiv preprint arXiv:1908.09492, 2019.
  41. 41.Xizhou Zhu, Han Hu, Stephen Lin, and Jifeng Dai. Deformable convnets v2: More deformable, better results. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018.

Citation

MLA
Yang, L., et al. “BEVHeight: A Robust Framework for Vision-based Roadside 3D Object Detection”. arXiv, 2023, http://arxiv.org/abs/2303.08498v2.
APA
Yang, L., Yu, K., Tang, T., Li, J., Yuan, K., Wang, L., Zhang, X., & Chen, P. (2023). BEVHeight: A Robust Framework for Vision-based Roadside 3D Object Detection. arXiv. http://arxiv.org/abs/2303.08498v2
Chicago
Yang, L., K. Yu, T. Tang, et al. 2023. “BEVHeight: A Robust Framework for Vision-based Roadside 3D Object Detection”. arXiv. http://arxiv.org/abs/2303.08498v2.
Harvard
Yang, L. et al. (2023) “BEVHeight: A Robust Framework for Vision-based Roadside 3D Object Detection”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2303.08498v2.
Vancouver
1. Yang L, Yu K, Tang T, Li J, Yuan K, Wang L, Zhang X, Chen P (2023) BEVHeight: A Robust Framework for Vision-based Roadside 3D Object Detection. arXiv

BibTeX

@article{yang2023bevheight,
  title = {BEVHeight: A Robust Framework for Vision-based Roadside 3D Object Detection},
  author = {Yang, Lei and Yu, Kaicheng and Tang, Tao and Li, Jun and Yuan, Kun and Wang, Li and Zhang, Xinyu and Chen, Peng},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2303.08498v2},
  eprint = {2303.08498}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE