BEVFusion: A Simple and Robust LiDAR-Camera Fusion Framework

Tingting LiangHongwei XieKaicheng YuZhongyu XiaZhiwei LinYongtao WangTao TangBing WangZhi Tang

article2022NeurIPS643 citations

Presents BEVFusion, a LiDAR-camera fusion framework for 3D object detection that eliminates camera dependence on LiDAR inputs, achieving superior performance on standard benchmarks while maintaining high accuracy during sensor malfunctions in autonomous driving scenarios.

Listen

Autonomous driving systems rely heavily on accurate three-dimensional object detection, commonly using a combination of cameras and laser scanners known as LiDAR. Traditional fusion methods use LiDAR data as queries to pull corresponding visual features from camera images. However, this creates a severe operational vulnerability: if the LiDAR sensor malfunctions, has a restricted viewing range, or fails to detect reflections from certain object surfaces, the entire detection pipeline fails. This fundamental limitation undermines vehicle safety and real-world deployment.

To solve this problem, the article presents BEVFusion, a simple and robust perception framework designed to decouple the camera and LiDAR processing paths. The primary objective is to build and evaluate an architecture where both camera images and LiDAR point clouds are processed into a unified bird's-eye-view space through two independent streams. A lightweight dynamic fusion module then merges these representations before passing them to standard detection heads, ensuring that a failure in one sensor stream does not disable the other.

The authors evaluated the framework on the large-scale nuScenes benchmark, which contains over a million annotated 3D bounding boxes across ten object categories. They integrated the framework with three distinct LiDAR backbones and tested performance under standard operating conditions as well as simulated hardware and environmental faults, such as restricted sensor viewing angles, dropped point reflections, and disabled camera streams.

The findings show that BEVFusion significantly improves overall performance and resilience. Under standard conditions, it achieved state-of-the-art detection precision, reaching up to 71.3% mean average precision (mAP) and outperforming existing fusion models. When integrated into standard LiDAR-only baselines, it boosted their precision by 3.0% to 18.4%. Most notably, under simulated LiDAR failure scenarios where points were lost, BEVFusion outperformed leading baseline methods by margins of 15.7% to 28.9% mAP. It also maintained superior accuracy when subjected to camera sensor dropouts, severe lighting shifts, and for distant objects beyond 30 meters.

These results demonstrate that multi-sensor autonomy architectures must avoid sequential dependencies that allow single-point hardware failures to compromise the entire perception system. By processing both modalities into a shared bird's-eye-view space independently, autonomous platforms can achieve higher detection accuracy during normal operation while maintaining critical safety fallbacks during sensor degradation or adverse weather.

Teams developing autonomous systems should adopt independent, dual-stream bird's-eye-view fusion architectures to eliminate single-point sensor vulnerabilities. For immediate deployment, engineering teams must focus on optimizing runtime latency; the current prototype takes approximately 1.5 seconds per frame due to processing bottlenecks in the image-to-3D projection module. Further development should explore concurrent processing pipelines, temporal multi-frame fusion, and fine-grained intermediate feature alignment to make the architecture viable for real-time vehicular control.

While the empirical evidence strongly supports the robustness of the framework across extensive benchmark tests, some boundaries remain. The model relies on supervised training distributions, and detection will inevitably fail if both sensor streams miss an object entirely. With appropriate runtime optimization and validation on target vehicle hardware, confidence in the architectural benefits of this approach remains high.

Cover for BEVFusion: A Simple and Robust LiDAR-Camera Fusion Framework

Abstract

Fusing the camera and LiDAR information has become a de-facto standard for 3D object detection tasks. Current methods rely on point clouds from the LiDAR sensor as queries to leverage the feature from the image space. However, people discovered that this underlying assumption makes the current fusion framework infeasible to produce any prediction when there is a LiDAR malfunction, regardless of minor or major. This fundamentally limits the deployment capability to realistic autonomous driving scenarios. In contrast, we propose a surprisingly simple yet novel fusion framework, dubbed BEVFusion, whose camera stream does not depend on the input of LiDAR data, thus addressing the downside of previous methods. We empirically show that our framework surpasses the state-of-the-art methods under the normal training settings. Under the robustness training settings that simulate various LiDAR malfunctions, our framework significantly surpasses the state-of-the-art methods by 15.7% to 28.9% mAP. To the best of our knowledge, we are the first to handle realistic LiDAR malfunction and can be deployed to realistic scenarios without any post-processing procedure. The code is available at this https URL.

Table of Contents

  • 1 Introduction
  • 2 Related Works
  • 3 BEVFusion: A General Framework for LiDAR-Camera Fusion
  • 3.1 Camera stream architecture: From multi-view images to BEV space
  • 3.2 LiDAR stream architecture: From point clouds to BEV space
  • 3.3 Dynamic fusion module
  • 3.4 Detection head
  • 4 Experiments
  • 4.1 Experimental settings
  • 4.2 Generalization ability
  • 4.3 Comparing with the state-of-the-art methods
  • 4.4 Robustness experiments
  • 4.4.1 Robustness experiments against LiDAR Malfunctions
  • 4.4.2 Robustness against camera malfunctions
  • 4.5 Ablation
  • 5 Conclusion
  • References
  • A Network architectures
  • B Experimental settings
  • B.1 Implementation details.
  • C More robustness analysis
  • C.1 Robustness of Dynamic Fusion Module
  • C.2 Robustness under both modality malfunctions
  • C.3 Robustness against Inferior Image Conditions
  • D Performance for different distance regions
  • E Latency
  • F Failure cases analysis

Knowls

  1. Knowl 1 — BEVFusion Disentangled Multi-Modal 3D Detection Architecture

    model/method

    BEVFusion is a LiDAR-camera fusion framework for 3D object detection that eliminates the asymmetric dependency of the camera stream on LiDAR input. Prior point-level or feature-level fusion approaches project raw LiDAR points or LiDAR-generated bounding-box proposals into the image space to sample camera features, causing complete perception failure whenever LiDAR data is missing or degraded.

    Instead, BEVFusion uses two fully independent parallel streams that encode raw multi-view camera images and raw LiDAR point clouds separately into feature representations within a shared bird's-eye view (BEV) space:

    • Camera BEV feature: FCamera∈RX×Y×CCameraF_{\text{Camera}} \in \mathbb{R}^{X \times Y \times C_{\text{Camera}}}
    • LiDAR BEV feature: FLiDAR∈RX×Y×CLiDARF_{\text{LiDAR}} \in \mathbb{R}^{X \times Y \times C_{\text{LiDAR}}}

    where XX and YY represent the spatial grid dimensions of the BEV map, and CCameraC_{\text{Camera}} and CLiDARC_{\text{LiDAR}} represent the respective channel dimensions. The two BEV feature representations are subsequently integrated through a dynamic fusion module into a unified representation Ffused∈RX×Y×CLiDARF_{\text{fused}} \in \mathbb{R}^{X \times Y \times C_{\text{LiDAR}}}, which is directly consumed by standard 3D bounding-box detection heads.

  2. Knowl 2 — Dynamic Fusion Module for BEV Feature Fusion

    equation

    The dynamic fusion module fuses the camera BEV feature map FCamera∈RX×Y×CCameraF_{\text{Camera}} \in \mathbb{R}^{X \times Y \times C_{\text{Camera}}} and LiDAR BEV feature map FLiDAR∈RX×Y×CLiDARF_{\text{LiDAR}} \in \mathbb{R}^{X \times Y \times C_{\text{LiDAR}}} into a single unified feature representation Ffused∈RX×Y×CLiDARF_{\text{fused}} \in \mathbb{R}^{X \times Y \times C_{\text{LiDAR}}}.

    The fusion process is defined by:

    Ffused=fadaptive(fstatic([FCamera,FLiDAR]))F_{\text{fused}} = f_{\text{adaptive}}(f_{\text{static}}([F_{\text{Camera}}, F_{\text{LiDAR}}]))

    where [⋅,⋅][\cdot, \cdot] denotes concatenation along the channel dimension, resulting in a tensor in RX×Y×(CCamera+CLiDAR)\mathbb{R}^{X \times Y \times (C_{\text{Camera}} + C_{\text{LiDAR}})}.

    The static fusion function fstaticf_{\text{static}} consists of a 3×33 \times 3 convolution layer that performs spatial and channel mixing while reducing the channel dimension from CCamera+CLiDARC_{\text{Camera}} + C_{\text{LiDAR}} to CLiDARC_{\text{LiDAR}}.

    Given the intermediate feature F∈RX×Y×CLiDARF \in \mathbb{R}^{X \times Y \times C_{\text{LiDAR}}}, the adaptive feature selection function fadaptivef_{\text{adaptive}} dynamically selects channels using a channel attention mechanism:

    fadaptive(F)=σ(Wfavg(F))⊙Ff_{\text{adaptive}}(F) = \sigma(W f_{\text{avg}}(F)) \odot F

    where favg(F)∈R1×1×CLiDARf_{\text{avg}}(F) \in \mathbb{R}^{1 \times 1 \times C_{\text{LiDAR}}} denotes global average pooling across spatial dimensions (X,Y)(X, Y), W∈RCLiDAR×CLiDARW \in \mathbb{R}^{C_{\text{LiDAR}} \times C_{\text{LiDAR}}} is a learnable linear transformation implemented via a 1×11 \times 1 convolution, σ(⋅)\sigma(\cdot) denotes the element-wise sigmoid activation function, and ⊙\odot represents channel-wise broadcast multiplication.

  3. Knowl 3 — Camera Stream Architecture with Adaptive FPN and Spatial-to-Channel BEV Encoder

    model/method

    The camera stream transforms multi-view 2D camera images into a 3D BEV representation FCamera∈RX×Y×CCameraF_{\text{Camera}} \in \mathbb{R}^{X \times Y \times C_{\text{Camera}}} across four sequential stages:

    1. Image-view Encoder: A Dual-Swin-Tiny 2D backbone extracts multi-scale feature maps from each camera view of input size H×W×3H \times W \times 3. A Feature Pyramid Network (FPN) produces multi-scale features F2,F3,F4,F5F_2, F_3, F_4, F_5 with shapes (H/4×W/4×C)(H/4 \times W/4 \times C), (H/8×W/8×C)(H/8 \times W/8 \times C), (H/16×W/16×C)(H/16 \times W/16 \times C), and (H/32×W/32×C)(H/32 \times W/32 \times C).

    2. Feature Adaptive Module (ADP): To align and merge multi-scale FPN features without spatial degradation, each level FiF_i passes through an adaptive module comprising bilinear upsampling to resolution H/4×W/4H/4 \times W/4, adaptive average pooling, and a 1×11 \times 1 convolution. The processed features are concatenated along the channel dimension and mixed via a 1×11 \times 1 convolution to produce a unified image feature map of shape H/4×W/4×CH/4 \times W/4 \times C.

    3. 2D-to-3D View Projector: Adopting the Lift-Splat-Shoot (LSS) approach, the module predicts a categorical depth distribution for each pixel, unprojects 2D features into 3D ego-car coordinates using known camera extrinsics and intrinsics, and aggregates them into a predefined 3D pseudo-voxel grid V∈RX×Y×Z×CV \in \mathbb{R}^{X \times Y \times Z \times C}.

    4. Spatial-to-Channel (S2C) BEV Encoder: Rather than downsampling through 3D pooling or strided 3D convolutions, the 4D pseudo-voxel tensor VV is flattened along the vertical axis by reshaping its depth dimension ZZ into channels: V∈RX×Y×(ZC)V \in \mathbb{R}^{X \times Y \times (ZC)}. Four successive 3×33 \times 3 2D convolutional layers gradually reduce the channel dimension to CCameraC_{\text{Camera}} at full BEV spatial resolution (X×Y)(X \times Y), preserving fine-grained spatial and semantic details.

  4. Knowl 4 — Two-Stage Training Pipeline and Sensor Malfunction Simulation Augmentations

    experimental setup

    BEVFusion employs a two-stage training scheme and specific data augmentation protocols for simulating real-world sensor corruptions:

    Training Scheme:

    • Stage 1 (Modality Pretraining): The LiDAR and camera streams are trained separately on their respective single-modality inputs following standard MMDetection3D configurations. The LiDAR stream utilizes global scaling, rotation, random flips along XX and YY axes, and Class-Balanced Grouping and Sampling (CBGS). The camera stream initializes its backbone and neck from a nuImages pre-trained Mask R-CNN Dual-Swin-Tiny model.
    • Stage 2 (Joint Fusion Training): The fused model inherits the pretrained weights. Joint training uses a batch size of 32 across 8 Nvidia V100 GPUs with an initial learning rate of 10−310^{-3}. When using PointPillars, the detector is trained for 12 epochs with learning rate drops of 10×10\times at epochs 8 and 11. When using CenterPoint or TransFusion-L, the model is trained for 6 epochs with 10×10\times drops at epochs 4 and 5, and the LiDAR stream weights are frozen.

    Robustness Augmentations:

    • Limited LiDAR Field-of-View (FOV): Truncates LiDAR point clouds to restricted azimuth ranges [−π/2,π/2][-\pi/2, \pi/2] (180∘180^\circ) or [−π/3,π/3][-\pi/3, \pi/3] (120∘120^\circ) to simulate front-facing semi-solid lidars or sensor hardware sector loss.
    • Object Point Dropping: To simulate low-reflectivity objects (e.g., in rain or dark conditions), each frame has a 0.50.5 probability of triggering dropping, and each annotated 3D bounding box within that frame has an independent 0.50.5 probability of having all LiDAR points within its volume removed.
    • Robustness Fine-tuning: LiDAR-camera detectors are fine-tuned on augmented datasets for 12 epochs with an initial learning rate of 10−410^{-4}, decayed by 10×10\times at epochs 8 and 11.
  5. Knowl 5 — 3D Object Detection Performance on the nuScenes Benchmark

    data/table

    On the nuScenes benchmark (evaluated using mean Average Precision, mAP, and nuScenes Detection Score, NDS, across 10 object classes), BEVFusion with a single detection head achieves state-of-the-art performance, surpassing previous single-modality and fusion approaches.

    Method Modality mAP (%) NDS (%)
    nuScenes Validation Set
    FUTR3D Camera + LiDAR 64.2 68.0
    BEVFusion Camera + LiDAR 67.9 71.0
    BEVFusion* (with BEV augmentation) Camera + LiDAR 69.6 72.1
    nuScenes Test Set
    PointPillars LiDAR 30.5 45.3
    CBGS LiDAR 52.8 63.3
    CenterPoint LiDAR 60.3 67.3
    TransFusion-L LiDAR 65.5 70.2
    PointPainting Camera + LiDAR 46.4 58.1
    3D-CVF Camera + LiDAR 52.7 62.3
    PointAugmenting Camera + LiDAR 66.8 71.0
    MVP Camera + LiDAR 66.4 70.5
    FusionPainting Camera + LiDAR 68.1 71.6
    TransFusion Camera + LiDAR 68.9 71.7
    BEVFusion (Ours) Camera + LiDAR 69.2 71.8
    BEVFusion (Ours)* (with BEV augmentation) Camera + LiDAR 71.3 73.3

    Without test-time augmentation or ensembling, BEVFusion (using TransFusion-L as its LiDAR stream) attains 69.2% mAP and 71.8% NDS, outperforming the two-stage TransFusion baseline (68.9% mAP, 71.7% NDS). When trained with BEV-space data augmentation, BEVFusion reaches 71.3% mAP and 73.3% NDS on the test set.

  6. Knowl 6 — Generalization Across Diverse LiDAR Backbones and Detection Heads

    data/table

    BEVFusion is modular and integrates across pillar-based, voxel-based, and transformer-based 3D LiDAR backbones and detection heads. Evaluated on the nuScenes validation set, fusing the camera stream via BEVFusion consistently improves single-modality LiDAR baselines:

    PointPillars CenterPoint TransFusion-L
    Modality mAP (%) NDS (%) mAP (%) NDS (%) mAP (%) NDS (%)
    Camera only 22.9 31.1 27.1 32.1 22.7 26.1
    LiDAR only 35.1 49.8 57.1 65.4 64.9 69.9
    BEVFusion (Joint) 53.5 60.4 64.2 68.0 67.9 71.0
    Improvement over LiDAR +18.4 +10.6 +7.1 +2.6 +3.0 +1.1

    BEVFusion provides substantial gains across all detector paradigms, lifting PointPillars by 18.4% mAP and 10.6% NDS, CenterPoint by 7.1% mAP and 2.6% NDS, and TransFusion-L by 3.0% mAP and 1.1% NDS.

  7. Knowl 7 — Detection Robustness Under LiDAR Sensor Failure and Object Reflection Loss

    data/table

    BEVFusion exhibits strong robustness to simulated LiDAR corruptions on the nuScenes validation set, whereas proposal-queried fusion methods (such as TransFusion LC) fail due to their strict reliance on LiDAR points for cross-modal querying.

    1. Limited LiDAR Field of View (FOV):

    PointPillars CenterPoint TransFusion
    FOV Metric LiDAR BEVFusion LiDAR BEVFusion LiDAR BEVFusion TransFusion LC
    [−π/2,π/2][-\pi/2, \pi/2] mAP (%) 12.4 36.8 (+24.4) 23.6 45.5 (+21.9) 27.8 46.4 (+18.6) 31.1
    NDS (%) 37.1 45.8 (+8.7) 48.0 54.9 (+6.9) 50.5 55.8 (+5.3) 49.2
    [−π/3,π/3][-\pi/3, \pi/3] mAP (%) 8.4 33.5 (+25.1) 15.9 40.9 (+25.0) 19.0 41.5 (+22.5) 21.0
    NDS (%) 34.3 42.1 (+7.8) 43.5 49.9 (+6.4) 45.3 50.8 (+5.5) 41.2

    2. Object Failure (50% Object Point Drop):

    PointPillars CenterPoint TransFusion
    Augmentation Metric LiDAR BEVFusion LiDAR BEVFusion LiDAR BEVFusion TransFusion LC
    Without Aug. mAP (%) 12.7 34.3 (+21.6) 31.3 40.2 (+8.9) 34.6 40.8 (+6.2) 38.1
    NDS (%) 36.6 49.1 (+12.5) 50.7 54.3 (+3.6) 53.6 56.0 (+2.4) 55.4
    With Aug. mAP (%) - 41.6 (+28.9) - 54.0 (+22.7) - 50.3 (+15.7) 37.2
    NDS (%) - 51.9 (+15.3) - 61.6 (+10.9) - 57.6 (+4.0) 51.1

    When fine-tuned with robustness augmentation, BEVFusion outperforms TransFusion LC by 13.1% mAP (50.3% vs. 37.2%) and boosts the LiDAR baseline by +15.7% to +28.9% mAP across models.

  8. Knowl 8 — Detection Robustness Under Camera Sensor Malfunctions

    data/table

    The robustness of BEVFusion and competing approaches was evaluated on the nuScenes validation set across three simulated camera failure modes: missing the front camera, preserving only the front camera, and experiencing a 50% stuck frame rate.

    Clean Missing Front Preserve Front 50% Stuck
    Approach mAP NDS mAP NDS mAP NDS mAP NDS
    DETR3D (Camera only) 34.9 43.4 25.8 39.2 3.3 20.5 17.3 32.3
    PointAugmenting 46.9 55.6 42.4 53.0 31.6 46.5 42.1 52.8
    MVX-Net 61.0 66.1 47.8 59.4 17.5 41.7 48.3 58.8
    TransFusion 66.9 70.9 65.3 70.1 64.4 69.3 65.9 70.2
    BEVFusion 67.9 71.0 65.9 70.7 65.1 69.9 66.2 70.3

    BEVFusion maintains superior performance across all camera failure scenarios. In the extreme case where five out of six surrounding cameras are missing (Preserve Front), BEVFusion retains 65.1% mAP and 69.9% NDS, outperforming DETR3D (3.3% mAP), MVX-Net (17.5% mAP), and PointAugmenting (31.6% mAP).

  9. Knowl 9 — Performance Evaluation Across Object Distance Ranges

    data/table

    Because LiDAR point clouds become sparser as distance increases, the independent camera BEV stream in BEVFusion provides complementary semantic cues at long ranges.

    PointPillars mAP (%) CenterPoint mAP (%) TransFusion-L mAP (%)
    Modality <15m<15\text{m} 15–30m15\text{--}30\text{m} >30m>30\text{m} <15m<15\text{m} 15–30m15\text{--}30\text{m} >30m>30\text{m} <15m<15\text{m} 15–30m15\text{--}30\text{m} >30m>30\text{m}
    Camera only 28.2 21.2 15.1 73.1 57.8 33.6 76.3 66.1 43.2
    LiDAR only 22.0 12.9 4.6 49.1 23.1 5.8 41.9 19.2 4.9
    BEVFusion 32.5 27.7 20.9 77.7 65.0 42.9 77.3 69.4 49.2
    Gain over LiDAR +10.5 +14.8 +16.3 +28.6 +41.9 +37.1 +35.4 +50.2 +44.3

    Comparing BEVFusion with TransFusion across distant regions (>30m>30\text{m}), BEVFusion achieves 49.2% mAP versus TransFusion's 43.7% (original) and 47.7% (re-implemented), confirming that multi-modal fusion in a shared BEV space offers the largest relative advantages where LiDAR returns are minimal.

  10. Knowl 10 — Ablation Analysis of Camera Stream and Dynamic Fusion Design Components

    empirical result

    Ablation studies on the nuScenes validation set analyze the individual contributions of the camera stream architecture and the components of the Dynamic Fusion Module:

    1. Camera Stream Modules (using PointPillars head):

      • Naive Baseline (ResNet50 + FPN + ResNet18 BEV encoder from LSS): achieves 13.9% mAP and 24.5% NDS.
      • Spatial-to-Channel (S2C) BEV Encoder: replacing ResNet18 BEV encoder with S2C reshaping and convolutions increases performance to 17.9% mAP and 27.0% NDS (+4.0% mAP, +2.5% NDS).
      • Adaptive Module (ADP) in FPN: adding ADP improves performance to 18.0% mAP and 27.1% NDS (+0.1% mAP, +0.1% NDS).
      • Dual-Swin-Tiny Backbone: replacing ResNet50 with Dual-Swin-Tiny brings performance to 22.9% mAP and 31.1% NDS (+4.9% mAP, +4.0% NDS).
    2. Dynamic Fusion Components:

      • LiDAR baseline: PointPillars (35.1% mAP, 49.8% NDS), CenterPoint (57.1% mAP, 65.4% NDS), TransFusion-L (64.9% mAP, 69.9% NDS).
      • Channel & Spatial Fusion (CSF): static 3×33 \times 3 convolution improves mAP to 51.6% (+16.5%) for PointPillars, 63.0% (+5.9%) for CenterPoint, and 67.3% (+2.4%) for TransFusion-L.
      • Adaptive Feature Selection (AFS): adding channel attention further increases mAP to 53.5% (+1.9%) for PointPillars, 64.2% (+1.2%) for CenterPoint, and 67.9% (+0.6%) for TransFusion-L.
  11. Knowl 11 — View Projection Latency Bottleneck and Simultaneous Modality Occlusion Failure

    limitation

    BEVFusion has two main limitations identified by empirical latency profiling and failure case analysis:

    1. Inference Latency Bottleneck: When evaluated on an Intel Xeon Gold 6126 CPU @ 2.60GHz and an Nvidia V100 GPU (batch size 1), total latency is 1464–1530 ms per frame. The primary computational bottleneck is the camera stream's Lift-Splat-Shoot (LSS) 2D-to-3D view projection module, which alone accounts for over 957 ms (>60% of total inference time), whereas the dynamic fusion module and LiDAR streams are substantially faster (LiDAR stream alone takes 189–264 ms).

    2. Simultaneous Sensor Failure Limit: While BEVFusion successfully recovers detection when either the LiDAR or camera stream fails independently, it cannot identify an object if it is simultaneously undetected or uncaptured by both modalities.

Coverage note — None was omitted; all key architectural components, training setups, robustness augmentations, benchmark evaluations, ablations, distance breakdowns, and runtime/failure analyses are covered.

References

  1. 1.Xuyang Bai, Zeyu Hu, Xinge Zhu, Qingqiu Huang, Yilun Chen, Hongbo Fu, and Chiew-Lan Tai. Transfusion: Robust lidar-camera fusion for 3d object detection with transformers. arXiv preprint arXiv:2203.11496, 2022.
  2. 2.Holger Caesar, Varun Bankiti, Alex H. Lang, Sourabh Vora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Giancarlo Baldan, and Oscar Beijbom. nuScenes: A multimodal dataset for autonomous driving. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2020.
  3. 3.Simon Chadwick, Will Maddern, and Paul Newman. Distant vehicle detection using radar and vision. In International Conference on Robotics and Automation (ICRA), 2019.
  4. 4.Xiaozhi Chen, Huimin Ma, Jixiang Wan, B. Li, and Tian Xia. Multi-view 3d object detection network for autonomous driving. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017.
  5. 5.Xuanyao Chen, Tianyuan Zhang, Yue Wang, Yilun Wang, and Hang Zhao. Futr3d: A unified sensor fusion framework for 3d detection. arXiv preprint arXiv:2203.10642, 2022.
  6. 6.Yukang Chen, Yanwei Li, Xiangyu Zhang, Jian Sun, and Jiaya Jia. Focal sparse convolutional networks for 3d object detection. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2022.
  7. 7.Zhiyu Chong, Xinzhu Ma, Hong Zhang, Yuxin Yue, Haojie Li, Zhihui Wang, and Wanli Ouyang. Monodistill: Learning spatial features for monocular 3d object detection. arXiv preprint arXiv:2201.10830, 2022.
  8. 8.MMDetection3D Contributors. Mmdetection3d: Open-mmlab next-generation platform for general 3d object detection. https://github.com/open-mmlab/mmdetection3d, 2020.
  9. 9.M. Everingham, L. Gool, Christopher K. I. Williams, J. Winn, and Andrew Zisserman. The pascal visual object classes (voc) challenge. International Journal on Computer Vision (IJCV), 2009.
  10. 10.Lue Fan, Xuan Xiong, Feng Wang, Naiyan Wang, and Zhaoxiang Zhang. Rangedet: In defense of range view for lidar-based 3d object detection. In IEEE International Conference on Computer Vision (ICCV), 2021.
  11. 11.Andreas Geiger, Philip Lenz, and R. Urtasun. Are we ready for autonomous driving? the KITTI vision benchmark suite. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2012.
  12. 12.Xiaoyang Guo, Shaoshuai Shi, Xiaogang Wang, and Hongsheng Li. Liga-stereo: Learning lidar geometry aware representations for stereo-based 3d detector. In IEEE International Conference on Computer Vision (ICCV), 2021.
  13. 13.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016.
  14. 14.Jie Hu, Li Shen, and Gang Sun. Squeeze-and-excitation networks. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018.
  15. 15.Junjie Huang and Guan Huang. Bevdet4d: Exploit temporal cues in multi-camera 3d object detection. arXiv preprint arXiv:2203.17054, 2022.
  16. 16.Junjie Huang, Guan Huang, Zheng Zhu, and Dalong Du. Bevdet: High-performance multi-camera 3d object detection in bird-eye-view. arXiv preprint arXiv:2112.11790, 2021.
  17. 17.Tengteng Huang, Zhe Liu, Xiwu Chen, and X. Bai. EPNet: Enhancing point features with image semantics for 3d object detection. In European Conference on Computer Vision (ECCV), 2020.
  18. 18.Vijay John and Seiichi Mita. Rvnet: deep sensor fusion of monocular camera and radar for image-based obstacle detection in challenging environments. In Pacific-Rim Symposium on Image and Video Technology (PSIVT), 2019.
  19. 19.Abhinav Kumar, Garrick Brazil, and Xiaoming Liu. Groomed-nms: Grouped mathematically differentiable nms for monocular 3d object detection. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2021.
  20. 20.Alex H. Lang, Sourabh Vora, Holger Caesar, Lubing Zhou, Jiong Yang, and Oscar Beijbom. PointPillars: Fast encoders for object detection from point clouds. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2019.
  21. 21.Yingwei Li, Adams Wei Yu, Tianjian Meng, Ben Caine, Jiquan Ngiam, Daiyi Peng, Junyang Shen, Bo Wu, Yifeng Lu, Denny Zhou, et al. Deepfusion: Lidar-camera deep fusion for multi-modal 3d object detection. arXiv preprint arXiv:2203.08195, 2022.
  22. 22.Zhichao Li, Feng Wang, and Naiyan Wang. Lidar r-cnn: An efficient and universal 3d object detector. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2021.
  23. 23.Zhiqi Li, Wenhai Wang, Hongyang Li, Enze Xie, Chonghao Sima, Tong Lu, Qiao Yu, and Jifeng Dai. Bevformer: Learning bird’s-eye-view representation from multi-camera images via spatiotemporal transformers. arXiv preprint arXiv:2203.17270, 2022.
  24. 24.Ming Liang, Binh Yang, Shenlong Wang, and R. Urtasun. Deep continuous fusion for multi-sensor 3d object detection. In European Conference on Computer Vision (ECCV), 2018.
  25. 25.Tingting Liang, Xiaojie Chu, Yudong Liu, Yongtao Wang, Zhi Tang, Wei Chu, Jingdong Chen, and Haibing Ling. Cbnet: A composite backbone network architecture for object detection. arXiv preprint arXiv:2107.00420, 2021.
  26. 26.Tsung-Yi Lin, Piotr Dollár, Ross B. Girshick, Kaiming He, Bharath Hariharan, and Serge J. Belongie. Feature pyramid networks for object detection. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017.
  27. 27.Zhijian Liu, Haotian Tang, Alexander Amini, Xingyu Yang, Huizi Mao, Daniela Rus, and Song Han. Bevfusion: Multi-task multi-sensor fusion with unified bird’s-eye view representation. arXiv preprint arXiv:2205.13542, 2022.
  28. 28.Zongdai Liu, Dingfu Zhou, Feixiang Lu, Jin Fang, and Liangjun Zhang. Autoshape: Real-time shape-aware monocular 3d object detection. In IEEE International Conference on Computer Vision (ICCV), 2021.
  29. 29.Yan Lu, Xinzhu Ma, Lei Yang, Tianzhu Zhang, Yating Liu, Qi Chu, Junjie Yan, and Wanli Ouyang. Geometry uncertainty projection network for monocular 3d object detection. In IEEE International Conference on Computer Vision (ICCV), 2021.
  30. 30.Ramin Nabati and Hairong Qi. Centerfusion: Center-based radar and camera fusion for 3d object detection. In IEEE Winter Conference on Applications of Computer Vision (WACV), 2021.
  31. 31.Felix Nobis, Maximilian Geisslinger, Markus Weber, Johannes Betz, and Markus Lienkamp. A deep learning-based radar and camera sensor fusion architecture for object detection. In Sensor Data Fusion: Trends, Solutions, Applications (SDF), 2019.
  32. 32.Su Pang, Daniel Morris, and Hayder Radha. Clocs: Camera-lidar object candidates fusion for 3d object detection. In IEEE International Conference on Intelligent Robots and Systems (IROS), 2020.
  33. 33.Dennis Park, Rares Ambrus, Vitor Guizilini, Jie Li, and Adrien Gaidon. Is pseudo-lidar needed for monocular 3d object detection? In IEEE International Conference on Computer Vision (ICCV), 2021.
  34. 34.Jonah Philion and S. Fidler. Lift, Splat, Shoot: Encoding images from arbitrary camera rigs by implicitly unprojecting to 3d. In European Conference on Computer Vision (ECCV), 2020.
  35. 35.Jonah Philion and Sanja Fidler. Lift, splat, shoot: Encoding images from arbitrary camera rigs by implicitly unprojecting to 3d. In European Conference on Computer Vision (ECCV), 2020.
  36. 36.C. Qi, W. Liu, Chenxia Wu, Hao Su, and L. Guibas. Frustum pointnets for 3d object detection from rgb-d data. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018.
  37. 37.C. Qi, Hao Su, Kaichun Mo, and L. Guibas. PointNet: Deep learning on point sets for 3d classification and segmentation. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017.
  38. 38.Charles Ruizhongtai Qi, Li Yi, Hao Su, and Leonidas J Guibas. Pointnet++: Deep hierarchical feature learning on point sets in a metric space. In Neural Information Processing Systems (NeurIPS), 2017.
  39. 39.Cody Reading, Ali Harakeh, Julia Chae, and Steven L Waslander. Categorical depth distribution network for monocular 3d object detection. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2021.
  40. 40.Thomas Roddick, Alex Kendall, and R. Cipolla. Orthographic feature transform for monocular 3d object detection. In British Machine Vision Conference (BMVC), 2019.
  41. 41.Shaoshuai Shi, Chaoxu Guo, Li Jiang, Zhe Wang, Jianping Shi, Xiaogang Wang, and Hongsheng Li. PV-RCNN: Point-voxel feature set abstraction for 3d object detection. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2020.
  42. 42.Shaoshuai Shi, Xiaogang Wang, and Hongsheng Li. PointRCNN: 3d object proposal generation and detection from point cloud. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2019.
  43. 43.Vishwanath A Sindagi, Yin Zhou, and Oncel Tuzel. Mvx-net: Multimodal voxelnet for 3d object detection. In International Conference on Robotics and Automation (ICRA), 2019.
  44. 44.Pei Sun, Henrik Kretzschmar, Xerxes Dotiwalla, Aurelien Chouard, Vijaysai Patnaik, P. Tsui, James Guo, Yin Zhou, Yuning Chai, Benjamin Caine, Vijay Vasudevan, Wei Han, Jiquan Ngiam, Hang Zhao, Aleksei Timofeev, S. Ettinger, Maxim Krivokon, A. Gao, Aditya Joshi, Y. Zhang, Jonathon Shlens, Zhifeng Chen, and Dragomir Anguelov. Scalability in perception for autonomous driving: Waymo open dataset. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2020.
  45. 45.Pei Sun, Weiyue Wang, Yuning Chai, Gamaleldin F. Elsayed, Alex Bewley, Xiao Zhang, Cristian Sminchisescu, and Drago Anguelov. RSN: Range sparse net for efficient, accurate lidar 3d object detection. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2021.
  46. 46.Sourabh Vora, Alex H Lang, Bassam Helou, and Oscar Beijbom. Pointpainting: Sequential fusion for 3d object detection. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2020.
  47. 47.Chunwei Wang, Chao Ma, Ming Zhu, and Xiaokang Yang. PointAugmenting: Cross-modal augmentation for 3d object detection. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2021.
  48. 48.Li Wang, Liang Du, Xiaoqing Ye, Yanwei Fu, Guodong Guo, Xiangyang Xue, Jianfeng Feng, and Li Zhang. Depth-conditioned dynamic message propagation for monocular 3d object detection. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2021.
  49. 49.Li Wang, Li Zhang, Yi Zhu, Zhi Zhang, Tong He, Mu Li, and Xiangyang Xue. Progressive coordinate transforms for monocular 3d object detection. In Neural Information Processing Systems (NeurIPS), 2021.
  50. 50.Tai Wang, ZHU Xinge, Jiangmiao Pang, and Dahua Lin. Probabilistic and geometric depth: Detecting objects in perspective. In Conference on Robot Learning (CoRL), 2022.
  51. 51.Tai Wang, Xinge Zhu, Jiangmiao Pang, and Dahua Lin. Fcos3d: Fully convolutional one-stage monocular 3d object detection. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2021.
  52. 52.Yue Wang, Alireza Fathi, Abhijit Kundu, David A. Ross, Caroline Pantofaru, Thomas A. Funkhouser, and Justin M. Solomon. Pillar-based object detection for autonomous driving. In European Conference on Computer Vision (ECCV), 2020.
  53. 53.Yue Wang, Vitor Campagnolo Guizilini, Tianyuan Zhang, Yilun Wang, Hang Zhao, and Justin Solomon. Detr3d: 3d object detection from multi-view images via 3d-to-2d queries. In Conference on Robot Learning (CoRL), 2022.
  54. 54.Enze Xie, Zhiding Yu, Daquan Zhou, Jonah Philion, Anima Anandkumar, Sanja Fidler, Ping Luo, and Jose M. Alvarez. M2\text{M}^{2}bev: Multi-camera joint 3d detection and segmentation with unified birds-eye view representation. arXiv preprint arXiv:2204.05088, 2022.
  55. 55.Shaoqing Xu, Dingfu Zhou, Jin Fang, Junbo Yin, Bin Zhou, and Liangjun Zhang. FusionPainting: Multimodal fusion with adaptive attention for 3d object detection. In IEEE International Conference on Intelligent Transportation Systems (ITSC), 2021.
  56. 56.Yan Yan, Yuxing Mao, and B. Li. SECOND: Sparsely embedded convolutional detection. Sensors, 2018.
  57. 57.Zetong Yang, Y. Sun, Shu Liu, and Jiaya Jia. 3DSSD: Point-based 3d single stage object detector. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2020.
  58. 58.Zeyu Yang, Jiaqi Chen, Zhenwei Miao, Wei Li, Xiatian Zhu, and Li Zhang. Deepinteraction: 3d object detection via modality interaction. In NeurIPS, 2022.
  59. 59.Tianwei Yin, Xingyi Zhou, and Philipp Krähenbühl. Center-based 3d object detection and tracking. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2021.
  60. 60.Tianwei Yin, Xingyi Zhou, and Philipp Krähenbühl. Multimodal virtual point 3d detection. In Neural Information Processing Systems (NeurIPS), 2021.
  61. 61.Jin Hyeok Yoo, Yeocheol Kim, Ji Song Kim, and J. Choi. 3D-CVF: Generating joint camera and lidar features using cross-view spatial feature fusion for 3d object detection. In European Conference on Computer Vision (ECCV), 2020.
  62. 62.Kaicheng Yu, Tang Tao, Hongwei Xie, Zhiwei Lin, Zhongwei Wu, Zhongyu Xia, Tingting Liang, Haiyang Sun, Jiong Deng, Dayang Hao, et al. Benchmarking the robustness of lidar-camera fusion for 3d object detection. arXiv preprint arXiv:2205.14951, 2022.
  63. 63.Yunpeng Zhang, Jiwen Lu, and Jie Zhou. Objects are different: Flexible monocular 3d object detection. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2021.
  64. 64.Yin Zhou, Pei Sun, Y. Zhang, Dragomir Anguelov, J. Gao, Tom Y. Ouyang, James Guo, Jiquan Ngiam, and Vijay Vasudevan. MVF: End-to-end multi-view fusion for 3d object detection in lidar point clouds. In Conference on Robot Learning (CoRL), 2019.
  65. 65.Yin Zhou and Oncel Tuzel. VoxelNet: End-to-end learning for point cloud based 3d object detection. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018.
  66. 66.Yunsong Zhou, Yuan He, Hongzi Zhu, Cheng Wang, Hongyang Li, and Qinhong Jiang. Monocular 3d object detection: An extrinsic parameter free approach. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2021.
  67. 67.Benjin Zhu, Zhengkai Jiang, Xiangxin Zhou, Zeming Li, and Gang Yu. Class-balanced grouping and sampling for point cloud 3d object detection. arXiv preprint arXiv:1908.09492, 2019.

Citation

MLA
Liang, T., et al. “BEVFusion: A Simple and Robust LiDAR-Camera Fusion Framework”. arXiv, 2022, http://arxiv.org/abs/2205.13790v3.
APA
Liang, T., Xie, H., Yu, K., Xia, Z., Lin, Z., Wang, Y., Tang, T., Wang, B., & Tang, Z. (2022). BEVFusion: A Simple and Robust LiDAR-Camera Fusion Framework. arXiv. http://arxiv.org/abs/2205.13790v3
Chicago
Liang, T., H. Xie, K. Yu, et al. 2022. “BEVFusion: A Simple and Robust LiDAR-Camera Fusion Framework”. arXiv. http://arxiv.org/abs/2205.13790v3.
Harvard
Liang, T. et al. (2022) “BEVFusion: A Simple and Robust LiDAR-Camera Fusion Framework”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2205.13790v3.
Vancouver
1. Liang T, Xie H, Yu K, Xia Z, Lin Z, Wang Y, Tang T, Wang B, Tang Z (2022) BEVFusion: A Simple and Robust LiDAR-Camera Fusion Framework. arXiv

BibTeX

@article{liang2022bevfusion,
  title = {BEVFusion: A Simple and Robust LiDAR-Camera Fusion Framework},
  author = {Liang, Tingting and Xie, Hongwei and Yu, Kaicheng and Xia, Zhongyu and Lin, Zhiwei and Wang, Yongtao and Tang, Tao and Wang, Bing and Tang, Zhi},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2205.13790v3},
  eprint = {2205.13790}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors