TransFusion: Robust LiDAR-Camera Fusion for 3D Object Detection with Transformers

Xuyang BaiZeyu HuXinge ZhuQingqiu HuangYilun ChenHongbo FuChiew-Lan Tai

article2022CVPR900 citations

Introduces TransFusion, a transformer-based 3D object detection framework that replaces rigid point-pixel calibrations with an adaptive soft-association mechanism to maintain state-of-the-art accuracy under poor illumination and sensor misalignment.

Listen

Reliable three-dimensional object detection is essential for the safe navigation of autonomous vehicles. While combining laser-based distance sensors (LiDAR) with visual cameras provides richer environmental understanding than using either sensor alone, existing fusion techniques remain fragile. Most current approaches rigidly match laser data points directly to image pixels using fixed mathematical calibration. When vehicles encounter poor lighting, missing camera frames, or minor mechanical vibrations that misalign the sensors, detection accuracy deteriorates significantly, posing safety risks for autonomous driving.

The article develops and evaluates a robust sensor fusion framework, named TransFusion, designed to maintain high 3D object detection accuracy under challenging conditions such as sensor misalignment and degraded image quality.

The authors implemented a two-stage computer vision architecture utilizing attention mechanisms. Instead of rigidly locking laser points to specific camera pixels, the system first generates preliminary 3D object locations using LiDAR data. A subsequent stage adaptively queries surrounding camera image regions to dynamically extract useful visual context, such as color and texture. The framework was evaluated across standard, large-scale autonomous driving benchmarks, specifically testing detection accuracy, tracking performance, and robustness against simulated camera dropouts, nighttime lighting, and physical sensor misalignments.

The evaluations yielded several notable findings. First, the proposed fusion method achieved top-ranked detection performance on the nuScenes benchmark (68.9% mean average precision and 71.7% detection score) without needing artificial post-processing filters. Second, the model demonstrated exceptional resilience to sensor degradation: when subjected to a 1-meter spatial misalignment between camera and laser sensors, the system's detection accuracy fell by only 0.49%, compared to drops of 2.33% to 2.85% in conventional rigid-association models. Third, during complete camera failure where all image inputs were dropped, the framework maintained a competitive 61.7% accuracy, whereas standard fusion models suffered severe performance drops between 17.2% and 23.8%. Finally, the system achieved substantial improvements in detecting distant objects beyond 30 meters and distinguishing visually nuanced categories, such as bicycles and construction vehicles.

These results demonstrate that soft, adaptive sensor fusion significantly enhances real-world reliability and vehicle safety without requiring complex calibration maintenance or separate fallback systems. Because the architecture naturally degrades gracefully to a pure laser-based mode when cameras fail, engineering teams can simplify system integration and reduce operational risks associated with hardware degradation or adverse weather.

Organizations developing autonomous perception systems should adopt flexible, attention-based soft association rather than rigid point-to-pixel mapping. Engineering teams can also streamline deployment by eliminating separate post-processing filtering stages and standardizing on feature extractors trained on instance segmentation. Future development should explore applying this soft-association mechanism to other perception tasks, such as 3D semantic segmentation, and optimize fusion strategies for denser point-cloud environments.

Confidence in these findings is high for sparse point-cloud environments and open benchmarks, backed by comprehensive ablation tests. However, decision-makers should note that performance gains are more modest in scenarios with naturally dense LiDAR data or coarse object categories, as observed on datasets like Waymo. Further pilot validation across specialized hardware setups and diverse operational domains is recommended before full-scale commercial deployment.

No sufficiently relevant recommendations were found.

Cover for TransFusion: Robust LiDAR-Camera Fusion for 3D Object Detection with Transformers

Abstract

LiDAR and camera are two important sensors for 3D object detection in autonomous driving. Despite the increasing popularity of sensor fusion in this field, the robustness against inferior image conditions, e.g., bad illumination and sensor misalignment, is under-explored. Existing fusion methods are easily affected by such conditions, mainly due to a hard association of LiDAR points and image pixels, established by calibration matrices. We propose TransFusion, a robust solution to LiDAR-camera fusion with a soft-association mechanism to handle inferior image conditions. Specifically, our TransFusion consists of convolutional backbones and a detection head based on a transformer decoder. The first layer of the decoder predicts initial bounding boxes from a LiDAR point cloud using a sparse set of object queries, and its second decoder layer adaptively fuses the object queries with useful image features, leveraging both spatial and contextual relationships. The attention mechanism of the transformer enables our model to adaptively determine where and what information should be taken from the image, leading to a robust and effective fusion strategy. We additionally design an image-guided query initialization strategy to deal with objects that are difficult to detect in point clouds. TransFusion achieves state-of-the-art performance on large-scale datasets. We provide extensive experiments to demonstrate its robustness against degenerated image quality and calibration errors. We also extend the proposed method to the 3D tracking task and achieve the 1st place in the leaderboard of nuScenes tracking, showing its effectiveness and generalization capability.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Methodology
  • 3.1 Preliminary: Transformer for 2D Detection
  • 3.2 Query Initialization
  • 3.3 Transformer Decoder and FFN
  • 3.4 LiDAR-Camera Fusion
  • 3.5 Label Assignment and Losses
  • 3.6 Image-Guided Query Initialization
  • 4 Implementation Details
  • 5 Experiments
  • 5.1 Main Results
  • 5.2 Robustness against Inferior Image Conditions
  • 5.3 Ablation Studies
  • 6 Conclusion
  • References
  • A Network Architectures
  • B Implementation Details
  • C Label Assignment Strategy
  • D Effect of NMS
  • E Pillar-based 3D Backbone
  • F Discussions of the 2D Network
  • G Adapt Queries at Test Time
  • H Dicussions on Waymo
  • I Qualitative Results

Knowls

  1. Knowl 1 — TransFusion uses sequential LiDAR detection followed by attentive image fusion

    model/method

    TransFusion uses convolutional backbones to encode a LiDAR point cloud as a bird’s-eye-view (BEV) feature map and camera views as image feature maps. Its detection head has two transformer decoder stages operating on a sparse set of object queries. The first stage uses LiDAR BEV features to predict initial 3D boxes. The second stage uses those queries and initial box predictions to retrieve image information through cross-attention, then predicts refined boxes and classes from the fused query features. The pipeline diagram depicts this sequence: BEV-derived queries and initial predictions guide attention to image features, while the initial LiDAR predictions remain available as a fallback for queries outside the camera field of view. Because attention softly associates each query with image features rather than requiring a fixed point-to-pixel pairing, the model can select relevant image content even when LiDAR returns are sparse.

  2. Knowl 2 — Input-dependent, category-aware queries provide strong initial box predictions

    model/method

    TransFusion initializes object queries from a class-specific center heatmap predicted from the LiDAR BEV feature map, rather than from learned, input-independent query positions. The heatmap has one channel per object category; its spatial locations and category channels form candidate queries. The detector selects the top NN candidates across categories, using local maxima that are at least as large as their eight spatial neighbors to avoid selecting many nearby candidates. Each selected candidate supplies a query position and a query feature from the BEV map. For category awareness, the model linearly projects the candidate’s one-hot class vector into the query-feature dimension and adds that embedding to the query feature. In the reported configurations, N=200N=200 for nuScenes and N=300N=300 for Waymo. The supplementary implementation disables the local-maximum test for small-object classes—pedestrians and traffic cones on nuScenes, and pedestrians and cyclists on Waymo—to avoid suppressing nearby instances.

  3. Knowl 3 — Spatially modulated cross-attention implements soft LiDAR–image association

    model/method

    For image fusion, TransFusion uses the initial 3D prediction and sensor calibration to identify the camera view containing a query, then cross-attends between that query and the corresponding image feature map. To bias attention toward the projected object without hard-cropping the image, spatially modulated cross-attention (SMCA) multiplies each query’s cross-attention map by a circular Gaussian centered at the projected 2D box center:

    Mij=exp⁡ ⁣(−(i−cx)2+(j−cy)2(σr)2).M_{ij}=\exp\!\left(-\frac{(i-c_x)^2+(j-c_y)^2}{(\sigma r)^2}\right).

    Here, ii and jj are spatial pixel indices in the image-feature attention map; (cx,cy)(c_x,c_y) is the projected center of the query’s predicted 3D box; rr is the radius of the minimum circumscribed circle around the projected 2D box corners, expressed in the same spatial coordinate system; and σ\sigma is a bandwidth parameter. The mask is applied elementwise across attention heads. It encourages attention to nearby, relevant image regions while allowing the model to adaptively use contextual image evidence, making fusion less dependent on precise calibration or uniformly reliable image features.

  4. Knowl 4 — Image guidance can add difficult-to-detect objects to query initialization

    model/method

    TransFusion supplements LiDAR-only query initialization with an image-guided strategy intended to recover objects that are difficult to detect in sparse point clouds. It first collapses each image feature map along its vertical dimension, reducing the image representation to image-column features. A cross-attention module then uses LiDAR BEV features as queries and the collapsed multiview image features as keys and values to produce a LiDAR–camera BEV feature map. The model predicts a class heatmap from this fused map and averages it with the LiDAR-only heatmap. It selects and initializes the final object queries from the averaged heatmap using the same candidate-selection strategy as for LiDAR-only initialization. The height collapse reduces computation under the authors’ assumption that camera geometry makes the relation between BEV locations and image columns tractable, and that usually at most one object lies along an image column; the resulting image guidance is intended as a hint about potential object locations rather than a lossless image-to-BEV representation.

  5. Knowl 5 — One-to-one matching supports end-to-end prediction without NMS

    model/method

    TransFusion assigns predicted boxes to ground-truth boxes through Hungarian bipartite matching. The matching cost is a weighted sum of classification, regression, and box-overlap costs:

    Cmatch=λ1Lcls(p,p^)+λ2Lreg(b,b^)+λ3Liou(b,b^).C_{\mathrm{match}}=\lambda_1 L_{\mathrm{cls}}(p,\hat p)+\lambda_2 L_{\mathrm{reg}}(b,\hat b)+\lambda_3 L_{\mathrm{iou}}(b,\hat b).

    Here, pp and p^\hat p denote predicted and ground-truth class labels or scores; bb and b^\hat b denote predicted and ground-truth boxes; LclsL_{\mathrm{cls}} is binary cross-entropy; LregL_{\mathrm{reg}} is the L1L_1 distance between BEV centers normalized to [0,1][0,1]; and LiouL_{\mathrm{iou}} is an intersection-over-union loss. For nuScenes, the matching-cost weights are λ1=0.15\lambda_1=0.15, λ2=0.25\lambda_2=0.25, and λ3=0.25\lambda_3=0.25. After matching, classification uses focal loss, box regression uses L1L_1 loss on matched positive pairs, and center-heatmap prediction uses a penalty-reduced focal loss. The same assignment and loss formulation is applied at both decoder stages. The box-prediction head decodes each query into center offsets, height, log dimensions, sine and cosine of yaw, optional planar velocity, and per-class scores. The model uses all predictions without non-maximum suppression (NMS): on the nuScenes validation set, mAP with versus without NMS was 65.58 versus 65.60 for TransFusion and 59.95 versus 59.98 for TransFusion-L, while CenterPoint’s mAP fell from 57.41 to 45.70.

  6. Knowl 6 — TransFusion improves nuScenes test-set detection over prior reported methods

    data/table

    On the nuScenes test set, TransFusion reports 68.9 mAP and 71.7 NDS, compared with 65.5 mAP and 70.2 NDS for its LiDAR-only variant, TransFusion-L. The fusion model therefore improves on its LiDAR-only counterpart by 3.4 mAP and 1.5 NDS. It also exceeds the listed prior fusion results from PointAugmenting and FusionPainting. The experiment uses LiDAR voxel size (0.075,0.075,0.2)(0.075,0.075,0.2) meters; the nuScenes image backbone is a pretrained, frozen DLA34 and the image resolution is 448×800448\times800. Training is staged: the LiDAR backbone and first decoder stage are trained for 20 epochs, followed by 6 epochs of LiDAR–camera fusion and image-guided initialization. The authors report no test-time augmentation or model ensemble for TransFusion; CenterPoint and PointAugmenting use double-flip testing in the comparison.

    Method mAP NDS
    CenterPoint (LiDAR) 60.3 67.3
    PointAugmenting (LiDAR + camera) 66.8 71.0
    FusionPainting (LiDAR + camera) 68.1 71.6
    TransFusion-L (LiDAR) 65.5 70.2
    TransFusion (LiDAR + camera) 68.9 71.7
  7. Knowl 7 — Ablations support the query, fusion, and image-guidance components

    empirical result

    On the nuScenes validation set, the proposed input-dependent, category-aware initialization achieved 60.0 mAP and 66.8 NDS with one decoder layer after 12 epochs. Removing category embeddings reduced the result to 54.3 mAP and 63.9 NDS. Replacing input-dependent positions with learned query positions yielded 24.0 mAP and 33.8 NDS with one layer; increasing this variant to three layers and 36 training epochs raised it to 46.9 mAP and 57.8 NDS, still below the proposed one-layer initialization. With the proposed initialization, using three layers yielded 59.9 mAP and 67.1 NDS.

    The fusion-component ablation further reports 60.0 mAP and 66.8 NDS for TransFusion-L, 61.6 and 67.4 without image-feature fusion, 64.8 and 69.3 without image-guided initialization, and 65.6 and 69.7 for the full model. The full model’s advantage over TransFusion-L was larger for distant objects: mAP rose from 70.4 to 75.5 below 15 m, from 59.5 to 66.9 at 15–30 m, and from 35.3 to 43.7 beyond 30 m. These validation results show that the initialization and image components contribute to performance, with the image-fusion benefit especially evident at longer ranges.

  8. Knowl 8 — Adaptive fusion retains accuracy under degraded or missing images

    empirical result

    The authors tested robustness on the nuScenes validation set using shortened training (12 first-stage epochs and no fade strategy) and two comparison fusions added to TransFusion-L: pointwise concatenation (CC) and PointAugmenting-style fusion (PA). For nighttime scenes, TransFusion improved mAP from the LiDAR-only baseline’s 49.2 to 55.2, a gain of 6.0; on daytime scenes it rose from 60.3 to 65.7, a gain of 5.4. The corresponding gains for CC were 0.2 at night and 3.1 by day, and for PA were 1.8 and 4.0.

    The authors also set randomly selected image features to zero at inference to simulate unavailable camera views. With zero, one, three, or all six nuScenes images dropped, mAP was 65.6, 65.1, 63.9, and 61.7 for TransFusion; 64.2, 61.6, 55.4, and 47.0 for PA; and 63.3, 59.8, 50.9, and 39.5 for CC. Thus, when all six images were unavailable, TransFusion lost 3.9 mAP, compared with losses of 17.2 for PA and 23.8 for CC. The authors attribute this resilience to generating LiDAR-based initial predictions before image fusion and adaptively selecting image information.

  9. Knowl 9 — Soft association reduces sensitivity to sensor calibration error

    empirical result

    To test calibration robustness on nuScenes, the authors randomly translated the camera-to-LiDAR transformation, creating sensor discrepancies up to 1 m. At a 1 m offset, TransFusion’s mAP decreased by 0.49 percentage points, compared with decreases of 2.33 points for PA and 2.85 points for CC. In TransFusion, calibration is used to project an object query onto an image to identify a relevant view and region; the attention mechanism can then adaptively locate useful image features around that projected position. The experiment supports the paper’s claim that soft association is less sensitive to inaccurate projected locations than hard point-to-pixel fusion.

  10. Knowl 10 — Evaluation shows transfer to Waymo detection and nuScenes tracking

    empirical result

    On Waymo validation, the reported LEVEL 2 mAPH was 65.5 overall for TransFusion and 64.9 for TransFusion-L; the class results for TransFusion were 65.1 for vehicles, 64.0 for pedestrians, and 67.4 for cyclists. CenterPoint’s reported overall result was 65.3, while PointAugmenting’s was 66.7. The authors describe TransFusion as competitive on Waymo and suggest, as possible explanations for its smaller fusion gain than on nuScenes, that Waymo has coarser categories and denser LiDAR, leaving less room for image fusion to help classification or localization.

    For nuScenes test-set tracking, the authors applied tracking-by-detection using the tracking algorithms adopted by CenterPoint. TransFusion achieved 71.8 AMOTA, with 96,775 true positives, 16,232 false positives, 21,846 false negatives, and 944 identity switches. This exceeded the listed AlphaTrack AMOTA of 69.3 and CenterPoint AMOTA of 63.8; the paper reports that TransFusion ranked first on the nuScenes tracking leaderboard at the time.

Coverage note — Detailed secondary studies of alternative 3D and 2D backbones, query-count sweeps, matching-cost sensitivity, and qualitative visualizations are omitted because they support implementation choices or illustrate behavior rather than adding a separate load-bearing contribution.

References

  1. 1.Holger Caesar, Varun Bankiti, Alex H. Lang, Sourabh Vora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Giancarlo Baldan, and Oscar Beijbom. nuScenes: A multi-modal dataset for autonomous driving. CVPR, 2020. 1
  2. 2.Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. DETR: End-to-end object detection with transformers. ECCV, 2020. 2, 3
  3. 3.Qi Chen, Lin Sun, Ernest C. H. Cheung, and A. Yuille. Every View Counts: Cross-view consistency in 3d object detection with hybrid-cylindrical-spherical voxelization. NeurIPS, 2020. 2
  4. 4.Qi Chen, Lin Sun, Zhixin Wang, K. Jia, and A. Yuille. Object as Hotspots: An anchor-free 3d object detection approach via firing of hotspots. ECCV, 2020. 2
  5. 5.Xiaozhi Chen, Huimin Ma, Jixiang Wan, B. Li, and Tian Xia. Multi-view 3d object detection network for autonomous driving. CVPR, 2017. 1, 2
  6. 6.MMDetection3D Contributors. MMDetection3D: OpenMMLab next-generation platform for general 3D object detection. https://github.com/open-mmlab/mmdetection3d, 2020. 5, 2
  7. 7.M. Everingham, L. Gool, Christopher K. I. Williams, J. Winn, and Andrew Zisserman. The pascal visual object classes (voc) challenge. IJCV, 2009. 6
  8. 8.Lue Fan, Xuan Xiong, Feng Wang, Naiyan Wang, and Zhaoxiang Zhang. RangeDet: In defense of range view for lidar-based 3d object detection. ICCV, 2021. 2
  9. 9.Peng Gao, Minghang Zheng, Xiaogang Wang, Jifeng Dai, and Hongsheng Li. Fast convergence of detr with spatially modulated co-attention. ICCV, 2021. 3, 4
  10. 10.Tengteng Huang, Zhe Liu, Xiwu Chen, and X. Bai. EPNet: Enhancing point features with image semantics for 3d object detection. ECCV, 2020. 1, 2
  11. 11.Aleksandr Kim, Aljosa Osep, and Laura Leal-Taixe. EagerMOT: 3d multi-object tracking via sensor fusion. ICRA, 2021. 7
  12. 12.Jason Ku, Melissa Mozifian, Jungwook Lee, Ali Harakeh, and Steven L. Waslander. Joint 3d proposal generation and object detection from view aggregation. IROS, 2018. 1
  13. 13.H. Kuhn. The hungarian method for the assignment problem. Naval Research Logistics Quarterly, 1955. 5
  14. 14.Alex H. Lang, Sourabh Vora, Holger Caesar, Lubing Zhou, Jiong Yang, and Oscar Beijbom. PointPillars: Fast encoders for object detection from point clouds. CVPR, 2019. 1, 2, 6
  15. 15.Zhichao Li, Feng Wang, and Naiyan Wang. LiDAR R-CNN: An efficient and universal 3d object detector. CVPR, 2021. 7
  16. 16.Ming Liang, Binh Yang, Yun Chen, Rui Hu, and R. Urtasun. Multi-task multi-sensor fusion for 3d object detection. CVPR, 2019. 1
  17. 17.Ming Liang, Binh Yang, Shenlong Wang, and R. Urtasun. Deep continuous fusion for multi-sensor 3d object detection. ECCV, 2018. 1
  18. 18.Tsung-Yi Lin, Priya Goyal, Ross B. Girshick, Kaiming He, and Piotr Dollar. Focal loss for dense object detection. ICCV, 2017. 5
  19. 19.Zili Liu, Guodong Xu, Honghui Yang, Minghao Chen, Kuoliang Wu, Zheng Yang, Haifeng Liu, and Deng Cai. Suppress-and-refine framework for end-to-end 3d object detection. arXiv, 2021. 2
  20. 20.Ze Liu, Zheng Zhang, Yue Cao, Han Hu, and Xin Tong. Group-free 3d object detection via transformers. ICCV, 2021. 2, 3
  21. 21.Jiageng Mao, Yujing Xue, Minzhe Niu, Haoyue Bai, Jiashi Feng, Xiaodan Liang, Hang Xu, and Chunjing Xu. Voxel transformer for 3d object detection. ICCV, 2021. 2
  22. 22.Gregory P. Meyer, Jake Charland, Darshan Hegde, Ankita Gajanan Laddha, and Carlos Vallespi-Gonzalez. Sensor fusion for joint 3d object detection and semantic segmentation. CVPRW, 2019. 1
  23. 23.Ishan Misra, Rohit Girdhar, and Armand Joulin. An End-to-End Transformer Model for 3D Object Detection. ICCV, 2021. 2, 3, 4, 5
  24. 24.Xuran Pan, Zhuofan Xia, Shiji Song, L. Li, and Gao Huang. 3d object detection with pointformer. CVPR, 2021. 2
  25. 25.Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer. Automatic differentiation in pytorch. NeurIPS-W, 2017. 5
  26. 26.Jonah Philion and S. Fidler. Lift, Splat, Shoot: Encoding images from arbitrary camera rigs by implicitly unprojecting to 3d. ECCV, 2020. 5
  27. 27.C. Qi, Xinlei Chen, O. Litany, and L. Guibas. ImVoteNet: Boosting 3d object detection in point clouds with image votes. CVPR, 2020. 2
  28. 28.C. Qi, O. Litany, Kaiming He, and L. Guibas. Deep hough voting for 3d object detection in point clouds. ICCV, 2019. 2
  29. 29.C. Qi, W. Liu, Chenxia Wu, Hao Su, and L. Guibas. Frustum pointnets for 3d object detection from rgb-d data. CVPR, 2018. 1, 2
  30. 30.C. Qi, Hao Su, Kaichun Mo, and L. Guibas. PointNet: Deep learning on point sets for 3d classification and segmentation. CVPR, 2017. 1
  31. 31.Shaoqing Ren, Kaiming He, Ross B. Girshick, and J. Sun. Faster R-CNN: Towards real-time object detection with region proposal networks. TPAMI, 2015. 1, 2
  32. 32.Thomas Roddick and R. Cipolla. Predicting semantic map representations from images using pyramid occupancy networks. CVPR, 2020. 5
  33. 33.Thomas Roddick, Alex Kendall, and R. Cipolla. Orthographic feature transform for monocular 3d object detection. BMVC, 2019. 5
  34. 34.Paul-Edouard Sarlin, Daniel DeTone, Tomasz Malisiewicz, and Andrew Rabinovich. SuperGlue: Learning feature matching with graph neural networks. CVPR, 2020. 4
  35. 35.Hualian Sheng, Sijia Cai, Yuan Liu, Bing Deng, Jianqiang Huang, Xiansheng Hua, and Min-Jian Zhao. Improving 3d object detection with channel-wise transformer. ICCV, 2021. 2
  36. 36.Shaoshuai Shi, Chaoxu Guo, Li Jiang, Zhe Wang, Jianping Shi, Xiaogang Wang, and Hongsheng Li. PV-RCNN: Point-voxel feature set abstraction for 3d object detection. CVPR, 2020. 2, 7
  37. 37.Shaoshuai Shi, Xiaogang Wang, and Hongsheng Li. PointRCNN: 3d object proposal generation and detection from point cloud. CVPR, 2019. 2
  38. 38.Shaoshuai Shi, Zhe Wang, Jianping Shi, Xiaogang Wang, and Hongsheng Li. From Points to Parts: 3d object detection from point cloud with part-aware and part-aggregation network. TPAMI, 2021. 2
  39. 39.Kiwoo Shin, Y. Kwon, and M. Tomizuka. RoarNet: A robust 3d object detection based on region approximation refinement. IV, 2019. 1, 2, 7
  40. 40.Vishwanath A. Sindagi, Yin Zhou, and Oncel Tuzel. MVX-Net: Multimodal voxelnet for 3d object detection. ICRA, 2019. 1, 2
  41. 41.Jiaming Sun, Zehong Shen, Yuang Wang, Hujun Bao, and Xiaowei Zhou. LoFTR: Detector-free local feature matching with transformers. CVPR, 2021. 4
  42. 42.Peize Sun, Yi Jiang, Enze Xie, Wenqi Shao, Zehuan Yuan, Changhu Wang, and Ping Luo. OneNet: What makes for end-to-end object detection? In ICML, 2021. 3
  43. 43.Pei Sun, Henrik Kretzschmar, Xerxes Dotiwalla, Aurelien Chouard, Vijaysai Patnaik, P. Tsui, James Guo, Yin Zhou, Yuning Chai, Benjamin Caine, Vijay Vasudevan, Wei Han, Jiquan Ngiam, Hang Zhao, Aleksei Timofeev, S. Ettinger, Maxim Krivokon, A. Gao, Aditya Joshi, Y. Zhang, Jonathon Shlens, Zhifeng Chen, and Dragomir Anguelov. Scalability in perception for autonomous driving: Waymo open dataset. CVPR, 2020. 1
  44. 44.Pei Sun, Weiyue Wang, Yuning Chai, Gamaleldin F. Elsayed, Alex Bewley, Xiao Zhang, Cristian Sminchisescu, and Drago Anguelov. RSN: Range sparse net for efficient, accurate lidar 3d object detection. CVPR, 2021. 2
  45. 45.Pei Sun, Rufeng Zhang, Yi Jiang, T. Kong, Chenfeng Xu, W. Zhan, M. Tomizuka, L. Li, Zehuan Yuan, C. Wang, and Ping Luo. Sparse R-CNN: End-to-end object detection with learnable proposals. CVPR, 2021. 2, 3
  46. 46.Ashish Vaswani, Noam M. Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. NeurIPS, 2017. 3, 2
  47. 47.Sourabh Vora, Alex H. Lang, Bassam Helou, and Oscar Beijbom. PointPainting: Sequential fusion for 3d object detection. CVPR, 2020. 1, 2, 4, 6, 7, 8
  48. 48.Chunwei Wang, Chao Ma, Ming Zhu, and Xiaokang Yang. PointAugmenting: Cross-modal augmentation for 3d object detection. CVPR, 2021. 1, 2, 4, 5, 6, 7
  49. 49.Yue Wang, Alireza Fathi, Abhijit Kundu, David A. Ross, Caroline Pantofaru, Thomas A. Funkhouser, and Justin M. Solomon. Pillar-based object detection for autonomous driving. In ECCV, 2020. 2
  50. 50.Yue Wang and Justin Solomon. Object DGCNN: 3d object detection using dynamic graphs. NeurIPS, 2021. 3
  51. 51.Liang Xie, Chao Xiang, Zhengxu Yu, Guodong Xu, Zheng Yang, Deng Cai, and Xiaofei He. PI-RCNN: An efficient multi-sensor 3d object detector with point-based attentive cont-conv fusion module. AAAI, 2020. 1
  52. 52.Shaoqing Xu, Dingfu Zhou, Jin Fang, Junbo Yin, Bin Zhou, and Liangjun Zhang. FusionPainting: Multimodal fusion with adaptive attention for 3d object detection. ITSC, 2021. 1, 6
  53. 53.Yan Yan, Yuxing Mao, and B. Li. SECOND: Sparsely embedded convolutional detection. Sensors, 2018. 2, 5
  54. 54.Binh Yang, Wenjie Luo, and R. Urtasun. PIXOR: Real-time 3d object detection from point clouds. CVPR, 2018. 2
  55. 55.Zetong Yang, Y. Sun, Shu Liu, and Jiaya Jia. 3DSSD: Point-based 3d single stage object detector. CVPR, 2020. 2
  56. 56.Zetong Yang, Y. Sun, Shu Liu, Xiaoyong Shen, and Jiaya Jia. STD: Sparse-to-dense 3d object detector for point cloud. ICCV, 2019. 2
  57. 57.Z. Yao, Jiangbo Ai, Boxun Li, and Chi Zhang. Efficient DETR: Improving end-to-end object detector with dense prior. arXiv, 2021. 3
  58. 58.Tianwei Yin, Xingyi Zhou, and Philipp Krahenbühl. Center-based 3d object detection and tracking. CVPR, 2021. 2, 4, 5, 6, 7
  59. 59.Tianwei Yin, Xingyi Zhou, and Philipp Krahenbühl. Multi-modal virtual point 3d detection. NeurIPS, 2021. 6
  60. 60.Jin Hyeok Yoo, Yeocheol Kim, Ji Song Kim, and J. Choi. 3D-CVF: Generating joint camera and lidar features using cross-view spatial feature fusion for 3d object detection. ECCV, 2020. 1, 6
  61. 61.F. Yu, Dequan Wang, and Trevor Darrell. Deep layer aggregation. CVPR, 2018. 5
  62. 62.Yihan Zeng, Chao Ma, Ming Zhu, Zhiming Fan, and Xiaokang Yang. Cross-modal 3d object detection and tracking for auto-driving. IROS, 2021. 7
  63. 63.Wenwei Zhang, Zhe Wang, and Chen Change Loy. Multi-modality cut and paste for 3d object detection. arXiv, 2020. 1
  64. 64.Lin Zhao, Hui Zhou, Xinge Zhu, Xiao Song, Hongsheng Li, and Wenbing Tao. LIF-Seg: Lidar and camera image fusion for 3d lidar semantic segmentation. arXiv, 2021. 1
  65. 65.Dingfu Zhou, Jin Fang, Xibin Song, Chenye Guan, Junbo Yin, Yuchao Dai, and Ruigang Yang. Iou loss for 2d/3d object detection. 3DV, 2019. 5
  66. 66.Xingyi Zhou, Dequan Wang, and Philipp Krahenbühl. Objects as points. arXiv, 2019. 3, 4
  67. 67.Yin Zhou, Pei Sun, Y. Zhang, Dragomir Anguelov, J. Gao, Tom Y. Ouyang, James Guo, Jiquan Ngiam, and Vijay Vasudevan. MVF: End-to-end multi-view fusion for 3d object detection in lidar point clouds. CoRL, 2019. 2
  68. 68.Yin Zhou and Oncel Tuzel. VoxelNet: End-to-end learning for point cloud based 3d object detection. CVPR, 2018. 1, 2, 5
  69. 69.Benjin Zhu, Zhengkai Jiang, Xiangxin Zhou, Zeming Li, and Gang Yu. Class-balanced grouping and sampling for point cloud 3d object detection. arXiv, 2019. 2, 5, 6
  70. 70.Xinge Zhu, Yuexin Ma, Tai Wang, Yan Xu, Jianping Shi, and Dahua Lin. SSN: Shape signature networks for multi-class object detection from point clouds. ECCV, 2020. 2
  71. 71.Xizhou Zhu, Weijie Su, Lewei Lu, Bin Li, Xiaogang Wang, and Jifeng Dai. Deformable DETR: Deformable transformers for end-to-end object detection. ICLR, 2021. 3, 2

Citation

MLA
Bai, X., et al. “TransFusion: Robust LiDAR-Camera Fusion for 3D Object Detection with Transformers”. arXiv, 2022, http://arxiv.org/abs/2203.11496v1.
APA
Bai, X., Hu, Z., Zhu, X., Huang, Q., Chen, Y., Fu, H., & Tai, C.-L. (2022). TransFusion: Robust LiDAR-Camera Fusion for 3D Object Detection with Transformers. arXiv. http://arxiv.org/abs/2203.11496v1
Chicago
Bai, X., Z. Hu, X. Zhu, et al. 2022. “TransFusion: Robust LiDAR-Camera Fusion for 3D Object Detection with Transformers”. arXiv. http://arxiv.org/abs/2203.11496v1.
Harvard
Bai, X. et al. (2022) “TransFusion: Robust LiDAR-Camera Fusion for 3D Object Detection with Transformers”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2203.11496v1.
Vancouver
1. Bai X, Hu Z, Zhu X, Huang Q, Chen Y, Fu H, Tai C-L (2022) TransFusion: Robust LiDAR-Camera Fusion for 3D Object Detection with Transformers. arXiv

BibTeX

@article{bai2022transfusion,
  title = {TransFusion: Robust LiDAR-Camera Fusion for 3D Object Detection with Transformers},
  author = {Bai, Xuyang and Hu, Zeyu and Zhu, Xinge and Huang, Qingqiu and Chen, Yilun and Fu, Hongbo and Tai, Chiew-Lan},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2203.11496v1},
  eprint = {2203.11496}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE