Voxel Field Fusion for 3D Object Detection

Yanwei LiXiaojuan QiYukang ChenLiwei WangZeming LiJian SunJiaya Jia

article2022CVPR117 citations

Proposes a cross-modality 3D object detection framework that maintains sensor consistency by projecting augmented camera features as rays into a voxel field with learnable sampling, achieving state-of-the-art results on the KITTI and nuScenes benchmarks.

Listen

Reliable three-dimensional object detection is vital for safety-critical systems such as autonomous vehicles. While laser-based distance sensors, known as LiDAR, offer precise geometry, their data becomes sparse over long distances or in occluded areas. Integrating camera images helps compensate for this deficiency by adding rich visual context. However, existing multi-modal systems suffer from representation gaps when projecting two-dimensional camera features into three-dimensional space and encounter misalignment issues during data augmentation training.

The article develops and evaluates Voxel Field Fusion, an integrated framework designed to maintain cross-modality consistency by projecting camera image features as continuous rays into a three-dimensional voxel grid. The authors assess this framework by integrating it with multiple standard detection networks across two widely recognized autonomous driving benchmarks: the KITTI and nuScenes datasets.

To overcome computational limits and data misalignment, the approach introduces three synchronized components. A mixed augmentor aligns data-level transformations across camera and LiDAR inputs. An intelligent sampler then selects high-importance image regions instead of processing entire frames, and a ray-wise fusion mechanism evaluates voxels along projection rays to populate spatial features into both occupied and empty voxels based on predicted probabilities.

The experimental findings show that the proposed framework delivers consistent performance improvements over existing baselines. First, the framework achieved leading results on the nuScenes test benchmark, reaching 68.4% mean Average Precision and a 72.4% NuScenes Detection Score, outperforming its base detector by 8.1% and 5.1% respectively. Second, the system demonstrated significant improvements in identifying difficult objects, yielding up to a 19% gain for ambiguous categories like motorcycles and bicycles on nuScenes, and a 6% boost for pedestrian detection on KITTI. Third, on the KITTI test set, the framework reached 79.29% accuracy on hard car detection cases, outperforming baseline models by 2.2%. Finally, in sparse data tests using a reduced 32-beam LiDAR setup, the method delivered a 2.85% overall improvement and a 3.27% gain on hard cases, demonstrating that ray-wise visual completion successfully compensates for missing sensor points.

These results indicate that ray-based fusion improves detection reliability without requiring specialized, high-cost sensors. By accurately identifying distant, occluded, and vulnerable road users, the framework reduces the risk of missed detections in safety-critical automated driving pipelines.

Engineering teams should consider adopting this ray-wise voxel fusion framework to upgrade existing 3D perception backbones. Technical leaders should also implement synchronized cross-modality data augmentation pipelines to prevent model degradation. Next steps include validating the framework on larger internal datasets and conducting runtime profiling to ensure ray construction meets real-time latency budgets on embedded vehicle hardware.

While confidence in the reported detection accuracy is high across the tested public datasets, the framework requires camera and LiDAR calibration parameters for projection mapping. Practitioners should account for hardware latency constraints when deploying the learnable sampler and ray-fusion modules in production environments.

Cover for Voxel Field Fusion for 3D Object Detection

Abstract

In this work, we present a conceptually simple yet effective framework for cross-modality 3D object detection, named voxel field fusion. The proposed approach aims to maintain cross-modality consistency by representing and fusing augmented image features as a ray in the voxel field. To this end, the learnable sampler is first designed to sample vital features from the image plane that are projected to the voxel grid in a point-to-ray manner, which maintains the consistency in feature representation with spatial context. In addition, ray-wise fusion is conducted to fuse features with the supplemental context in the constructed voxel field. We further develop mixed augmentor to align feature-variant transformations, which bridges the modality gap in data augmentation. The proposed framework is demonstrated to achieve consistent gains in various benchmarks and outperforms previous fusion-based methods on KITTI and nuScenes datasets. Code is made available at https://github.com/dvlab-research/VFF.1

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Voxel Field Fusion
  • 3.1. Mixed Augmentor
  • 3.2. Voxel Field Construction
  • 3.3. Ray-voxel Interaction
  • 3.4. Optimization Objectives
  • 4. Experiments
  • 4.1. Experimental Setup
  • 4.2. Component-wise Analysis
  • 4.3. Main Results
  • 5. Conclusion
  • 6. Acknowledgments
  • References

Knowls

  1. Knowl 1 — Voxel field fusion framework

    model/method

    Voxel field fusion (VFF) is an end-to-end cross-modality 3D object-detection framework for synchronized camera images and LiDAR point clouds. During training, a mixed augmentor applies corresponding transformations to both modalities. Separate feature encoders then produce image and voxel features, a projection matrix establishes image–voxel correspondence, and VFF represents each selected image feature as a ray through the 3D voxel space rather than attaching it to a single LiDAR point. A learnable sampler selects image features in high-response regions, and ray-wise fusion injects those features into the voxel bins along each ray. The resulting voxel features are passed to a voxel-based detection backbone and classification, regression, and direction heads. The page-3 pipeline diagram depicts this sequence from camera/LiDAR input through augmentation, feature encoding, ray construction, sampling, ray-wise fusion, and detection. VFF can be instantiated with PV-RCNN, Voxel R-CNN, or CenterPoint backbones.

  2. Knowl 2 — Ray-wise image–voxel fusion

    model/method

    For a sampled image pixel pip_i and its voxel ray RiR_i, VFF predicts how strongly the image feature belongs to every voxel along the ray, allowing image context to complete sparse or empty LiDAR regions. Let Fl,iIF^I_{l,i} be the image feature at encoder stage ll, let vjv_j be a voxel at coordinates (xj,yj,zj)(x_j,y_j,z_j) on RiR_i, and let ϕl(xj,yj,zj)=MLP⁡l([xj,yj,zj])\phi_l(x_j,y_j,z_j)=\operatorname{MLP}_l([x_j,y_j,z_j]) be a learned positional feature. The voxel response is

    ωj=sigmoid⁡ ⁣(⟨Fl,iI,ϕl(xj,yj,zj)⟩).\omega_j=\operatorname{sigmoid}\!\left(\left\langle F^I_{l,i},\phi_l(x_j,y_j,z_j)\right\rangle\right).

    The fused voxel feature is

    F~lV(xj,yj,zj)=FlV(xj,yj,zj)+ωj gl ⁣([Fl,iI,ϕl(xj,yj,zj)]),\widetilde F^V_l(x_j,y_j,z_j)=F^V_l(x_j,y_j,z_j)+\omega_j\,g_l\!\left([F^I_{l,i},\phi_l(x_j,y_j,z_j)]\right),

    where FlVF^V_l is the original voxel feature and glg_l is a learned convolution or feature-mixing function. An empty voxel has FlV=0F^V_l=0, so the operation also performs feature completion. Unlike single-point fusion or local fusion inside a fixed-radius neighborhood, ray-wise fusion evaluates the entire ray and uses the predicted response to select only the highest-scoring voxels; the implementation retains a number equal to one quarter of the original non-empty voxels. During inference, ray features are selected when ωj>0.05\omega_j>0.05.

  3. Knowl 3 — Mixed augmentor for cross-modality alignment

    model/method

    The mixed augmentor aligns camera and LiDAR transformations during training through two augmentation groups. Sample-added augmentation uses 3D ground-truth sampling for LiDAR and copy-paste augmentation for RGB: each inserted 3D object is cropped within its projected 2D bounding box and pasted into the image in depth or cropping order. Points occluded by nearer inserted objects are removed to avoid ambiguous correspondences. Sample-static augmentation applies paired operations without adding objects: LiDAR flipping is paired with image flipping, LiDAR rescaling with image rescaling, and LiDAR rotation with image reprojection. Image-level flipping and rescaling are important because asynchronous transformations otherwise alter the spatial context processed by pretrained 2D convolutions. This design maintains correspondence directly at the feature-variant image level instead of modifying only the LiDAR coordinates after image features have been computed.

  4. Knowl 4 — Voxel-field and point-to-ray construction

    definition

    A voxel field is a function over a 3D voxel space VV. For a voxel bin vv with center coordinates (x,y,z)(x,y,z) and image-feature stage ll, the voxel representation is written as

    Fl,vV=Fl(x,y,z).F^V_{l,v}=F_l(x,y,z).

    Let pip_i be the homogeneous image coordinate of pixel ii, let vjv_j be the homogeneous coordinate of voxel-bin center jj, and let TVoxel→ImageT_{\mathrm{Voxel}\to\mathrm{Image}} be the calibrated voxel-to-image projection matrix. The ray associated with pip_i is the set

    Ri={vj∈V  |  pi=vjTTVoxel→Image}.R_i=\left\{v_j\in V\;\middle|\;p_i=v_j^{\mathsf T}T_{\mathrm{Voxel}\to\mathrm{Image}}\right\}.

    Thus, every voxel bin that projects to the same image pixel belongs to the same ray. Without sampling, as many as W×HW\times H rays could be constructed for an image of width WW pixels and height HH pixels, making dense image-to-voxel rendering expensive. The point-to-ray representation preserves neighboring image context while maintaining the calibrated spatial relation between the image and voxel modalities; the page-4 ray diagrams visualize this extension from a single projected point to all aligned voxels along depth.

  5. Knowl 5 — Learnable importance sampler

    algorithm

    VFF constructs only a limited number of rays by learning which image features are useful for 3D detection.

    Input: image features FlIF^I_l, projected image pixels PP, window size w=64w=64, and ray budget nn. Output: sampled image pixels P^\widehat P and their rays.

    Partition the image plane into non-overlapping windows of size w by w.
    Discard windows containing no projected LiDAR pixel.
    For every remaining candidate pixel p_i, compute a response a_i with stacked convolutions and a sigmoid.
    Keep pixels whose response exceeds 0.5.
    Uniformly sample at most n pixels from the retained high-response pixels.
    Construct one voxel ray for each sampled pixel.
    Return the sampled pixels and their rays.

    In equation form, with fsf_s denoting the learned stacked convolutions, δ\delta the sigmoid, 1[⋅]\mathbf{1}[\cdot] an indicator, and Un\mathcal U_n uniform sampling with budget nn,

    P^=Un ⁣({pi∈P:1[δ ⁣(fs(Fl,iI))>0.5]=1}).\widehat P=\mathcal U_n\!\left(\left\{p_i\in P:\mathbf{1}\left[\delta\!\left(f_s(F^I_{l,i})\right)>0.5\right]=1\right\}\right).

    The sampler is supervised to emphasize foreground regions, so it avoids the useless projected pixels retained by uniformity-, density-, or sparsity-based heuristics while constructing fewer rays.

  6. Knowl 6 — Supervision for sampling and ray responses

    equation

    VFF supervises its two learned selection mechanisms with foreground-centered Gaussian targets. For an image pixel with coordinates (u,v)(u,v), the target for the learnable sampler is

    Yl,u,v=exp⁡ ⁣(−(u−u^i)2+(v−v^i)22σi2),Y_{l,u,v}=\exp\!\left(-\frac{(u-\hat u_i)^2+(v-\hat v_i)^2}{2\sigma_i^2}\right),

    where (u^i,v^i)(\hat u_i,\hat v_i) is the center of object ii's 2D bounding box and σi\sigma_i is an object-size-adaptive standard deviation. For a voxel vjv_j on a ray, let (x^j,y^j,z^j)(\hat x_j,\hat y_j,\hat z_j) be the coordinates of its LiDAR-containing anchor voxel and let σj\sigma_j be a size-adaptive standard deviation. The ray-response target is

    ω^j=exp⁡ ⁣(−(x−x^j)2+(y−y^j)2+(z−z^j)22σj2),\widehat\omega_j=\exp\!\left(-\frac{(x-\hat x_j)^2+(y-\hat y_j)^2+(z-\hat z_j)^2}{2\sigma_j^2}\right),

    for voxels within Euclidean radius rr of the anchor; voxels farther than rr receive target 00. The VFF auxiliary loss is

    LVFF=λs BCE⁡ ⁣(fs(FlI),Yl)+λrm∑i=1mFL⁡(ωi,ω^i),\mathcal L_{\mathrm{VFF}}=\lambda_s\,\operatorname{BCE}\!\left(f_s(F^I_l),Y_l\right)+\frac{\lambda_r}{m}\sum_{i=1}^{m}\operatorname{FL}(\omega_i,\widehat\omega_i),

    where mm is the number of sampled rays, BCE⁡\operatorname{BCE} is binary cross-entropy, FL⁡\operatorname{FL} is focal loss, and λs=2\lambda_s=2 and λr=5\lambda_r=5 in all experiments. The complete training objective is the raw detector loss Ldet\mathcal L_{\mathrm{det}} plus LVFF\mathcal L_{\mathrm{VFF}}.

  7. Knowl 7 — KITTI and nuScenes experimental protocol

    experimental setup

    VFF was evaluated on synchronized multimodal autonomous-driving data. KITTI contains 7,481 training samples and 7,518 test samples; the usual split used here has 3,712 training and 3,769 validation samples. nuScenes contains 1,000 scenes split into 700 training, 150 validation, and 150 test scenes, with 10 object categories, a 32-beam LiDAR, and six cameras covering 360 degrees. PV-RCNN and Voxel R-CNN were used as KITTI backbones, while CenterPoint was used on nuScenes. The implementation follows each backbone's standard architecture and training settings, uses three convolutions for the sampler and MLP feature transformations, and uses an individual MLP for each camera view. Fusion is inserted at stage 1 by default, and the feature-encoder stage index is l=1l=1. In inference, non-empty voxel features and ray voxels with response ω>0.05\omega>0.05 are considered. KITTI reports 3D and bird's-eye-view average precision at the stated class-specific IoU thresholds; nuScenes reports mean average precision (mAP), nuScenes detection score (NDS), and per-category average precision.

  8. Knowl 8 — KITTI validation ablations identify the effective components

    data/table

    The component ablations use the KITTI validation set, PV-RCNN, and AP⁡3D\operatorname{AP}_{3D} for cars at IoU =0.7=0.7 with the R40 protocol. They show that paired augmentation, learned importance sampling, full-ray fusion, and radius-one Gaussian supervision are the strongest choices.

    Could not parse LaTeX table
    Could not parse LaTeX table
    Could not parse LaTeX table
    Could not parse LaTeX table

    Image-level alignment is especially effective for flipping: the aligned operation reaches 84.15%84.15\% moderate AP versus 82.50%82.50\% for point-cloud reprojection and 81.82%81.82\% without static augmentation. On 32-beam downsampled LiDAR, VFF improves the KITTI moderate score from 79.51%79.51\% to 82.36%82.36\% and the hard score from 76.54%76.54\% to 79.81%79.81\%, supporting the intended completion of sparse voxel regions.

  9. Knowl 9 — KITTI detection results

    empirical result

    On the KITTI validation set, VFF consistently improves both voxel-based backbones under the car AP⁡3D\operatorname{AP}_{3D} R40 metric at IoU =0.7=0.7.

    Could not parse LaTeX table

    On the KITTI test set, VFF reaches 89.58/81.97/79.1789.58/81.97/79.17 AP for easy/moderate/hard cases with PV-RCNN and 89.50/82.09/79.2989.50/82.09/79.29 with Voxel R-CNN. The corresponding LiDAR-only backbones score 90.25/81.43/76.8290.25/81.43/76.82 and 90.90/81.62/77.0690.90/81.62/77.06, respectively. The VFF Voxel R-CNN result therefore improves the hard-case score by 2.232.23 AP over its LiDAR-only baseline, reported by the paper as a 2.2%2.2\% gain, and exceeds the earlier fusion-based EPNet and 3D-CVF test results on hard cases. On KITTI validation, the gain is particularly large for pedestrians: VFF raises easy/moderate/hard AP from 66.04/59.19/54.1566.04/59.19/54.15 to 73.26/65.11/60.0373.26/65.11/60.03.

  10. Knowl 10 — nuScenes detection results

    empirical result

    On the nuScenes test set, VFF combined with CenterPoint achieves the strongest reported multimodal result in the paper, with 68.4%68.4\% mAP and 72.4%72.4\% NDS. The per-category average precisions are 86.886.8 for Car, 58.158.1 for Truck, 70.270.2 for Bus, 61.061.0 for Trailer, 32.132.1 for Construction Vehicle, 87.187.1 for Pedestrian, 78.578.5 for Motorcycle, 52.952.9 for Bicycle, 83.883.8 for Traffic Cone, and 73.973.9 for Barrier.

    Could not parse LaTeX table

    Relative to the CenterPoint backbone, VFF adds 8.18.1 mAP points and 5.15.1 NDS points. The paper reports especially large gains for ambiguous categories such as motorcycles and bicycles, demonstrating that ray-based image context can compensate for LiDAR sparsity and class ambiguity.

Coverage note — Lower-priority analyses of fusion-stage placement, pretrained 2D task choice, and the complete set of category-specific and baseline comparison rows were omitted because the ten knowls retain the method, optimization, core ablations, experimental protocol, and principal KITTI/nuScenes results.

References

  1. 1.Garrick Brazil and Xiaoming Liu. M3d-rpn: Monocular 3d region proposal network for object detection. In ICCV, 2019.
  2. 2.Holger Caesar, Varun Bankiti, Alex H Lang, Sourabh Vora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Giancarlo Baldan, and Oscar Beijbom. nuscenes: A multi-modal dataset for autonomous driving. In CVPR, 2020.
  3. 3.Liang-Chieh Chen, George Papandreou, Florian Schroff, and Hartwig Adam. Rethinking atrous convolution for semantic image segmentation. arXiv:1706.05587, 2017.
  4. 4.Rui Chen, Songfang Han, Jing Xu, and Hao Su. Point-based multi-view stereo network. In ICCV, 2019.
  5. 5.Xiaozhi Chen, Kaustav Kundu, Yukun Zhu, Andrew G Berneshawi, Huimin Ma, Sanja Fidler, and Raquel Urtasun. 3d object proposals for accurate object class detection. In NeurIPS, 2015.
  6. 6.Xiaozhi Chen, Huimin Ma, Ji Wan, Bo Li, and Tian Xia. Multi-view 3d object detection network for autonomous driving. In CVPR, 2017.
  7. 7.Yilun Chen, Shu Liu, Xiaoyong Shen, and Jiaya Jia. Dsgn: Deep stereo geometry network for 3d object detection. In CVPR, 2020.
  8. 8.Jiajun Deng, Shaoshuai Shi, Peiwei Li, Wengang Zhou, Yanyong Zhang, and Houqiang Li. Voxel r-cnn: Towards high performance voxel-based 3d object detection. In AAAI, 2021.
  9. 9.Nikita Dvornik, Julien Mairal, and Cordelia Schmid. Modeling visual context is key to augmenting object detection datasets. In ECCV, 2018.
  10. 10.Andreas Geiger, Philip Lenz, and Raquel Urtasun. Are we ready for autonomous driving? the kitti vision benchmark suite. In CVPR, 2012.
  11. 11.Chenhang He, Hui Zeng, Jianqiang Huang, Xian-Sheng Hua, and Lei Zhang. Structure aware single-stage 3d object detection from point cloud. In CVPR, 2020.
  12. 12.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In CVPR, 2016.
  13. 13.Tengteng Huang, Zhe Liu, Xiwu Chen, and Xiang Bai. Epnet: Enhancing point features with image semantics for 3d object detection. In ECCV, 2020.
  14. 14.Jason Ku, Melissa Mozifian, Jungwook Lee, Ali Harakeh, and Steven L Waslander. Joint 3d proposal generation and object detection from view aggregation. In IROS, 2018.
  15. 15.Alex H Lang, Sourabh Vora, Holger Caesar, Lubing Zhou, Jiong Yang, and Oscar Beijbom. Pointpillars: Fast encoders for object detection from point clouds. In CVPR, 2019.
  16. 16.Ming Liang, Bin Yang, Yun Chen, Rui Hu, and Raquel Urtasun. Multi-task multi-sensor fusion for 3d object detection. In CVPR, 2019.
  17. 17.Ming Liang, Bin Yang, Shenlong Wang, and Raquel Urtasun. Deep continuous fusion for multi-sensor 3d object detection. In ECCV, 2018.
  18. 18.Tsung-Yi Lin, Priya Goyal, Ross Girshick, Kaiming He, and Piotr Dollar. Focal loss for dense object detection. In ICCV, 2017.
  19. 19.Lingjie Liu, Jiatao Gu, Kyaw Zaw Lin, Tat-Seng Chua, and Christian Theobalt. Neural sparse voxel fields. In NeurIPS, 2020.
  20. 20.Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoorthi, and Ren Ng. Nerf: Representing scenes as neural radiance fields for view synthesis. In ECCV, 2020.
  21. 21.Charles R Qi, Xinlei Chen, Or Litany, and Leonidas J Guibas. Imvotenet: Boosting 3d object detection in point clouds with image votes. In CVPR, 2020.
  22. 22.Charles R Qi, Or Litany, Kaiming He, and Leonidas J Guibas. Deep hough voting for 3d object detection in point clouds. In ICCV, 2019.
  23. 23.Charles R Qi, Wei Liu, Chenxia Wu, Hao Su, and Leonidas J Guibas. Frustum pointnets for 3d object detection from rgb-d data. In CVPR, 2018.
  24. 24.Charles R Qi, Li Yi, Hao Su, and Leonidas J Guibas. Pointnet++: Deep hierarchical feature learning on point sets in a metric space. In NeurIPS, 2017.
  25. 25.Cody Reading, Ali Harakeh, Julia Chae, and Steven L Waslander. Categorical depth distribution network for monocular 3d object detection. In CVPR, 2021.
  26. 26.Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. Faster r-cnn: towards real-time object detection with region proposal networks. TPAMI, 2016.
  27. 27.Shaoshuai Shi, Chaoxu Guo, Li Jiang, Zhe Wang, Jianping Shi, Xiaogang Wang, and Hongsheng Li. Pv-rcnn: Point-voxel feature set abstraction for 3d object detection. In CVPR, 2020.
  28. 28.Shaoshuai Shi, Xiaogang Wang, and Hongsheng Li. Pointrcnn: 3d object proposal generation and detection from point cloud. In CVPR, 2019.
  29. 29.Shaoshuai Shi, Zhe Wang, Jianping Shi, Xiaogang Wang, and Hongsheng Li. From points to parts: 3d object detection from point cloud with part-aware and part-aggregation network. TPAMI, 2020.
  30. 30.Xuepeng Shi, Zhixiang Chen, and Tae-Kyun Kim. Distance-normalized unified representation for monocular 3d object detection. In ECCV, 2020.
  31. 31.Andrea Simonelli, Samuel Rota Bulo, Lorenzo Porzi, Manuel López-Antequera, and Peter Kontschieder. Disentangling monocular 3d object detection. In ICCV, 2019.
  32. 32.Vishwanath A Sindagi, Yin Zhou, and Oncel Tuzel. Mvx-net: Multimodal voxelnet for 3d object detection. In ICRA, 2019.
  33. 33.Sourabh Vora, Alex H Lang, Bassam Helou, and Oscar Beijbom. Pointpainting: Sequential fusion for 3d object detection. In CVPR, 2020.
  34. 34.Chunwei Wang, Chao Ma, Ming Zhu, and Xiaokang Yang. Pointaugmenting: Cross-modal augmentation for 3d object detection. In CVPR, 2021.
  35. 35.Tai Wang, Xinge Zhu, Jiangmiao Pang, and Dahua Lin. Fcos3d: Fully convolutional one-stage monocular 3d object detection. arXiv:2104.10956, 2021.
  36. 36.Yan Wang, Wei-Lun Chao, Divyansh Garg, Bharath Hariharan, Mark Campbell, and Kilian Q Weinberger. Pseudo-lidar from visual depth estimation: Bridging the gap in 3d object detection for autonomous driving. In CVPR, 2019.
  37. 37.Shaoqing Xu, Dingfu Zhou, Jin Fang, Junbo Yin, Zhou Bin, and Liangjun Zhang. Fusionpainting: Multimodal fusion with adaptive attention for 3d object detection. In ITSC, 2021.
  38. 38.Yan Yan, Yuxing Mao, and Bo Li. Second: Sparsely embedded convolutional detection. Sensors, 2018.
  39. 39.Bin Yang, Ming Liang, and Raquel Urtasun. Hdnet: Exploiting hd maps for 3d object detection. In CoRL, 2018.
  40. 40.Bin Yang, Wenjie Luo, and Raquel Urtasun. Pixor: Real-time 3d object detection from point clouds. In CVPR, 2018.
  41. 41.Zetong Yang, Yanan Sun, Shu Liu, and Jiaya Jia. 3dssd: Point-based 3d single stage object detector. In CVPR, 2020.
  42. 42.Zetong Yang, Yanan Sun, Shu Liu, Xiaoyong Shen, and Jiaya Jia. Std: Sparse-to-dense 3d object detector for point cloud. In ICCV, 2019.
  43. 43.Yao Yao, Zixin Luo, Shiwei Li, Tian Fang, and Long Quan. Mvsnet: Depth inference for unstructured multi-view stereo. In ECCV, 2018.
  44. 44.Tianwei Yin, Xingyi Zhou, and Philipp Krahenbuhl. Center-based 3d object detection and tracking. In CVPR, 2021.
  45. 45.Tianwei Yin, Xingyi Zhou, and Philipp Krahenbühl. Multi-modal virtual point 3d detection. In NeurIPS, 2021.
  46. 46.Jin Hyeok Yoo, Yecheol Kim, Jisong Kim, and Jun Won Choi. 3d-cvf: Generating joint camera and lidar features using cross-view spatial feature fusion for 3d object detection. In ECCV, 2020.
  47. 47.Yurong You, Yan Wang, Wei-Lun Chao, Divyansh Garg, Geoff Pleiss, Bharath Hariharan, Mark Campbell, and Kilian Q Weinberger. Pseudo-lidar++: Accurate depth for 3d object detection in autonomous driving. In ICLR, 2020.
  48. 48.Alex Yu, Vickie Ye, Matthew Tancik, and Angjoo Kanazawa. pixelnerf: Neural radiance fields from one or few images. In CVPR, 2021.
  49. 49.Sangdoo Yun, Dongyoon Han, Seong Joon Oh, Sanghyuk Chun, Junsuk Choe, and Youngjoon Yoo. Cutmix: Regularization strategy to train strong classifiers with localizable features. In ICCV, 2019.
  50. 50.Wenwei Zhang, Zhe Wang, and Chen Change Loy. Exploring data augmentation for multi-modality 3d object detection. arXiv:2012.12741, 2020.
  51. 51.Yin Zhou and Oncel Tuzel. Voxelnet: End-to-end learning for point cloud based 3d object detection. In CVPR, 2018.
  52. 52.Benjin Zhu, Zhengkai Jiang, Xiangxin Zhou, Zeming Li, and Gang Yu. Class-balanced grouping and sampling for point cloud 3d object detection. arXiv:1908.09492, 2019.

Citation

MLA
Li, Y., et al. “Voxel Field Fusion for 3D Object Detection”. arXiv, 2022, http://arxiv.org/abs/2205.15938v1.
APA
Li, Y., Qi, X., Chen, Y., Wang, L., Li, Z., Sun, J., & Jia, J. (2022). Voxel Field Fusion for 3D Object Detection. arXiv. http://arxiv.org/abs/2205.15938v1
Chicago
Li, Y., X. Qi, Y. Chen, et al. 2022. “Voxel Field Fusion for 3D Object Detection”. arXiv. http://arxiv.org/abs/2205.15938v1.
Harvard
Li, Y. et al. (2022) “Voxel Field Fusion for 3D Object Detection”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2205.15938v1.
Vancouver
1. Li Y, Qi X, Chen Y, Wang L, Li Z, Sun J, Jia J (2022) Voxel Field Fusion for 3D Object Detection. arXiv

BibTeX

@article{li2022voxel,
  title = {Voxel Field Fusion for 3D Object Detection},
  author = {Li, Yanwei and Qi, Xiaojuan and Chen, Yukang and Wang, Liwei and Li, Zeming and Sun, Jian and Jia, Jiaya},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2205.15938v1},
  eprint = {2205.15938}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE