OcTr: Octree-Based Transformer for 3D Object Detection

Chao ZhouYanan ZhangJiaxin ChenDi Huang

article2023CVPR71 citations

Proposes an octree-based Transformer architecture that captures coarse-to-fine global context through hierarchical dynamic self-attention, reducing computational complexity to linear while achieving state-of-the-art 3D object detection on large-scale LiDAR point clouds.

Listen

Autonomous driving systems rely on three-dimensional object detection from laser scanner point clouds to safely recognize obstacles in real time. However, detecting objects that are distant or partially blocked remains a critical challenge due to sparse and incomplete sensor measurements. While attention-based deep learning models known as Transformers can capture broader contextual relationships across an entire scene, standard designs demand prohibitive computational memory and processing time when applied to large-scale three-dimensional data.

The article introduces and evaluates an Octree-based Transformer framework, termed OcTr, designed to capture comprehensive global context while maintaining high computational efficiency for three-dimensional perception.

The researchers developed a hierarchical attention mechanism that structures spatial features into a multi-scale pyramid and adaptively selects only the most informative regions from top to bottom, reducing overall computational complexity from quadratic to linear. They also integrated a hybrid semantic positional embedding and masking technique that guides the model to focus on meaningful foreground objects rather than empty background space. The proposed framework was benchmarked against existing state-of-the-art methods using two major public autonomous driving datasets, the Waymo Open Dataset and the KITTI dataset, evaluating performance across multiple vehicle, pedestrian, and cyclist categories.

The findings establish that OcTr achieves top-tier detection accuracy across standard benchmarks while maintaining a lightweight footprint. On the Waymo dataset, the system outperformed existing baselines across all evaluated object categories, demonstrating significant improvements on difficult, sparsely sampled pedestrian targets by over 2.5 percentage points in precision. Notably, detection gains were most pronounced for distant objects located beyond 50 meters, where OcTr exceeded previous best-performing methods by more than 2.2 to 2.6 percentage points. Furthermore, resource evaluations confirmed that OcTr operates with fewer parameters (2.9 million) and lower memory occupancy than comparable Transformer-based alternatives, and ablation studies confirmed that its modular architecture consistently enhances both single-stage and two-stage detection pipelines.

These results demonstrate that long-range contextual modeling can be achieved efficiently without requiring excessive onboard computing hardware. By significantly boosting detection accuracy for distant and partially hidden hazards, the framework offers valuable safety improvements and operational reliability for autonomous navigation pipelines.

Organizations developing perception systems for autonomous vehicles or mobile robotics should consider adopting coarse-to-fine attention mechanisms and semantic masking to enhance detection ranges. Future implementation efforts should include real-world road testing across diverse weather conditions, as well as integrating multi-sensor streams such as cameras and radar before full production deployment.

arXiv: 2303.12621
Cover for OcTr: Octree-Based Transformer for 3D Object Detection

Abstract

A key challenge for LiDAR-based 3D object detection is to capture sufficient features from large scale 3D scenes especially for distant or/and occluded objects. Albeit recent efforts made by Transformers with the long sequence modeling capability, they fail to properly balance the accuracy and efficiency, suffering from inadequate receptive fields or coarse-grained holistic correlations. In this paper, we propose an Octree-based Transformer, named OcTr, to address this issue. It first constructs a dynamic octree on the hierarchical feature pyramid through conducting self-attention on the top level and then recursively propagates to the level below restricted by the octants, which captures rich global context in a coarse-to-fine manner while maintaining the computational complexity under control. Furthermore, for enhanced foreground perception, we propose a hybrid positional embedding, composed of the semantic-aware positional embedding and attention mask, to fully exploit semantic and geometry clues. Extensive experiments are conducted on the Waymo Open Dataset and KITTI Dataset, and OcTr reaches newly state-of-the-art results.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 2.1. 3D Object Detection from Point Clouds
  • 2.2. Transformer in 3D Vision
  • 3. Method
  • 3.1. Framework
  • 3.2. Self-attention Revisit
  • 3.3. Octree Attention
  • 3.4. Semantic Positional Embedding
  • 4. Experiments
  • 4.1. Datasets and Implementation Details
  • 4.2. Results on WOD
  • 4.3. Results on KITTI
  • 4.4. Ablation study
  • 5. Conclusion
  • Acknowledgment
  • References

Knowls

  1. Knowl 1 — OcTr voxel-based detection architecture

    model/method

    OcTr is a voxel-based 3D object detector that replaces the conventional sparse-convolution-only 3D backbone with Octree Transformer Blocks (OTBs). The input LiDAR point cloud is voxelized into regular sparse grids, and sparse 3D convolutions produce patch-embedded voxel tokens. OTBs apply the proposed Octree Attention mechanism to learn global context, after which voxel features are downsampled, projected into a bird’s-eye-view representation, and processed by a multiscale dense 2D backbone. An anchor-based or anchor-free region-proposal head generates 3D detections, and an optional RoI head performs two-stage refinement. The design can be inserted into different voxel-based detectors by replacing their 3D backbone.

  2. Knowl 2 — Dynamic attention-guided octree construction

    model/method

    OctAttn constructs a differentiable, data-dependent octree over a multiscale sparse-voxel feature pyramid rather than using a fixed geometric partition. Let F0∈RM×dF_0\in\mathbb{R}^{M\times d} be the feature matrix of MM nonempty input voxels, let I0∈RM×3I_0\in\mathbb{R}^{M\times 3} contain their voxel coordinates, let dd be the feature dimension, and let NN be the pyramid height. At pyramid level n∈{0,…,N−1}n\in\{0,\ldots,N-1\}, coordinates are downsampled and features are aggregated as

    In=⌊I02n⌋,Fn=BN⁡ ⁣(Smax⁡(F0,In)),I_n=\left\lfloor\frac{I_0}{2^n}\right\rfloor,\qquad F_n=\operatorname{BN}\!\left(S_{\max}(F_0,I_n)\right),

    where Smax⁡S_{\max} is max scattering of fine-grid features into the coordinates InI_n, and BN⁡\operatorname{BN} is batch normalization. The coarsest feature map FN−1F_{N-1} is padded and reorganized into dense per-batch tokens FˉN−1∈RB×mN−1×d\bar F_{N-1}\in\mathbb{R}^{B\times m_{N-1}\times d}, where BB is batch size and mN−1m_{N-1} is the maximum number of nonempty voxels in a batch at the coarsest level. Multi-head self-attention is applied only at this top level; for head h∈{1,…,H}h\in\{1,\ldots,H\}, with learnable query, key, and value projections Wq(h),Wk(h),Wv(h)W_q^{(h)},W_k^{(h)},W_v^{(h)}, it produces attention scores

    AN−1=∑h=1Hsoftmax⁡ ⁣(FˉN−1Wq(h)(FˉN−1Wk(h))Td).A_{N-1}=\sum_{h=1}^{H}\operatorname{softmax}\!\left(\frac{\bar F_{N-1}W_q^{(h)}(\bar F_{N-1}W_k^{(h)})^{\mathsf T}}{\sqrt d}\right).

    For every top-level query token, the kk highest-scoring tokens in its attention row are selected as the relevant token group, and their indices define the octants that are recursively subdivided at finer pyramid levels. Thus, the octree is learned from feature relevance and can adapt its fine-grained resolution to the scene content.

  3. Knowl 3 — Coarse-to-fine sparse attention propagation

    model/method

    After selecting relevant octants at the coarsest pyramid level, OctAttn propagates information from level n+1n+1 to level nn for n=N−2,…,0n=N-2,\ldots,0. For each query at level nn, only KK voxels are uniformly sampled from the regions selected by its parent-level top-kk indices. The level-nn query sequence is Qˉn∈RB×mn×d\bar Q_n\in\mathbb{R}^{B\times m_n\times d}, while the sampled key and value sequences are Kˉn,Vˉn∈RB×mn×K×d\bar K_n,\bar V_n\in\mathbb{R}^{B\times m_n\times K\times d}. Cross-attention is then computed as

    An=∑h=1Hsoftmax⁡ ⁣(QˉnWq(h)(KˉnWk(h))Td),A_n=\sum_{h=1}^{H}\operatorname{softmax}\!\left(\frac{\bar Q_nW_q^{(h)}(\bar K_nW_k^{(h)})^{\mathsf T}}{\sqrt d}\right), Fˉn′=∑h=1HWh[softmax⁡ ⁣(QˉnWq(h)(KˉnWk(h))Td)VˉnWv(h)],\bar F'_n=\sum_{h=1}^{H}W_h\left[\operatorname{softmax}\!\left(\frac{\bar Q_nW_q^{(h)}(\bar K_nW_k^{(h)})^{\mathsf T}}{\sqrt d}\right)\bar V_nW_v^{(h)}\right],

    where HH is the number of heads, WhW_h is the output projection for head hh, and dd is the feature dimension. The restricted cross-attention is equivalent to applying a sparse mask to a full self-attention matrix, so fine-level queries retain access to globally selected context while avoiding all-to-all interactions. Because hard top-kk selection blocks gradients, training uses a Gumbel-top-kk relaxation. For candidate score An,iA_{n,i}, independently sampled Gumbel noise gig_i, temperature τ\tau, and mnm_n candidates, the differentiable normalized selection score is

    pi=exp⁡((An,i+gi)/τ)∑j=1mnexp⁡((An,j+gj)/τ)∈[0,1].p_i=\frac{\exp((A_{n,i}+g_i)/\tau)}{\sum_{j=1}^{m_n}\exp((A_{n,j}+g_j)/\tau)}\in[0,1].

    The noisy relaxed scores are used during training, while the original attention scores are used for inference. This top-down procedure gives OctAttn a global receptive field with progressively finer local representations.

  4. Knowl 4 — Octree Transformer block and linearized complexity

    theoretical result

    An Octree Transformer Block fuses the outputs from all pyramid levels after they have been scattered and upsampled back to the finest voxel locations. If Fˉn′∈Rm0×d\bar F'_n\in\mathbb{R}^{m_0\times d} denotes the compact upsampled output from level nn, m0m_0 is the number of fine-level nonempty voxels, ∣| denotes concatenation, and FC⁡\operatorname{FC} is a feature-dimension projection, the block computes

    F~=FC⁡ ⁣(FˉN−1′∣FˉN−2′∣⋯∣Fˉ0′)+LePE⁡(F0),\widetilde F=\operatorname{FC}\!\left(\bar F'_{N-1}|\bar F'_{N-2}|\cdots|\bar F'_0\right)+\operatorname{LePE}(F_0), F~′=BN⁡ ⁣(FFN⁡(F~))+F~.\widetilde F'=\operatorname{BN}\!\left(\operatorname{FFN}(\widetilde F)\right)+\widetilde F.

    Here FFN⁡\operatorname{FFN} is the feed-forward network, BN⁡\operatorname{BN} is batch normalization, and LePE⁡\operatorname{LePE} is a locally enhanced positional embedding implemented with submanifold sparse convolutions. LePE supplies local-neighbor interactions that complement OctAttn’s long-range context.

    If MM is the number of nonempty fine-level voxels, KK is the number of sampled attending voxels per fine-level query, NN is the pyramid height, and ω>1\omega>1 is the average sparse-voxel downsampling ratio between levels, OctAttn has time complexity

    O ⁣((MωN−1)2+∑n=0N−2KMωn)=O ⁣((MωN−1)2+ωω−1KM(1−ω1−N)).O\!\left(\left(\frac{M}{\omega^{N-1}}\right)^2+\sum_{n=0}^{N-2}\frac{KM}{\omega^n}\right) =O\!\left(\left(\frac{M}{\omega^{N-1}}\right)^2+\frac{\omega}{\omega-1}KM\left(1-\omega^{1-N}\right)\right).

    The quadratic term is confined to the coarsest level, while attention at finer levels is linear in the number of voxels when KK and the pyramid settings are fixed.

  5. Knowl 5 — Hybrid semantic and geometric positional embedding

    model/method

    OcTr introduces a hybrid positional embedding to reduce the dominance of background voxel pairs and to combine geometry with foreground semantics. A submanifold sparse-convolution branch predicts a sigmoid foreground score for every voxel and is trained with focal loss using ground-truth foreground supervision. For a voxel feature f∈Rdf\in\mathbb{R}^{d}, voxel-center coordinates (x,y,z)(x,y,z), and scalar foreground score ss, Semantic Absolute Positional Embedding (SAPE) concatenates the four positional/semantic values with the feature and applies a bias-free linear projection:

    SAPE⁡(f,x,y,z,s)=FC⁡d+4→d([x,y,z,s ∣ f])=FC⁡d→d(f)+FC⁡4→d(x,y,z,s).\operatorname{SAPE}(f,x,y,z,s)=\operatorname{FC}_{d+4\to d}([x,y,z,s\,|\,f]) =\operatorname{FC}_{d\to d}(f)+\operatorname{FC}_{4\to d}(x,y,z,s).

    SAPE is applied to the multiscale features; the scatter operations used to form coarse levels also preserve aggregated spatial and semantic information. OcTr additionally applies a Semantic Attention Mask (SAM). Let A∈RNq×NkA\in\mathbb{R}^{N_q\times N_k} be a pre-softmax attention matrix, Sq∈RNqS_q\in\mathbb{R}^{N_q} and Sk∈RNkS_k\in\mathbb{R}^{N_k} be query and key/value foreground scores, δq\delta_q and δk\delta_k be their thresholds, 1(⋅)\mathbf{1}(\cdot) be an elementwise indicator, and Γ\Gamma be a sufficiently large masking scalar. The masked attention is

    A′=softmax⁡ ⁣(−Γ[1Nq×Nk−1(Sq≥δq)1(Sk≥δk)T]+A).A'=\operatorname{softmax}\!\left(-\Gamma\left[\mathbf{1}_{N_q\times N_k}-\mathbf{1}(S_q\ge\delta_q)\mathbf{1}(S_k\ge\delta_k)^{\mathsf T}\right]+A\right).

    Consequently, attention links between queries or keys classified as inferior foreground candidates are suppressed, and the resulting mask is broadcast to all attention heads. The complete detector is optimized with Ldet=Lrpn+Lrcnn+Lseg\mathcal{L}_{\mathrm{det}}=\mathcal{L}_{\mathrm{rpn}}+\mathcal{L}_{\mathrm{rcnn}}+\mathcal{L}_{\mathrm{seg}}, with Lrcnn\mathcal{L}_{\mathrm{rcnn}} omitted for one-stage detectors.

  6. Knowl 6 — Waymo Open Dataset detection performance

    empirical result

    On the Waymo Open Dataset validation split of 202 sequences, OcTr was trained on 20% of the samples using single-frame, three-class detection. Waymo reports both mean average precision (mAP) and heading-weighted mAP (mAPH); LEVEL 1 (L1) contains boxes with more than five LiDAR points and LEVEL 2 (L2) contains boxes with one to five points. The IoU thresholds are 0.7 for vehicles and 0.5 for pedestrians and cyclists.

    OcTr obtained the following mAP/mAPH scores: vehicle L1 78.12/77.6378.12/77.63, vehicle L2 69.79/69.3469.79/69.34, pedestrian L1 80.76/74.3980.76/74.39, pedestrian L2 72.48/66.5272.48/66.52, cyclist L1 72.58/71.5072.58/71.50, and cyclist L2 69.93/68.9069.93/68.90. Against the strongest comparable three-class baseline, PVRCNN++, whose corresponding scores were vehicle 77.82/77.3277.82/77.32 and 69.07/68.6269.07/68.62, pedestrian 77.99/71.3677.99/71.36 and 69.92/63.7469.92/63.74, and cyclist 71.80/70.7171.80/70.71 and 69.31/68.2669.31/68.26, OcTr improves every listed L1/L2 result. The paper also reports a single-class SST result, but that result is not directly comparable to the jointly trained three-class models.

  7. Knowl 7 — Improved long-range vehicle detection on Waymo

    empirical result

    For vehicle detection on the Waymo validation set, OcTr was evaluated with 3D mAP at L1 and L2 difficulty over distance ranges of 0–30 m, 30–50 m, and 50 m to infinity. OcTr achieved L1 mAP values of 92.9992.99 overall, 92.9992.99 at 0–30 m, 77.6677.66 at 30–50 m, and 58.0258.02 beyond 50 m. Its L2 values were 70.5070.50 overall, 91.7891.78 at 0–30 m, 71.2871.28 at 30–50 m, and 45.4645.46 beyond 50 m.

    The strongest previously reported values in the same comparison were L1 77.8277.82 overall, 93.1393.13 at 0–30 m, 77.2177.21 at 30–50 m, and 54.7554.75 beyond 50 m, and L2 70.2170.21 overall, 92.4192.41 at 0–30 m, 70.1070.10 at 30–50 m, and 43.2043.20 beyond 50 m. Thus, OcTr’s principal advantage is at 30–50 m and beyond 50 m rather than in the nearest range, supporting the claim that its coarse-to-fine global context is especially useful for sparse distant objects.

  8. Knowl 8 — KITTI benchmark performance

    empirical result

    On the KITTI 3D detection benchmark, OcTr used the official car, pedestrian, and cyclist IoU thresholds of 0.70.7, 0.50.5, and 0.50.5, respectively. For single-stage validation models using the same anchor-based proposal head, OcTr achieved car mAP values of 88.4388.43, 78.5778.57, and 77.1677.16 for easy, moderate, and hard difficulty; pedestrian values of 61.4961.49, 57.1757.17, and 52.3552.35; and cyclist values of 85.2985.29, 70.4470.44, and 66.1766.17. The corresponding SECOND baseline scores were car 88.61/78.62/77.2288.61/78.62/77.22, pedestrian 56.55/52.98/47.7356.55/52.98/47.73, and cyclist 80.58/67.15/63.1080.58/67.15/63.10, so OcTr’s largest single-stage gains occur for pedestrians and cyclists.

    For the two-stage car detector, OcTr obtained test-set mAP values of 90.8890.88 easy, 82.6482.64 moderate, and 77.7777.77 hard, with a three-level mean of 83.7683.76. On the validation set it obtained 89.8089.80, 86.9786.97, and 79.2879.28, with a mean of 85.3585.35. The test moderate score exceeds the reported Focals Conv score of 82.2882.28, and the validation mean exceeds the reported second-best mean of 84.8684.86.

  9. Knowl 9 — Ablations validate OctAttn, positional components, and top-k selection

    empirical result

    Ablations on the Waymo validation set show that the proposed attention and positional components are complementary. With the same detector, vehicle/pedestrian L1/L2 mAP was 73.3/65.173.3/65.1 and 68.1/60.468.1/60.4 for OctAttn, compared with 71.4/63.671.4/63.6 and 65.7/57.965.7/57.9 for Performer, 71.7/63.571.7/63.5 and 64.3/56.164.3/56.1 for ACT, 69.4/61.569.4/61.5 and 65.0/57.065.0/57.0 for VoTr, and 68.2/59.868.2/59.8 and 64.9/56.764.9/56.7 for nearest-KK sampling.

    Starting from the positional-embedding baseline, vehicle/pedestrian L1/L2 mAP was 71.35/63.3071.35/63.30 and 65.75/57.8965.75/57.89. Adding LePE increased these values to 72.34/64.3272.34/64.32 and 66.56/58.6266.56/58.62; adding LePE with SAPE gave 72.64/64.4672.64/64.46 and 66.62/58.8366.62/58.83; adding LePE with SAM gave 72.86/64.4072.86/64.40 and 67.79/59.9067.79/59.90; and using LePE, SAPE, and SAM together gave the full OcTr scores of 73.28/65.0573.28/65.05 and 68.08/60.3668.08/60.36.

    The number kk of top-level selected regions also affects accuracy: with k=1k=1, 44, 88, and 1616, vehicle L1/L2 mAP was respectively 70.38/62.2070.38/62.20, 72.58/64.4272.58/64.42, 73.28/65.0573.28/65.05, and 73.25/65.0173.25/65.01, while pedestrian L1/L2 mAP was 64.19/56.4364.19/56.43, 66.21/58.4266.21/58.42, 68.08/60.3668.08/60.36, and 67.89/60.1067.89/60.10. Performance therefore saturates near k=8k=8.

  10. Knowl 10 — Detector scalability and resource trade-off

    empirical result

    OcTr improves several detector architectures on the Waymo validation set, showing that the 3D backbone is not tied to a particular proposal design. With SECOND, vehicle and pedestrian L1/L2 mAP changed from 70.96/62.5870.96/62.58 and 65.23/57.2265.23/57.22 to 73.28/65.0573.28/65.05 and 68.08/60.3668.08/60.36. With PV-RCNN, the corresponding changes were from 75.41/67.4475.41/67.44 and 71.98/63.7071.98/63.70 to 76.77/68.3176.77/68.31 and 73.22/64.3073.22/64.30. With PV-RCNN++, they were from 77.82/69.0777.82/69.07 and 77.99/69.9277.99/69.92 to 78.01/69.6078.01/69.60 and 80.75/72.4580.75/72.45.

    On KITTI using single-stage detectors evaluated on a GTX 2080 Ti, OcTr-SSD used 2.92.9 million parameters, 6464 ms latency, and 2.52.5 GB memory. SECOND used 5.35.3 million parameters, 4848 ms, and 2.32.3 GB; VoTR-SSD used 4.84.8 million, 6767 ms, and 3.03.0 GB; and VoxSeT-SSD used 3.03.0 million, 3737 ms, and 3.63.6 GB. OcTr therefore has the fewest parameters and lower memory use than the Transformer alternatives, while trading some latency for its richer voxel representation and detection accuracy.

Coverage note — Qualitative detection visualizations, the full per-baseline benchmark tables, and routine implementation details were omitted because they add no distinct load-bearing method or result beyond the summarized quantitative evidence.

References

  1. 1.Joshua Ainslie, Santiago Ontanon, Chris Alberti, Vaclav Cvicek, Zachary Fisher, Philip Pham, Anirudh Ravula, Sumit Sanghai, Qifan Wang, and Li Yang. Etc: Encoding long and structured inputs in transformers. In EMNLP, 2020.
  2. 2.Iz Beltagy, Matthew E Peters, and Arman Cohan. Longformer: The long-document transformer. arXiv preprint arXiv:2004.05150, 2020.
  3. 3.Chen Chen, Zhe Chen, Jing Zhang, and Dacheng Tao. Sasa: Semantics-augmented set abstraction for point-based 3d object detection. In AAAI, volume 1, 2022.
  4. 4.Yukang Chen, Yanwei Li, Xiangyu Zhang, Jian Sun, and Jiaya Jia. Focal sparse convolutional networks for 3d object detection. In CVPR, pages 5428–5437, 2022.
  5. 5.Krzysztof Marcin Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Andreea Gane, Tamas Sarlós, Peter Hawkins, Jared Quincy Davis, Afroz Mohiuddin, Lukasz Kaiser, David Benjamin Belanger, Lucy J. Colwell, and Adrian Weller. Rethinking attention with performers. In ICLR, 2021.
  6. 6.Jiajun Deng, Shaoshuai Shi, Peiwei Li, Wengang Zhou, Yanyong Zhang, and Houqiang Li. Voxel r-cnn: Towards high performance voxel-based 3d object detection. In AAAI, volume 35, pages 1201–1209, 2021.
  7. 7.Xiaoyi Dong, Jianmin Bao, Dongdong Chen, Weiming Zhang, Nenghai Yu, Lu Yuan, Dong Chen, and Baining Guo. Cswin transformer: A general vision transformer backbone with cross-shaped windows. In CVPR, pages 12124–12134, 2022.
  8. 8.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. In ICLR, 2020.
  9. 9.Lue Fan, Ziqi Pang, Tianyuan Zhang, Yu-Xiong Wang, Hang Zhao, Feng Wang, Naiyan Wang, and Zhaoxiang Zhang. Embracing single stride 3d object detector with sparse transformer. In CVPR, pages 8458–8468, 2022.
  10. 10.Andreas Geiger, Philip Lenz, and Raquel Urtasun. Are we ready for autonomous driving? the kitti vision benchmark suite. In CVPR, pages 3354–3361. IEEE, 2012.
  11. 11.Meng-Hao Guo, Jun-Xiong Cai, Zheng-Ning Liu, Tai-Jiang Mu, Ralph R Martin, and Shi-Min Hu. Pct: Point cloud transformer. CVM, 7(2):187–199, 2021.
  12. 12.Chenhang He, Ruihuang Li, Shuai Li, and Lei Zhang. Voxel set transformer: A set-to-set approach to 3d object detection from point clouds. In CVPR, pages 8417–8427, 2022.
  13. 13.Chenhang He, Hui Zeng, Jianqiang Huang, Xian-Sheng Hua, and Lei Zhang. Structure aware single-stage 3d object detection from point cloud. In CVPR, pages 11873–11882, 2020.
  14. 14.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In CVPR, pages 770–778, 2016.
  15. 15.Jordan SK Hu. Point density-aware voxels for lidar 3d object detection. In CVPR, pages 8469–8478, 2022.
  16. 16.Sergey Ioffe and Christian Szegedy. Batch normalization: Accelerating deep network training by reducing internal covariate shift. In ICML, pages 448–456. PMLR, 2015.
  17. 17.Eric Jang, Shixiang Gu, and Ben Poole. Categorical reparameterization with gumbel-softmax. In ICLR, 2016.
  18. 18.Alex H Lang, Sourabh Vora, Holger Caesar, Lubing Zhou, Jiong Yang, and Oscar Beijbom. Pointpillars: Fast encoders for object detection from point clouds. In CVPR, pages 12697–12705, 2019.
  19. 19.Juho Lee, Yoonho Lee, Jungtaek Kim, Adam Kosiorek, Seungjin Choi, and Yee Whye Teh. Set transformer: A framework for attention-based permutation-invariant neural networks. In ICML, pages 3744–3753. PMLR, 2019.
  20. 20.Zhichao Li, Feng Wang, and Naiyan Wang. Lidar r-cnn: An efficient and universal 3d object detector. In CVPR, pages 7546–7555, 2021.
  21. 21.Youwei Liang, Chongjian Ge, Zhan Tong, Yibing Song, Jue Wang, and Pengtao Xie. Not all patches are what you need: Expediting vision transformers via token reorganizations. In ICLR, 2022.
  22. 22.Tsung-Yi Lin, Priya Goyal, Ross Girshick, Kaiming He, and Piotr Dollár. Focal loss for dense object detection. In ICCV, pages 2980–2988, 2017.
  23. 23.Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. In ICCV, pages 10012–10022, 2021.
  24. 24.Jiageng Mao. Voxel transformer for 3d object detection. In ICCV, pages 3164–3173, 2021.
  25. 25.Jiageng Mao, Minzhe Niu, Haoyue Bai, Xiaodan Liang, Hang Xu, and Chunjing Xu. Pyramid r-cnn: Towards better performance and adaptability for 3d object detection. In ICCV, pages 2723–2732, 2021.
  26. 26.Julien NP Martel, David B Lindell, Connor Z Lin, Eric R Chan, Marco Monteiro, and Gordon Wetzstein. Acorn: Adaptive coordinate networks for neural scene representation. In ACM SIGGRAPH, 2021.
  27. 27.Zheng Minghang, Gao Peng, Zhang Renrui, Li Kunchang, Li Hongsheng, and Dong Hao. End-to-end object detection with adaptive clustering transformer. In BMVC, page 226, 2021.
  28. 28.Ishan Misra, Rohit Girdhar, and Armand Joulin. An end-to-end transformer model for 3d object detection. In ICCV, pages 2906–2917, 2021.
  29. 29.Xuran Pan, Zhuofan Xia, Shiji Song, Li Erran Li, and Gao Huang. 3d object detection with pointformer. In CVPR, pages 7463–7472, 2021.
  30. 30.Chunghyun Park, Yoonwoo Jeong, Minsu Cho, and Jaesik Park. Fast point transformer. In CVPR, pages 16949–16958, 2022.
  31. 31.Charles R Qi, Or Litany, Kaiming He, and Leonidas J Guibas. Deep hough voting for 3d object detection in point clouds. In ICCV, pages 9277–9286, 2019.
  32. 32.Charles R Qi, Wei Liu, Chenxia Wu, Hao Su, and Leonidas J Guibas. Frustum pointnets for 3d object detection from rgb-d data. In CVPR, pages 918–927, 2018.
  33. 33.Charles Ruizhongtai Qi, Li Yi, Hao Su, and Leonidas J Guibas. Pointnet++: Deep hierarchical feature learning on point sets in a metric space. In NeurIPS, volume 30, 2017.
  34. 34.Jiezhong Qiu, Hao Ma, Omer Levy, Scott Wen-tau Yih, Sinong Wang, and Jie Tang. Blockwise self-attention for long document understanding. arXiv preprint arXiv:1911.02972, 2019.
  35. 35.Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. Faster r-cnn: Towards real-time object detection with region proposal networks. In NeurIPS, volume 28, 2015.
  36. 36.Gernot Riegler, Ali Osman Ulusoy, and Andreas Geiger. Octnet: Learning deep 3d representations at high resolutions. In CVPR, pages 3577–3586, 2017.
  37. 37.Hualian Sheng, Sijia Cai, Yuan Liu, Bing Deng, Jianqiang Huang, Xian-Sheng Hua, and Min-Jian Zhao. Improving 3d object detection with channel-wise transformer. In ICCV, pages 2743–2752, 2021.
  38. 38.Shaoshuai Shi, Chaoxu Guo, Li Jiang, Zhe Wang, Jianping Shi, Xiaogang Wang, and Hongsheng Li. Pv-rcnn: Point-voxel feature set abstraction for 3d object detection. In CVPR, pages 10529–10538, 2020.
  39. 39.Shaoshuai Shi, Li Jiang, Jiajun Deng, Zhe Wang, Chaoxu Guo, Jianping Shi, Xiaogang Wang, and Hongsheng Li. Pv-rcnn++: Point-voxel feature set abstraction with local vector representation for 3d object detection. IEEE TPAMI, 2021.
  40. 40.Shaoshuai Shi, Xiaogang Wang, and Hongsheng Li. Pointr-cnn: 3d object proposal generation and detection from point cloud. In CVPR, pages 770–779, 2019.
  41. 41.Shaoshuai Shi, Zhe Wang, Jianping Shi, Xiaogang Wang, and Hongsheng Li. From points to parts: 3d object detection from point cloud with part-aware and part-aggregation network. IEEE TPAMI, 43(8):2647–2664, 2020.
  42. 42.Weijing Shi and Raj Rajkumar. Point-gnn: Graph neural network for 3d object detection in a point cloud. In CVPR, pages 1711–1719, 2020.
  43. 43.Pei Sun, Henrik Kretzschmar, Xerxes Dotiwalla, Aurelien Chouard, Vijaysai Patnaik, Paul Tsui, James Guo, Yin Zhou, Yuning Chai, Benjamin Caine, et al. Scalability in perception for autonomous driving: Waymo open dataset. In CVPR, pages 2446–2454, 2020.
  44. 44.Jia-Heng Tang, Weikai Chen, Jie Yang, Bo Wang, Songrun Liu, Bo Yang, and Lin Gao. Octfield: Hierarchical implicit functions for 3d modeling. In NeurIPS, 2021.
  45. 45.Shitao Tang, Jiahui Zhang, Siyu Zhu, and Ping Tan. Quadtree attention for vision transformers. In ICLR, 2022.
  46. 46.OpenPCDet Development Team. Openpcdet: An open-source toolbox for 3d object detection from point clouds. https://github.com/open-mmlab/OpenPCDet, 2020.
  47. 47.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In NeurIPS, volume 30, 2017.
  48. 48.Peng-Shuai Wang, Yang Liu, Yu-Xiao Guo, Chun-Yu Sun, and Xin Tong. O-cnn: Octree-based convolutional neural networks for 3d shape analysis. ACM TOG, 36(4):1–11, 2017.
  49. 49.Peng-Shuai Wang, Chun-Yu Sun, Yang Liu, and Xin Tong. Adaptive o-cnn: A patch-based deep representation of 3d shapes. ACM TOG, 37(6):1–11, 2018.
  50. 50.Yue Wang, Yongbin Sun, Ziwei Liu, Sanjay E Sarma, Michael M Bronstein, and Justin M Solomon. Dynamic graph cnn for learning on point clouds. ACM TOG, 38(5):1–12, 2019.
  51. 51.Tete Xiao, Mannat Singh, Eric Mintun, Trevor Darrell, Piotr Dollár, and Ross Girshick. Early convolutions help transformers see better. In NeurIPS, volume 34, pages 30392–30400, 2021.
  52. 52.Yan Yan, Yuxing Mao, and Bo Li. Second: Sparsely embedded convolutional detection. Sensors, 18(10):3337, 2018.
  53. 53.Zetong Yang, Yanan Sun, Shu Liu, and Jiaya Jia. 3dssd: Point-based 3d single stage object detector. In CVPR, pages 11040–11048, 2020.
  54. 54.Zetong Yang, Yanan Sun, Shu Liu, Xiaoyong Shen, and Jiaya Jia. Std: Sparse-to-dense 3d object detector for point cloud. In CVPR, pages 1951–1960, 2019.
  55. 55.Tianwei Yin, Xingyi Zhou, and Philipp Krahenbuhl. Center-based 3d object detection and tracking. In CVPR, pages 11784–11793, 2021.
  56. 56.Manzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie, Chris Alberti, Santiago Ontanon, Philip Pham, Anirudh Ravula, Qifan Wang, Li Yang, et al. Big bird: Transformers for longer sequences. In NeurIPS, volume 33, pages 17283–17297, 2020.
  57. 57.Cheng Zhang, Haocheng Wan, Xinyi Shen, and Zizhao Wu. Patchformer: An efficient point transformer with patch attention. In CVPR, pages 11799–11808, 2022.
  58. 58.Yanan Zhang, Jiaxin Chen, and Di Huang. Cat-det: Contrastively augmented transformer for multi-modal 3d object detection. In CVPR, pages 908–917, 2022.
  59. 59.Yifan Zhang, Qingyong Hu, Guoquan Xu, Yanxin Ma, Jianwei Wan, and Yulan Guo. Not all points are equal: Learning highly efficient point-based detectors for 3d lidar point clouds. In CVPR, pages 18953–18962, 2022.
  60. 60.Yanan Zhang, Di Huang, and Yunhong Wang. Pc-rgnn: Point cloud completion and graph neural network for 3d object detection. In AAAI, volume 35, pages 3430–3437, 2021.
  61. 61.Hengshuang Zhao, Li Jiang, Jiaya Jia, Philip HS Torr, and Vladlen Koltun. Point transformer. In ICCV, pages 16259–16268, 2021.
  62. 62.Wu Zheng, Weiliang Tang, Sijin Chen, Li Jiang, and Chi-Wing Fu. Cia-ssd: Confident iou-aware single-stage object detector from point cloud. In AAAI, volume 35, pages 3555–3562, 2021.
  63. 63.Yin Zhou and Oncel Tuzel. Voxelnet: End-to-end learning for point cloud based 3d object detection. In CVPR, pages 4490–4499, 2018.

Citation

MLA
Zhou, C., et al. “OcTr: Octree-based Transformer for 3D Object Detection”. arXiv, 2023, http://arxiv.org/abs/2303.12621v1.
APA
Zhou, C., Zhang, Y., Chen, J., & Huang, D. (2023). OcTr: Octree-based Transformer for 3D Object Detection. arXiv. http://arxiv.org/abs/2303.12621v1
Chicago
Zhou, C., Y. Zhang, J. Chen, and D. Huang. 2023. “OcTr: Octree-based Transformer for 3D Object Detection”. arXiv. http://arxiv.org/abs/2303.12621v1.
Harvard
Zhou, C. et al. (2023) “OcTr: Octree-based Transformer for 3D Object Detection”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2303.12621v1.
Vancouver
1. Zhou C, Zhang Y, Chen J, Huang D (2023) OcTr: Octree-based Transformer for 3D Object Detection. arXiv

BibTeX

@article{zhou2023octr,
  title = {OcTr: Octree-based Transformer for 3D Object Detection},
  author = {Zhou, Chao and Zhang, Yanan and Chen, Jiaxin and Huang, Di},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2303.12621v1},
  eprint = {2303.12621}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE