SparseOcc: Rethinking Sparse Latent Representation for Vision-Based Semantic Occupancy Prediction

Pin TangZhongdao WangGuoqing WangJilai ZhengXiangxuan RenBailan FengChao Ma

article2024CVPR96 citations

Proposes an efficient vision-based 3D semantic occupancy prediction network that replaces dense volume processing with a purely sparse latent representation, cutting computation by up to 74.9% while improving accuracy by preventing feature hallucinations in empty space.

Listen

Vision-based 3D semantic occupancy prediction is critical for autonomous vehicles to understand surrounding geometry and identify dynamic and static objects. However, existing approaches rely on dense 3D representations that incur heavy computational and memory costs, or they compress 3D space into flattened 2D projections like Bird's Eye View, which sacrifices fine-grained spatial accuracy. Because the vast majority of 3D driving scenes are physically empty, these conventional architectures perform redundant calculations over empty space.

The article demonstrates an efficient vision-based occupancy network called SparseOcc, which replaces dense 3D volumes with a purely sparse representation to improve both computational efficiency and prediction accuracy.

The research evaluated SparseOcc on standard autonomous driving benchmarks, specifically nuScenes-Occupancy and SemanticKITTI, using multi-view and monocular camera inputs. The architecture transforms 2D image features into 3D space, stores only the occupied locations, and processes them through three core components: a 3D sparse latent diffuser with spatially decomposed convolutional kernels for scene completion, a multi-scale sparse feature pyramid fused via linear interpolation, and a sparse transformer head that segments occupied voxels while assigning empty space to a single unified token.

The experimental findings highlight substantial efficiency and performance gains. On the nuScenes-Occupancy benchmark, SparseOcc reduced computation floating-point operations by 59.8% to 74.9% and memory usage by 31.6% to 40.9% compared to dense and projected baselines. It reduced 3D inference latency to 0.19 seconds and overall latency to 0.25 seconds, outperforming dense baselines that required over two seconds. Simultaneously, semantic accuracy improved from 12.8% to 14.1% mean Intersection over Union over the top-performing dense model, driven by the sparse structure's inherent ability to avoid false predictions on empty space. On the SemanticKITTI benchmark, the method achieved a competitive 13.12% semantic accuracy while utilizing only 44.2% of the computation of prior transformer baselines.

These results demonstrate that sparse 3D representations eliminate the historic trade-off between computational overhead and geometric fidelity. For autonomous driving programs, adopting sparse processing reduces onboard hardware costs and power consumption, while lowering latency to improve real-time driving safety.

Engineering teams developing vision-based perception systems should adopt SparseOcc as a foundational architecture for occupancy prediction. When deploying the system, developers must balance image resolution and feature density, as excessively dense inputs can cause spatial hallucinations that degrade geometry scores unless diffusion layers are appropriately tuned. Future efforts should focus on validating the model on physical vehicle hardware across varying weather and edge-case operational conditions.

No sufficiently relevant recommendations were found.

Cover for SparseOcc: Rethinking Sparse Latent Representation for Vision-Based Semantic Occupancy Prediction

Abstract

Vision-based perception for autonomous driving requires an explicit modeling of a 3D space, where 2D latent representations are mapped and subsequent 3D operators are applied. However, operating on dense latent spaces introduces a cubic time and space complexity, which limits its scalability in terms of perception range or spatial resolution. Existing approaches compress the dense representation using projections like Bird’s Eye View (BEV) or Tri-Perspective View (TPV). Although efficient, these projections result in information loss, especially for tasks like semantic occupancy prediction. To address this, we propose SparseOcc, an efficient occupancy network inspired by sparse point cloud processing. It utilizes a lossless sparse latent representation with three key innovations. Firstly, a 3D sparse diffuser performs latent completion using spatially decomposed 3D sparse convolutional kernels. Secondly, a feature pyramid and sparse interpolation enhance scales with information from others. Finally, the transformer head is redesigned as a sparse variant. SparseOcc achieves a remarkable 74.9% reduction on FLOPs over the dense baseline. Interestingly, it also improves accuracy, from 12.8% to 14.1% mIOU, which in part can be attributed to the sparse representation’s ability to avoid hallucinations on empty voxels.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 2.1. 3D Scene Representation
  • 2.2. Scene Completion in Occupancy Prediction
  • 3. Proposed Approach
  • 3.1. Preliminary: Sparse Representation
  • 3.2. Sparse Latent Diffuser
  • 3.3. Sparse Feature Pyramid
  • 3.4. Sparse Transformer Head
  • 3.5. Objective Function
  • 4. Experiments
  • 4.1. Experimental Setup
  • 4.2. Benchmark Performance
  • 4.3. Qualitative Evaluation
  • 4.4. Ablation Studies
  • 5. Conclusion
  • References

Knowls

  1. Knowl 1 — Sparse Latent Representation Formulation in SparseOcc

    model/method

    In SparseOcc, 2D perspective image features from multi-view or monocular cameras are lifted to a 3D grid using estimated depth maps following the Lift-Splat-Shoot (LSS) paradigm, producing a dense voxel feature volume V∈RH×W×D×C\mathbf{V} \in \mathbb{R}^{H \times W \times D \times C}, where H,W,DH, W, D are spatial grid dimensions and CC is the feature channel dimension.

    Because LSS projects features along rays into empty 3D space, approximately 80%80\% of voxels in V\mathbf{V} contain no information. SparseOcc converts V\mathbf{V} into a lossless sparse tensor in Coordinate (COO) format by collecting only the NN active (non-empty) voxels:

    V={(pi=[xi,yi,zi]∈R3,fi∈RC)∣i=1,2,…,N}\mathcal{V} = \left\{ (\mathbf{p}_i = [x_i, y_i, z_i] \in \mathbb{R}^3, \mathbf{f}_i \in \mathbb{R}^C) \mid i = 1, 2, \dots, N \right\}

    where pi\mathbf{p}_i represents the 3D voxel coordinates and fi\mathbf{f}_i denotes the corresponding feature vector. All downstream 3D operations in SparseOcc operate exclusively on this sparse tensor V\mathcal{V}, eliminating computation and memory allocation on unoccupied space.

  2. Knowl 2 — Sparse Latent Diffuser with Decomposed Orthogonal Kernels

    model/method

    Because ray-casted 3D features only cover visible surfaces, scene completion requires diffusing features to adjacent unobserved regions while maintaining spatial sparsity. The Sparse Latent Diffuser achieves this via two sequential sub-blocks:

    1. Sparse Completion Block: Employs regular 3D sparse convolution to propagate non-empty features to their direct neighbors. To match geometric priors in autonomous driving scenes (such as horizontally planar roads vs. vertically tall buildings) and reduce computation, a standard k×k×kk \times k \times k kernel is decomposed into three sequential 1D/2D kernels of sizes k×k×1k \times k \times 1, k×1×kk \times 1 \times k, and 1×k×k1 \times k \times k. This decomposes kernel complexity from O(k3)\mathcal{O}(k^3) to O(3k2)\mathcal{O}(3k^2). To limit uncontrolled densification, only a single decomposed completion layer is used per diffuser stage.

    2. Contextual Aggregation Block: Aggregates local context without expanding the set of active voxels by utilizing 3D sparse submanifold convolutions, where output locations are active if and only if the corresponding input location is active. Following Cylindrical convolution designs, the k×k×kk \times k \times k submanifold convolution is decomposed into two parallel asymmetric branches: one branch consists of consecutive 1×k×k1 \times k \times k and k×1×kk \times 1 \times k kernels, while the second branch consists of consecutive k×1×kk \times 1 \times k and 1×k×k1 \times k \times k kernels, with their outputs summed.

    In the deepest scales where height resolution DD is small, the 3D completion convolution is replaced with a 2D 3×33 \times 3 sparse convolution, and the aggregation block is replaced with two parallel 2D 5×55 \times 5 submanifold convolutions.

  3. Knowl 3 — Multi-Scale Sparse Feature Pyramid and Sparse Interpolation Decoder

    model/method

    To capture both small foreground objects and expansive static elements without excessive diffusion layers, SparseOcc constructs a multi-scale sparse feature pyramid {Vl}l=1L\{\mathcal{V}_l\}_{l=1}^L (with L=4L=4). The encoder stacks LL sparse latent diffuser stages, each followed by a sparse convolution with stride 2 for spatial downsampling.

    To fuse inter-scale representations without the heavy memory and computational overhead of 3D multi-scale deformable attention (MSDeformAttn3D), the Sparse Voxel Decoder uses linear sparse interpolation and learned scale weighting. For any scale l∈{1,…,L}l \in \{1, \dots, L\}, the augmented multi-scale feature V^l\hat{\mathcal{V}}_l is computed as:

    V^l=∑j≠lWj⋅Interp(Vj,Vl)\hat{\mathcal{V}}_l = \sum_{j \neq l} W_j \cdot \text{Interp}(\mathcal{V}_j, \mathcal{V}_l)

    where WjW_j is a learnable scale weight matrix/scalar for the jj-th feature scale, and Interp(Vj,Vl)\text{Interp}(\mathcal{V}_j, \mathcal{V}_l) denotes linear interpolation projecting the sparse coordinates and features of scale Vj\mathcal{V}_j to the active voxel coordinates of target scale Vl\mathcal{V}_l. This enriches high-resolution representations with coarse semantic context and denser completions from lower-resolution scales.

  4. Knowl 4 — 3D Sparse Transformer Prediction Head

    model/method

    SparseOcc formulates 3D semantic occupancy prediction as a sparse mask classification task using a 3D transformer head:

    1. Coarse Binary Filtering: A linear binary classifier predicts occupancy for all voxels in the sparse pyramid feature V^l\hat{\mathcal{V}}_l. Voxels classified as non-occupied are filtered out, leaving NlN_l occupied voxels V^l={(pi,fi)∣i=1,…,Nl}\hat{\mathcal{V}}_l = \{(\mathbf{p}_i, \mathbf{f}_i)\mid i=1,\dots,N_l\}. A single learnable embedding token pϕ∈RC\mathbf{p}_\phi \in \mathbb{R}^C is used to represent all empty voxels collectively.

    2. Sparse Query Decoding: Given NqN_q learnable object queries Q∈RNq×C\mathbf{Q} \in \mathbb{R}^{N_q \times C}, the model performs outer products between Q\mathbf{Q} and {fi∣(pi,fi)∈V^l}∪{pϕ}\{\mathbf{f}_i \mid (\mathbf{p}_i, \mathbf{f}_i) \in \hat{\mathcal{V}}_l\} \cup \{\mathbf{p}_\phi\}. This produces sparse occupied masks Mocc∈RNq×Nl\mathbf{M}_\text{occ} \in \mathbb{R}^{N_q \times N_l} and an empty mask Mϕ∈RNq×1\mathbf{M}_\phi \in \mathbb{R}^{N_q \times 1}. The full dense 3D mask M∈RNq×H×W×D\mathbf{M} \in \mathbb{R}^{N_q \times H \times W \times D} is reconstructed via a gradient-preserving scatter operation using coordinates p\mathbf{p}. This sparse formulation reduces the mask decoding complexity from dense O(NqHWDC)\mathcal{O}(N_q H W D C) to O(NlNqC+HWD)\mathcal{O}(N_l N_q C + H W D), where Nl≪HWDN_l \ll H W D. Concurrently, a linear classifier predicts semantic class probabilities over CC classes for each query.

    3. Query Updating via 3D Masked Attention: Across transformer layers l=1,…,Ll = 1, \dots, L, queries are updated using attention constrained by the resized predicted binary mask from layer l−1l-1:

    Ql=softmax(Ml−1+WqQl−1(WkVˉl)T)WvVˉl+Ql−1\mathbf{Q}_l = \text{softmax}\left( \mathbf{M}_{l-1} + \mathbf{W}_q \mathbf{Q}_{l-1} (\mathbf{W}_k \bar{\mathbf{V}}_l)^T \right) \mathbf{W}_v \bar{\mathbf{V}}_l + \mathbf{Q}_{l-1}

    where Wq,Wk,Wv\mathbf{W}_q, \mathbf{W}_k, \mathbf{W}_v are linear projection layers, Vˉl\bar{\mathbf{V}}_l is the dense representation reconstructed from V^l\hat{\mathcal{V}}_l, and the spatial attention mask is defined as:

    Ml−1(x,y,z)={0if σ(Ml−1′(x,y,z))≥0.5−∞otherwise\mathbf{M}_{l-1}(x,y,z) = \begin{cases} 0 & \text{if } \sigma(\mathbf{M}'_{l-1}(x,y,z)) \ge 0.5 \\ -\infty & \text{otherwise} \end{cases}

    with σ(⋅)\sigma(\cdot) being the sigmoid function and Ml−1′=maxpooling(Ml−1)\mathbf{M}'_{l-1} = \text{maxpooling}(\mathbf{M}_{l-1}) downsampling the mask to match the spatial resolution of Vˉl\bar{\mathbf{V}}_l.

  5. Knowl 5 — SparseOcc Multi-Task Loss Objective

    equation

    Training SparseOcc employs Hungarian bipartite matching between the NqN_q predicted 3D masks (and their class predictions) and ground-truth semantic occupancy masks. The total loss L\mathcal{L} is defined as:

    L=Lmask+Lcls+Ldepth+Lseg\mathcal{L} = \mathcal{L}_\text{mask} + \mathcal{L}_\text{cls} + \mathcal{L}_\text{depth} + \mathcal{L}_\text{seg}

    where:

    • Lmask\mathcal{L}_\text{mask} is the binary mask loss (cross-entropy and Dice loss) computed on matched query-ground truth pairs,
    • Lcls\mathcal{L}_\text{cls} is the classification cross-entropy loss over CC semantic categories for the matched queries,
    • Ldepth\mathcal{L}_\text{depth} is the depth supervision loss between estimated image depth maps and projected LiDAR ground-truth depth for the 2D-to-3D view transformation,
    • Lseg\mathcal{L}_\text{seg} is the binary cross-entropy segmentation loss supervising the coarse voxel occupancy filter in the sparse transformer head.
  6. Knowl 6 — Semantic Occupancy Benchmark Performance on nuScenes-Occupancy

    data/table

    On the nuScenes-Occupancy validation set, SparseOcc was evaluated against dense, TPV, and prior vision-based methods for geometric IoU (occupancy geometry), semantic mIoU across 16 classes, FLOPs, training GPU memory, and inference latency (3D feature processing latency / overall latency).

    Method Input IoU (%) mIoU (%) FLOPs Memory 3D / Overall Latency
    MonoScene Camera 18.4 6.9 - - -
    TPVFormer Camera 15.3 7.8 1132G 20G 0.57s / 0.73s
    OpenOccupancy Camera 19.3 10.3 1716G 19G 0.84s / 1.22s
    C-CONet Camera 20.1 12.8 1810G 21G 2.18s / 2.58s
    SparseOcc (ours) Camera 21.8 14.1 455G 13G 0.19s / 0.25s

    SparseOcc reduces FLOPs by 74.9%74.9\% (from 1810G to 455G) and 3D inference latency by 91.3%91.3\% (from 2.18s to 0.19s) compared to the dense state-of-the-art C-CONet, while improving geometric IoU by +1.7%+1.7\% and semantic mIoU by +1.3%+1.3\%.

  7. Knowl 7 — Semantic Scene Completion Benchmark on SemanticKITTI

    data/table

    On the monocular SemanticKITTI validation set (evaluating geometric IoU and semantic mIoU over 19 classes), SparseOcc achieves accuracy competitive with dense/transformer models while requiring substantially lower computational complexity.

    Method Input Geometric IoU (%) Semantic mIoU (%) FLOPs
    LMSCNet* Camera 28.61 6.70 -
    3DSketch* Camera 33.30 7.50 -
    AICNet* Camera 29.59 8.31 -
    JS3C-Net* Camera 38.98 10.31 -
    MonoScene Camera 36.86 11.08 -
    TPVFormer Camera 35.61 11.36 946G
    OccFormer Camera 36.50 13.46 889G
    SparseOcc (ours) Camera 36.48 13.12 393G

    (Asterisks denote RGB-input baseline variants reported by MonoScene).

    SparseOcc achieves 13.12%13.12\% mIoU, closely matching the 13.46%13.46\% mIoU of OccFormer (which uses heavy 3D Swin Transformer attention), while operating at only 44.2%44.2\% of OccFormer's FLOPs (393G vs. 889G).

  8. Knowl 8 — Ablation of Sparse Completion, Voxel Decoder, and Prediction Head

    empirical result

    Ablation experiments conducted on the SemanticKITTI validation set reveal the impact of architectural choices in SparseOcc:

    1. Sparse Completion Block Kernel Decomposition: A single layer of decomposed orthogonal kernels (3×3×13\times3\times1, 3×1×33\times1\times3, 1×3×11\times3\times1) achieves 36.5%36.5\% IoU and 13.1%13.1\% mIoU, outperforming both no completion (35.5%35.5\% IoU, 12.1%12.1\% mIoU) and a single regular 3×3×33\times3\times3 sparse convolution (35.8%35.8\% IoU, 12.2%12.2\% mIoU). Increasing the number of decomposed convolution layers beyond 1 provides no additional performance gain (36.4%36.4\% IoU, 12.7%12.7\% mIoU at 3 layers).

    2. Sparse Voxel Decoder vs. 3D Attention Decoders: The proposed linear interpolation sparse decoder achieves 36.5%36.5\% IoU and 13.1%13.1\% mIoU using 13.3G GPU memory and 279G FLOPs. This performs comparably to a 6-layer MSDeformAttn3D decoder (36.7%36.7\% IoU, 13.3%13.3\% mIoU) while saving 6.5G GPU memory and 100G FLOPs. It significantly outperforms 3D FPN (34.4%34.4\% IoU, 9.8%9.8\% mIoU).

    3. Segmentation Prediction Head: A dense linear head achieves higher geometric IoU (36.8%36.8\%) but lower semantic mIoU (11.8%11.8\%) due to simple voxel-wise cross-entropy. A dense transformer head achieves 36.2%36.2\% IoU and 13.2%13.2\% mIoU with 19.9G memory and 19.0G FLOPs. The Sparse Transformer Head attains 36.5%36.5\% IoU and 13.1%13.1\% mIoU while reducing memory to 13.3G and head FLOPs to 13.5G.

  9. Knowl 9 — Sensitivity to High-Resolution Input Images and Over-Dense Hallucination

    limitation

    When scaling input image resolution on nuScenes-Occupancy from 704×256704 \times 256 to 1600×9001600 \times 900 using a ResNet-50 backbone, SparseOcc's semantic mIoU improves from 14.1%14.1\% to 14.6%14.6\% due to higher 2D semantic feature fidelity. However, geometric IoU drops from 21.8%21.8\% to 20.4%20.4\%.

    This degradation is attributed to hallucination errors: higher image resolutions generate an over-dense set of 3D sparse features during ray casting, which interacts with the sparse completion diffuser layers to activate empty voxels incorrectly. Deactivating the sparse completion block in the final sparse diffuser layer mitigates this hallucination effect.

Coverage note — None was omitted. All principal contributions, model modules (diffuser, pyramid, transformer head), loss functions, benchmark tables, ablations, and identified limitations are covered.

References

  1. 1.Jens Behley, Martin Garbade, Andres Milioto, Jan Quenzel, Sven Behnke, Cyrill Stachniss, and Jurgen Gall. Semantickitti: A dataset for semantic scene understanding of lidar sequences. In ICCV, 2019.
  2. 2.Jens Behley, Martin Garbade, Andres Milioto, Jan Quenzel, Sven Behnke, Cyrill Stachniss, and Jurgen Gall. Semantickitti: A dataset for semantic scene understanding of lidar sequences. In ICCV, 2019.
  3. 3.Holger Caesar, Varun Bankiti, Alex H Lang, Sourabh Vora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Giancarlo Baldan, and Oscar Beijbom. Nuscenes: A multimodal dataset for autonomous driving. In CVPR, 2020.
  4. 4.Anh-Quan Cao and Raoul de Charette. Monoscene: Monocular 3d semantic scene completion. In CVPR, 2022.
  5. 5.Xiaokang Chen, Kwan-Yee Lin, Chen Qian, Gang Zeng, and Hongsheng Li. 3d sketch-aware semantic scene completion via semi-supervised structure prior. In CVPR, 2020.
  6. 6.Bowen Cheng, Ishan Misra, Alexander G Schwing, Alexander Kirillov, and Rohit Girdhar. Masked-attention mask transformer for universal image segmentation. In CVPR, 2022.
  7. 7.Spconv Contributors. Spconv: Spatially sparse convolution library. https://github.com/traveller59/spconv, 2022.
  8. 8.Chongjian Ge, Junsong Chen, Enze Xie, Zhongdao Wang, Lanqing Hong, Huchuan Lu, Zhenguo Li, and Ping Luo. Metabev: Solving sensor failures for 3d detection and map segmentation. In ICCV, 2023.
  9. 9.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In CVPR, 2016.
  10. 10.Anthony Hu, Zak Murez, Nikhil Mohan, Sofıa Dudas, Jeffrey Hawke, Vijay Badrinarayanan, Roberto Cipolla, and Alex Kendall. Fiery: future instance prediction in bird’seye view from surround monocular cameras. In ICCV, 2021.
  11. 11.Yihan Hu, Jiazhi Yang, Li Chen, Keyu Li, Chonghao Sima, Xizhou Zhu, Siqi Chai, Senyao Du, Tianwei Lin, Wenhai Wang, et al. Planning-oriented autonomous driving. In CVPR, 2023.
  12. 12.Junjie Huang, Guan Huang, Zheng Zhu, and Dalong Du. Bevdet: High-performance multi-camera 3d object detection in bird-eye-view. arXiv:2112.11790, 2021.
  13. 13.Yuanhui Huang, Wenzhao Zheng, Yunpeng Zhang, Jie Zhou, and Jiwen Lu. Tri-perspective view for vision-based 3d semantic occupancy prediction. arXiv:2302.07817, 2023.
  14. 14.Xiaosong Jia, Yulu Gao, Li Chen, Junchi Yan, Patrick Langechuan Liu, and Hongyang Li. Driveadapter: Breaking the coupling barrier of perception and planning in end-to-end autonomous driving. In CVPR, 2023.
  15. 15.Alex H Lang, Sourabh Vora, Holger Caesar, Lubing Zhou, Jiong Yang, and Oscar Beijbom. Pointpillars: Fast encoders for object detection from point clouds. In CVPR, 2019.
  16. 16.Jie Li, Kai Han, Peng Wang, Yu Liu, and Xia Yuan. Anisotropic convolutional networks for 3d semantic scene completion. In CVPR, 2020.
  17. 17.Yinhao Li, Zheng Ge, Guanyi Yu, Jinrong Yang, Zengran Wang, Yukang Shi, Jianjian Sun, and Zeming Li. Bevdepth: Acquisition of reliable depth for multi-view 3d object detection. arXiv:2206.10092, 2022.
  18. 18.Yiming Li, Zhiding Yu, Christopher Choy, Chaowei Xiao, Jose M Alvarez, Sanja Fidler, Chen Feng, and Anima Anandkumar. Voxformer: Sparse voxel transformer for camerabased 3d semantic scene completion. In CVPR, 2023.
  19. 19.Zhiqi Li, Wenhai Wang, Hongyang Li, Enze Xie, Chonghao Sima, Tong Lu, Qiao Yu, and Jifeng Dai. Bevformer: Learning bird’s-eye-view representation from multi-camera images via spatiotemporal transformers. arXiv:2203.17270, 2022.
  20. 20.Tsung-Yi Lin, Piotr Dollar, Ross Girshick, Kaiming He, Bharath Hariharan, and Serge Belongie. Feature pyramid networks for object detection. In CVPR, 2017.
  21. 21.Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. In ICCV, 2021.
  22. 22.Zhijian Liu, Haotian Tang, Alexander Amini, Xinyu Yang, Huizi Mao, Daniela Rus, and Song Han. Bevfusion: Multitask multi-sensor fusion with unified bird’s-eye view representation. arXiv:2205.13542, 2022.
  23. 23.Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. In ICLR, 2019.
  24. 24.Lang Peng, Zhirong Chen, Zhangjie Fu, Pengpeng Liang, and Erkang Cheng. Bevsegformer: Bird’s eye view semantic segmentation from arbitrary camera rigs. arXiv:2203.04050, 2022.
  25. 25.Jonah Philion and Sanja Fidler. Lift, splat, shoot: Encoding images from arbitrary camera rigs by implicitly unprojecting to 3d. In ECCV, 2020.
  26. 26.Thomas Roddick and Roberto Cipolla. Predicting semantic map representations from images using pyramid occupancy networks. In CVPR, 2020.
  27. 27.Luis Roldao, Raoul de Charette, and Anne Verroust-Blondet. Lmscnet: Lightweight multiscale 3d semantic completion. In 3DV, 2020.
  28. 28.Shuran Song, Fisher Yu, Andy Zeng, Angel X Chang, Manolis Savva, and Thomas Funkhouser. Semantic scene completion from a single depth image. In CVPR, 2017.
  29. 29.Mingxing Tan and Quoc Le. Efficientnet: Rethinking model scaling for convolutional neural networks. In ICML, 2019.
  30. 30.Pin Tang, Hai-Ming Xu, and Chao Ma. Prototransfer: Crossmodal prototype transfer for point cloud segmentation. In ICCV, 2023.
  31. 31.Xiaoyu Tian, Tao Jiang, Longfei Yun, Yue Wang, Yilun Wang, and Hang Zhao. Occ3d: A large-scale 3d occupancy prediction benchmark for autonomous driving. arXiv:2304.14365, 2023.
  32. 32.Wenwen Tong, Chonghao Sima, Tai Wang, Li Chen, Silei Wu, Hanming Deng, Yi Gu, Lewei Lu, Ping Luo, Dahua Lin, et al. Scene as occupancy. In ICCV, 2023.
  33. 33.Xiaofeng Wang, Zheng Zhu, Wenbo Xu, Yunpeng Zhang, Yi Wei, Xu Chi, Yun Ye, Dalong Du, Jiwen Lu, and Xingang Wang. Openoccupancy: A large scale benchmark for surrounding semantic occupancy perception. In ICCV, 2023.
  34. 34.Yi Wei, Linqing Zhao, Wenzhao Zheng, Zheng Zhu, Jie Zhou, and Jiwen Lu. Surroundocc: Multi-camera 3d occupancy prediction for autonomous driving. In CVPR, 2023.
  35. 35.Zhaoyang Xia, Youquan Liu, Xin Li, Xinge Zhu, Yuexin Ma, Yikang Li, Yuenan Hou, and Yu Qiao. Scpnet: Semantic scene completion on point cloud. In CVPR, 2023.
  36. 36.Shenghua Xu, Xinyue Cai, Bin Zhao, Li Zhang, Hang Xu, Yanwei Fu, and Xiangyang Xue. Rclane: Relay chain prediction for lane detection. In ECCV, 2022.
  37. 37.Xu Yan, Jiantao Gao, Jie Li, Ruimao Zhang, Zhen Li, Rui Huang, and Shuguang Cui. Sparse single sweep lidar point cloud segmentation via learning contextual shape priors from scene completion. In AAAI, 2021.
  38. 38.Xu Yan, Jiantao Gao, Chaoda Zheng, Chao Zheng, Ruimao Zhang, Shuguang Cui, and Zhen Li. 2dpass: 2d priors assisted semantic segmentation on lidar point clouds. In ECCV, 2022.
  39. 39.Yan Yan, Yuxing Mao, and Bo Li. Second: Sparsely embedded convolutional detection. Sensors, 2018.
  40. 40.Weijia Zhang, Dongnan Liu, Chao Ma, and Weidong Cai. Alleviating foreground sparsity for semi-supervised monocular 3d object detection. In WACV, 2024.
  41. 41.Yunpeng Zhang, Zheng Zhu, and Dalong Du. Occformer: Dual-path transformer for vision-based 3d semantic occupancy prediction. arXiv:2304.05316, 2023.
  42. 42.Brady Zhou and Philipp Krahenbuhl. Cross-view transform- ers for real-time map-view semantic segmentation. In CVPR, 2022.
  43. 43.Shengchao Zhou, Weizhou Liu, Chen Hu, Shuchang Zhou, and Chao Ma. Unidistill: A universal cross-modality knowledge distillation framework for 3d object detection in bird’seye view. In CVPR, 2023.
  44. 44.Xizhou Zhu, Weijie Su, Lewei Lu, Bin Li, Xiaogang Wang, and Jifeng Dai. Deformable detr: Deformable transformers for end-to-end object detection. In ICLR, 2020.
  45. 45.Xizhou Zhu, Weijie Su, Lewei Lu, Bin Li, Xiaogang Wang, and Jifeng Dai. Deformable DETR: Deformable transformers for end-to-end object detection. In ICLR, 2021.
  46. 46.Xinge Zhu, Hui Zhou, Tai Wang, Fangzhou Hong, Yuexin Ma, Wei Li, Hongsheng Li, and Dahua Lin. Cylindrical and asymmetrical 3d convolution networks for lidar segmentation. In CVPR, 2021.
  47. 47.Sicheng Zuo, Wenzhao Zheng, Yuanhui Huang, Jie Zhou, and Jiwen Lu. Pointocc: Cylindrical tri-perspective view for point-based 3d semantic occupancy prediction. arXiv:2308.16896, 2023.

Citation

MLA
Tang, P., et al. “SparseOcc: Rethinking Sparse Latent Representation for Vision-Based Semantic Occupancy Prediction”. IEEE Conference on Computer Vision and Pattern Recognition 2024 (CVPR 2024), 2024, http://arxiv.org/abs/2404.09502v1.
APA
Tang, P., Wang, Z., Wang, G., Zheng, J., Ren, X., Feng, B., & Ma, C. (2024). SparseOcc: Rethinking Sparse Latent Representation for Vision-Based Semantic Occupancy Prediction. IEEE Conference on Computer Vision and Pattern Recognition 2024 (CVPR 2024). http://arxiv.org/abs/2404.09502v1
Chicago
Tang, P., Z. Wang, G. Wang, et al. 2024. “SparseOcc: Rethinking Sparse Latent Representation for Vision-Based Semantic Occupancy Prediction”. IEEE Conference on Computer Vision and Pattern Recognition 2024 (CVPR 2024). http://arxiv.org/abs/2404.09502v1.
Harvard
Tang, P. et al. (2024) “SparseOcc: Rethinking Sparse Latent Representation for Vision-Based Semantic Occupancy Prediction”, IEEE Conference on Computer Vision and Pattern Recognition 2024 (CVPR 2024) [Preprint]. Available at: http://arxiv.org/abs/2404.09502v1.
Vancouver
1. Tang P, Wang Z, Wang G, Zheng J, Ren X, Feng B, Ma C (2024) SparseOcc: Rethinking Sparse Latent Representation for Vision-Based Semantic Occupancy Prediction. IEEE Conference on Computer Vision and Pattern Recognition 2024 (CVPR 2024)

BibTeX

@article{tang2024sparseocc,
  title = {SparseOcc: Rethinking Sparse Latent Representation for Vision-Based Semantic Occupancy Prediction},
  author = {Tang, Pin and Wang, Zhongdao and Wang, Guoqing and Zheng, Jilai and Ren, Xiangxuan and Feng, Bailan and Ma, Chao},
  year = {2024},
  journal = {IEEE Conference on Computer Vision and Pattern Recognition 2024 (CVPR 2024)},
  url = {http://arxiv.org/abs/2404.09502v1},
  eprint = {2404.09502}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE