SparseOcc: Rethinking Sparse Latent Representation for Vision-Based Semantic Occupancy Prediction
Pin TangZhongdao WangGuoqing WangJilai ZhengXiangxuan RenBailan FengChao Ma
Proposes an efficient vision-based 3D semantic occupancy prediction network that replaces dense volume processing with a purely sparse latent representation, cutting computation by up to 74.9% while improving accuracy by preventing feature hallucinations in empty space.
Vision-based 3D semantic occupancy prediction is critical for autonomous vehicles to understand surrounding geometry and identify dynamic and static objects. However, existing approaches rely on dense 3D representations that incur heavy computational and memory costs, or they compress 3D space into flattened 2D projections like Bird's Eye View, which sacrifices fine-grained spatial accuracy. Because the vast majority of 3D driving scenes are physically empty, these conventional architectures perform redundant calculations over empty space.
The article demonstrates an efficient vision-based occupancy network called SparseOcc, which replaces dense 3D volumes with a purely sparse representation to improve both computational efficiency and prediction accuracy.
The research evaluated SparseOcc on standard autonomous driving benchmarks, specifically nuScenes-Occupancy and SemanticKITTI, using multi-view and monocular camera inputs. The architecture transforms 2D image features into 3D space, stores only the occupied locations, and processes them through three core components: a 3D sparse latent diffuser with spatially decomposed convolutional kernels for scene completion, a multi-scale sparse feature pyramid fused via linear interpolation, and a sparse transformer head that segments occupied voxels while assigning empty space to a single unified token.
The experimental findings highlight substantial efficiency and performance gains. On the nuScenes-Occupancy benchmark, SparseOcc reduced computation floating-point operations by 59.8% to 74.9% and memory usage by 31.6% to 40.9% compared to dense and projected baselines. It reduced 3D inference latency to 0.19 seconds and overall latency to 0.25 seconds, outperforming dense baselines that required over two seconds. Simultaneously, semantic accuracy improved from 12.8% to 14.1% mean Intersection over Union over the top-performing dense model, driven by the sparse structure's inherent ability to avoid false predictions on empty space. On the SemanticKITTI benchmark, the method achieved a competitive 13.12% semantic accuracy while utilizing only 44.2% of the computation of prior transformer baselines.
These results demonstrate that sparse 3D representations eliminate the historic trade-off between computational overhead and geometric fidelity. For autonomous driving programs, adopting sparse processing reduces onboard hardware costs and power consumption, while lowering latency to improve real-time driving safety.
Engineering teams developing vision-based perception systems should adopt SparseOcc as a foundational architecture for occupancy prediction. When deploying the system, developers must balance image resolution and feature density, as excessively dense inputs can cause spatial hallucinations that degrade geometry scores unless diffusion layers are appropriately tuned. Future efforts should focus on validating the model on physical vehicle hardware across varying weather and edge-case operational conditions.
- Paper: 3D Semantic Segmentation with Submanifold Sparse Convolutional Networks, Benjamin Graham et al. (2017). It introduces submanifold sparse convolutions that preserve spatial sparsity across deep layers, providing the foundational sparse 3D operations underlying SparseOcc's architecture.
- Paper: Semantic Scene Completion from a Single Depth Image, Shuran Song et al. (2016). It formulates the problem of semantic scene completion to jointly predict volumetric 3D occupancy and semantic categories, defining the fundamental task addressed by SparseOcc.
- Paper: BEVFormer: Learning Bird's-Eye-View Representation from Multi-Camera Images via Spatiotemporal Transformers, Zhiqi Li et al. (2022). It establishes spatiotemporal transformer mechanisms for multi-camera 3D perception that SparseOcc contrasts against to eliminate dense bird's-eye-view compression.
- Paper: 4D Spatio-Temporal ConvNets: Minkowski Convolutional Neural Networks, Christopher Choy et al. (2019). It generalizes sparse tensor convolutions to higher-dimensional spaces via the Minkowski Engine, providing essential computational techniques for processing sparse 3D representations efficiently.
- Paper: Lift, Splat, Shoot: Encoding Images From Arbitrary Camera Rigs by Implicitly Unprojecting to 3D, Jonah Philion et al. (2020). It introduces the lift-splat paradigm for projecting 2D camera features into 3D coordinate space, serving as a key baseline transformation for vision-based 3D perception.
- Paper: OctNet: Learning Deep 3D Representations at High Resolutions, Gernot Riegler et al. (2016). It demonstrates how exploiting spatial sparsity in 3D data overcomes cubic memory scaling, providing architectural motivation for sparse volumetric representations.
- Paper: Sparse R-CNN: End-to-End Object Detection with Learnable Proposals, Peize Sun et al. (2020). It pioneers fully sparse token-based detection pipelines with learnable proposals, directly inspiring SparseOcc's sparse transformer head.
No sufficiently relevant recommendations were found.
