Semantic Scene Completion from a Single Depth Image

Shuran SongFisher YuAndy ZengAngel X. ChangManolis SavvaThomas Funkhouser

article2016CVPR1,441 citations

Introduces an end-to-end 3D convolutional network that jointly predicts complete volumetric occupancy and semantic labels from a single depth image, alongside the large-scale SUNCG dataset for learning 3D contextual scene completion.

Listen

Autonomous systems and robotics require an accurate understanding of 3D physical spaces to navigate, avoid obstacles, and interact with objects. Existing computer vision methods typically observe only visible surfaces or predict spatial geometry without identifying object categories. The article introduces a unified framework to solve "semantic scene completion" by simultaneously estimating physical 3D volumetric occupancy—including occluded spaces behind visible objects—and identifying semantic object categories directly from a single depth image.

The researchers developed the Semantic Scene Completion Network (SSCNet), an end-to-end 3D deep convolutional neural network. To train and evaluate the network, the authors created SUNCG, a large-scale synthetic dataset consisting of over 45,600 manually designed indoor 3D environments containing more than 400,000 rooms with dense, voxel-level ground truth annotations. The network converts 2D depth maps into 3D volumetric representations using a modified flipped truncated signed distance function, applies dilated 3D convolutions to expand the spatial receptive field, and aggregates multi-scale contextual features across different object sizes.

The findings show that jointly predicting geometry and semantic labels significantly outperforms performing each task in isolation. On real-world NYU depth benchmarks, SSCNet achieved an average semantic scene completion intersection-over-union score of 30.5%, outperforming prior model-fitting approaches (19.6%) and bounding-box methods (12.0%). Pre-training on the synthetic SUNCG dataset produced an absolute 10.3% performance boost on real data over training on real data alone. Furthermore, SSCNet generated predictions in approximately 7 seconds per image, which is roughly 18 times faster than previous 3D mesh model-fitting approaches that require 127 seconds.

These results demonstrate that synthetic 3D datasets can effectively bridge the data shortage required to train deep 3D networks for real-world robotics. Furthermore, the substantial runtime reduction makes unified 3D spatial reasoning much more feasible for automated systems, reducing the compute and latency constraints typically associated with complex geometric model fitting.

Organizations developing spatial computing, robotics, or computer vision applications should adopt joint completion-and-segmentation architectures and leverage large-scale synthetic training data to reduce physical data collection costs. However, current limitations include the absence of color (RGB) inputs—which causes difficulty in identifying transparent or depth-missing objects like windows—and lower output voxel resolutions due to GPU memory limits, which can obscure small objects. Future developments should incorporate multi-modal color data and test higher-resolution spatial representations.

Cover for Semantic Scene Completion from a Single Depth Image

Abstract

This paper focuses on semantic scene completion, a task for producing a complete 3D voxel representation of volumetric occupancy and semantic labels for a scene from a single-view depth map observation. Previous work has considered scene completion and semantic labeling of depth maps separately. However, we observe that these two problems are tightly intertwined. To leverage the coupled nature of these two tasks, we introduce the semantic scene completion network (SSCNet), an end-to-end 3D convolutional network that takes a single depth image as input and simultaneously outputs occupancy and semantic labels for all voxels in the camera view frustum. Our network uses a dilation-based 3D context module to efficiently expand the receptive field and enable 3D context learning. To train our network, we construct SUNCG - a manually created large-scale dataset of synthetic 3D scenes with dense volumetric annotations. Our experiments demonstrate that the joint model outperforms methods addressing each task in isolation and outperforms alternative approaches on the semantic scene completion task.

Citation

MLA
Song, S., et al. “Semantic Scene Completion from a Single Depth Image”. arXiv, 2016, http://arxiv.org/abs/1611.08974v1.
APA
Song, S., Yu, F., Zeng, A., Chang, A. X., Savva, M., & Funkhouser, T. (2016). Semantic Scene Completion from a Single Depth Image. arXiv. http://arxiv.org/abs/1611.08974v1
Chicago
Song, S., F. Yu, A. Zeng, A. X. Chang, M. Savva, and T. Funkhouser. 2016. “Semantic Scene Completion from a Single Depth Image”. arXiv. http://arxiv.org/abs/1611.08974v1.
Harvard
Song, S. et al. (2016) “Semantic Scene Completion from a Single Depth Image”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1611.08974v1.
Vancouver
1. Song S, Yu F, Zeng A, Chang AX, Savva M, Funkhouser T (2016) Semantic Scene Completion from a Single Depth Image. arXiv

BibTeX

@article{song2016semantic,
  title = {Semantic Scene Completion from a Single Depth Image},
  author = {Song, Shuran and Yu, Fisher and Zeng, Andy and Chang, Angel X. and Savva, Manolis and Funkhouser, Thomas},
  year = {2016},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1611.08974v1},
  eprint = {1611.08974}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE