Semantic Scene Completion from a Single Depth Image
Shuran SongFisher YuAndy ZengAngel X. ChangManolis SavvaThomas Funkhouser
Introduces an end-to-end 3D convolutional network that jointly predicts complete volumetric occupancy and semantic labels from a single depth image, alongside the large-scale SUNCG dataset for learning 3D contextual scene completion.
Autonomous systems and robotics require an accurate understanding of 3D physical spaces to navigate, avoid obstacles, and interact with objects. Existing computer vision methods typically observe only visible surfaces or predict spatial geometry without identifying object categories. The article introduces a unified framework to solve "semantic scene completion" by simultaneously estimating physical 3D volumetric occupancy—including occluded spaces behind visible objects—and identifying semantic object categories directly from a single depth image.
The researchers developed the Semantic Scene Completion Network (SSCNet), an end-to-end 3D deep convolutional neural network. To train and evaluate the network, the authors created SUNCG, a large-scale synthetic dataset consisting of over 45,600 manually designed indoor 3D environments containing more than 400,000 rooms with dense, voxel-level ground truth annotations. The network converts 2D depth maps into 3D volumetric representations using a modified flipped truncated signed distance function, applies dilated 3D convolutions to expand the spatial receptive field, and aggregates multi-scale contextual features across different object sizes.
The findings show that jointly predicting geometry and semantic labels significantly outperforms performing each task in isolation. On real-world NYU depth benchmarks, SSCNet achieved an average semantic scene completion intersection-over-union score of 30.5%, outperforming prior model-fitting approaches (19.6%) and bounding-box methods (12.0%). Pre-training on the synthetic SUNCG dataset produced an absolute 10.3% performance boost on real data over training on real data alone. Furthermore, SSCNet generated predictions in approximately 7 seconds per image, which is roughly 18 times faster than previous 3D mesh model-fitting approaches that require 127 seconds.
These results demonstrate that synthetic 3D datasets can effectively bridge the data shortage required to train deep 3D networks for real-world robotics. Furthermore, the substantial runtime reduction makes unified 3D spatial reasoning much more feasible for automated systems, reducing the compute and latency constraints typically associated with complex geometric model fitting.
Organizations developing spatial computing, robotics, or computer vision applications should adopt joint completion-and-segmentation architectures and leverage large-scale synthetic training data to reduce physical data collection costs. However, current limitations include the absence of color (RGB) inputs—which causes difficulty in identifying transparent or depth-missing objects like windows—and lower output voxel resolutions due to GPU memory limits, which can obscure small objects. Future developments should incorporate multi-modal color data and test higher-resolution spatial representations.
- Paper: 3D ShapeNets: A deep representation for volumetric shapes, Zhirong Wu et al. (2014). This paper introduced deep volumetric representations and shape completion from 2.5D depth observations, providing the conceptual foundation for volumetric neural representations in SSCNet.
- Paper: Multi-Scale Context Aggregation by Dilated Convolutions, Fisher Yu et al. (2016). This work pioneered dilated convolutions for expanding receptive fields and aggregating multi-scale context, which SSCNet directly adapts into its 3D dilated context module.
- Paper: SUN RGB-D: A RGB-D scene understanding benchmark suite, Shuran Song et al. (2015). This paper established the standard RGB-D indoor scene understanding dataset and 3D bounding evaluation tasks that motivate full 3D semantic scene completion.
- Paper: Indoor Segmentation and Support Inference from RGBD Images, Nathan Silberman et al. (2012). This paper introduced the NYU Depth v2 dataset and established foundational tasks for indoor segmentation and geometric reasoning from single depth images.
- Paper: Fully convolutional networks for semantic segmentation, Jonathan Long et al. (2015). This work pioneered end-to-end fully convolutional networks for dense semantic prediction that SSCNet generalizes from 2D pixel grids to 3D voxel spaces.
- Paper: Learning Rich Features from RGB-D Images for Object Detection and Segmentation, Saurabh Gupta et al. (2014). This work developed geocentric surface encoding and feature extraction from RGB-D images for 3D indoor scene understanding.
- Paper: A volumetric method for building complex models from range images, Brian Curless et al. (1996). This seminal paper formulated voxelized signed distance fields and volumetric integration for 3D range images, defining the representation used to voxelize depth frustums.
- Paper: 3D Semantic Segmentation with Submanifold Sparse Convolutional Networks, Benjamin Graham et al. (2017). This paper introduces submanifold sparse convolutions to overcome the memory and computational bottlenecks of dense 3D voxel convolutions used in networks like SSCNet.
- Paper: OctNet: Learning Deep 3D Representations at High Resolutions, Gernot Riegler et al. (2016). This work introduces hybrid grid-octree structures to scale deep 3D convolutional architectures to much higher spatial resolutions than dense volumetric grids allow.
- Paper: ScanNet: Richly-Annotated 3D Reconstructions of Indoor Scenes, Angela Dai et al. (2017). This work provides a massive benchmark of richly annotated real-world 3D indoor scenes to train and evaluate volumetric semantic understanding networks beyond synthetic data.
- Paper: SemanticKITTI: A Dataset for Semantic Scene Understanding of LiDAR Sequences, Jens Behley et al. (2019). This paper establishes a large-scale outdoor LiDAR dataset and formalizes sequential outdoor semantic scene completion inspired by indoor SSCNet formulations.
- Paper: Matterport3D: Learning from RGB-D Data in Indoor Environments, Angel Chang et al. (2017). This work advances indoor 3D understanding with large-scale real-world RGB-D building scans and benchmarks dense semantic voxel labeling.
- Paper: 4D Spatio-Temporal ConvNets: Minkowski Convolutional Neural Networks, Christopher Choy et al. (2019). This paper generalizes 3D volumetric convolutions to 4D spatio-temporal sparse tensor convolutions for understanding dynamic 3D scenes.
- Paper: Occupancy Networks: Learning 3D Reconstruction in Function Space, Lars Mescheder et al. (2018). This work addresses the resolution and memory limitations of discrete voxel grids by learning continuous implicit occupancy representations.
