3D Semantic Segmentation with Submanifold Sparse Convolutional Networks

Benjamin GrahamMartin EngelckeLaurens van der Maaten

article2017CVPR1,890 citations

Develops sparse convolutional operations that maintain spatial sparsity across network layers, enabling highly efficient and accurate 3D semantic segmentation on point cloud data.

Listen

Modern 3D spatial data—such as point clouds gathered from LiDAR scanners and depth sensors—is inherently sparse, with meaningful measurements occupying only a small fraction of a 3D grid. Standard deep learning techniques designed for dense 2D images scale poorly to 3D spaces because computation and memory requirements grow exponentially with each added dimension. Furthermore, existing sparse convolutional approaches suffer from a "dilation" problem where standard filters rapidly fill empty surrounding space, causing sparsity to vanish after only a few layers and restricting the depth and efficiency of the network.

To overcome these limitations, the article evaluates Submanifold Sparse Convolutional Networks (SSCNs). The core objective is to demonstrate that restricting convolutional outputs strictly to active input coordinates preserves data sparsity throughout deep network architectures, significantly cutting computational overhead while boosting segmentation accuracy on 3D point cloud datasets.

Evaluating this approach, the researchers introduced two complementary operators: sparse convolutions with downsampling to connect separate spatial components, and submanifold sparse convolutions with stride one that maintain identical sparsity patterns across layers. They implemented these using hash tables and rule-book lookups executed as dense matrix multiplications on graphics processing units (GPUs). The framework was tested across multiple network architectures against dense 3D networks, 2D multi-view projections, and traditional feature baselines using the ShapeNet dataset (16 object categories and 50 part labels across 16,881 models) and indoor scene parsing on the NYU Depth v2 dataset.

The findings show substantial performance and efficiency gains across multiple benchmarks. First, SSCNs established a new state of the art on the ShapeNet part-segmentation challenge with an intersection-over-union score of 85.98%, outperforming all prior competitive entries by at least 0.49%. Second, under identical computational budgets (FLOPs), SSCNs outperformed standard 3D dense and 2D multi-view baselines by 6% to 8% in segmentation accuracy. Third, on NYU Depth v2 indoor scene parsing, an SSCN model achieved up to 68.5% pixel accuracy—a 7.0% improvement over standard 2D fully convolutional networks—while reducing required floating-point operations from 28.5 billion down to 4.5 billion (over an 80% reduction) and slashing memory usage by more than 65%.

These results demonstrate that high-resolution 3D point clouds can be processed natively and deeply without discarding spatial structure or incurring prohibitive computational costs. By eliminating unnecessary calculations in empty space, organizations can deploy high-performing 3D perception models on constrained hardware, reducing cloud inference expenses and lowering latency for real-time applications such as robotics, autonomous driving, and augmented reality.

Engineering teams developing 3D computer vision systems should adopt submanifold sparse convolutions when deploying deep networks for sparse point cloud segmentation. To maximize accuracy, architectures should combine submanifold layers with multi-scale downsampling (such as U-Nets or Fully Convolutional Networks) rather than relying strictly on single-scale representations. When applying these models, practitioners should note that input point clouds require voxelization, which imposes a discrete resolution scale. While confidence in the benchmarked segmentation gains is high, teams extending these methods to uncalibrated or open-world environments should conduct pilot evaluations to account for variable sensor noise, density variations, and dynamic scene rotations.

Cover for 3D Semantic Segmentation with Submanifold Sparse Convolutional Networks

Abstract

Convolutional networks are the de-facto standard for analyzing spatio-temporal data such as images, videos, and 3D shapes. Whilst some of this data is naturally dense (e.g., photos), many other data sources are inherently sparse. Examples include 3D point clouds that were obtained using a LiDAR scanner or RGB-D camera. Standard "dense" implementations of convolutional networks are very inefficient when applied on such sparse data. We introduce new sparse convolutional operations that are designed to process spatially-sparse data more efficiently, and use them to develop spatially-sparse convolutional networks. We demonstrate the strong performance of the resulting models, called submanifold sparse convolutional networks (SSCNs), on two tasks involving semantic segmentation of 3D point clouds. In particular, our models outperform all prior state-of-the-art on the test set of a recent semantic segmentation competition.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Spatial sparsity for ConvNets
  • 4 Submanifold Convolutional Networks
  • 4.1 Sparse Convolutional Operations
  • 4.2 Implementation
  • 5 Submanifold FCNs and U-Nets for Semantic Segmentation
  • 6 Experiments
  • 6.1 Dataset
  • 6.2 Details of Experimental Setup
  • 6.3 Baselines
  • 6.4 Results
  • 6.5 Results on Competition Data
  • 6.6 Semantic Segmentation of Scenes
  • 7 Conclusions
  • References

Knowls

  1. Knowl 1 — Submanifold Sparse Convolution

    model/method

    When standard or naive sparse convolutions are applied repeatedly to low-dimensional structures (such as 1D curves or 2D surfaces) embedded in higher-dimensional grids (d≥3d \ge 3), the set of active sites expands rapidly into surrounding empty space (the submanifold dilation problem), causing sparsity to disappear within a few layers.

    Submanifold sparse convolution, denoted SSC(m,n,f)\text{SSC}(m, n, f), addresses this by restricting the spatial support of output activations strictly to the set of active input sites. Given an odd filter size ff, an input feature dimension mm, and an output feature dimension nn in a dd-dimensional space:

    1. The input is zero-padded by (f−1)/2(f - 1) / 2 on each side so that the spatial grid size is preserved.
    2. An output site is active if and only if the site at the corresponding central spatial coordinate in the input is active.
    3. For every active output site yy, its output feature vector hy∈Rnh_y \in \mathbb{R}^n is computed by convolving filter weight matrices Wi∈Rm×nW_i \in \mathbb{R}^{m \times n} (i∈{0,1,…,f−1}di \in \{0, 1, \dots, f-1\}^d) over only the active input sites present in its fdf^d receptive field:

    hy=∑i∈{0,1,…,f−1}d s.t. x=y+i−(f−1)/2 is activehxWih_y = \sum_{i \in \{0, 1, \dots, f-1\}^d \text{ s.t. } x = y + i - (f-1)/2 \text{ is active}} h_x W_i

    Because inactive sites remain inactive across layers, the spatial sparsity pattern and active site count aa remain invariant throughout arbitrary depth until an explicit downsampling or pooling operation is applied.

  2. Knowl 2 — Sparse Convolution Implementation via Hash Tables and Rule Books

    algorithm

    To implement sparse convolutions SC(m,n,f,s)\text{SC}(m, n, f, s) and submanifold sparse convolutions SSC(m,n,f)\text{SSC}(m, n, f) without computing or storing values in empty space, spatial data is decomposed into two data structures:

    1. A hash table mapping integer coordinate tuples of active grid sites to unique row indices in {0,…,a−1}\{0, \dots, a-1\}, where aa is the number of active sites.
    2. A feature matrix X∈Ra×mX \in \mathbb{R}^{a \times m} storing the feature vector of each active site along its corresponding row.

    Computations are organized using a rule book R=(Ri:i∈F)R = (R_i : i \in F), where F={0,1,…,f−1}dF = \{0, 1, \dots, f-1\}^d denotes the spatial filter grid. Each RiR_i is an integer matrix whose rows (j,k)(j, k) record that the jj-th active input site contributes via filter weight offset ii to the kk-th active output site.

    Input: Input hash table HinH_{in}, input feature matrix Xin∈Rain×mX_{in} \in \mathbb{R}^{a_{in} \times m}, filter spatial offset set F={0,…,f−1}dF = \{0, \dots, f-1\}^d, stride ss, filter weight matrices Wi∈Rm×nW_i \in \mathbb{R}^{m \times n} for i∈Fi \in F, convolution mode (SC or SSC)
    Output: Output hash table HoutH_{out}, output feature matrix Xout∈Raout×nX_{out} \in \mathbb{R}^{a_{out} \times n}
    Initialize Hout←empty hash tableH_{out} \leftarrow \text{empty hash table}
    Initialize rule book Ri←empty list for each i∈FR_i \leftarrow \text{empty list for each } i \in F
    aout←0a_{out} \leftarrow 0
    for each active coordinate pxp_x with row index j=Hin(px)j = H_{in}(p_x) do
        for each spatial offset i∈Fi \in F do
            Compute target output coordinate pyp_y such that pxp_x falls into offset ii of pyp_y
            if mode is SSC and py≠pxp_y \ne p_x then
                continue
            if py∉Houtp_y \notin H_{out} then
                Hout(py)←aoutH_{out}(p_y) \leftarrow a_{out}
                k←aoutk \leftarrow a_{out}
                aout←aout+1a_{out} \leftarrow a_{out} + 1
            else
                k←Hout(py)k \leftarrow H_{out}(p_y)
            Append pair (j,k)(j, k) to rule book list RiR_i
    Initialize Xout←zeros(aout,n)X_{out} \leftarrow \text{zeros}(a_{out}, n)
    for each i∈Fi \in F do
        if RiR_i is not empty then
            Extract input row indices JJ and output row indices KK from RiR_i
            Xout[K,:]←Xout[K,:]+Xin[J,:]⋅WiX_{out}[K, :] \leftarrow X_{out}[K, :] + X_{in}[J, :] \cdot W_i
    return Hout,XoutH_{out}, X_{out}

    For SSC layers, Hout=HinH_{out} = H_{in} and the rule book can be constructed once and reused across all subsequent SSC layers until a downsampling, strided, or pooling layer is reached. Constructing hash tables and rule books takes O(a)O(a) time.

  3. Knowl 3 — Sparse Convolution and Associated Sparse Operators

    model/method

    In addition to submanifold sparse convolutions, spatially-sparse networks utilize a suite of sparse operators operating over dd-dimensional grids:

    • Sparse Convolution SC(m,n,f,s)\text{SC}(m, n, f, s): A convolution with filter size ff, stride ss, mm input planes, and nn output planes. An output site is active if at least one active input site lies in its fdf^d receptive field. For an input of spatial size ℓ\ell, the output spatial size is (ℓ−f+s)/s(\ell - f + s) / s. Ground states of inactive sites are assumed to be zero and discarded from computation.
    • Activation Functions: Pointwise nonlinearities (e.g., ReLU) evaluated strictly on the rows of the active feature matrix.
    • Batch Normalization: Standard batch normalization where mean and variance statistics are estimated and applied exclusively across the set of active sites.
    • Max-Pooling MP(f,s)\text{MP}(f, s): Evaluated on receptive fields of size fdf^d with stride ss, setting output active vectors to the elementwise maximum of the active input feature vectors in the receptive field and the zero vector.
    • Average-Pooling AP(f,s)\text{AP}(f, s): Calculates f−df^{-d} times the sum of the active input feature vectors within the fdf^d receptive field.
    • Deconvolution DC(m,n,f,s)\text{DC}(m, n, f, s): The exact topological inverse of SC(m,n,f,s)\text{SC}(m, n, f, s), inverting input-output site connectivity such that output active sites match the input active sites of the forward SC operation.
  4. Knowl 4 — Computational and Memory Costs of Regular vs Sparse vs Submanifold Convolutions

    theoretical result

    For a single spatial location in a dd-dimensional space using a filter of spatial size f=3f = 3, padding stride s=1s = 1, mm input feature channels, and nn output feature channels, the computational cost (FLOPs) and memory consumption depend on whether the target output site is active and the number of active input neighbors a∈[0,3d]a \in [0, 3^d]:

    Target Site Active? Type Regular (C) Sparse (SC) Submanifold Sparse (SSC)
    Yes FLOPs 3dmn3^d m n amna m n amna m n
    Memory nn nn nn
    No, a>0a > 0 FLOPs 3dmn3^d m n amna m n 00
    Memory nn nn 00
    No, a=0a = 0 FLOPs 3dmn3^d m n 00 00
    Memory nn 00 00

    While regular convolution executes 3dmn3^d m n FLOPs regardless of data sparsity, SSC requires zero computation and zero memory at any site that was not already active in the input layer, preventing the active site count from expanding.

  5. Knowl 5 — Submanifold Sparse Network Architectures for 3D Segmentation

    model/method

    Submanifold Sparse Convolutional Networks (SSCNs) process sparse 3D point cloud data via three primary architectural paradigms built from pre-activated SSC(⋅,⋅,3)\text{SSC}(\cdot, \cdot, 3) convolutions (each preceded by Batch Normalization and ReLU):

    1. Single-Scale Architecture (C3): Stacks 2, 4, or 6 pre-activated SSC(⋅,⋅,3)\text{SSC}(\cdot, \cdot, 3) layers with constant spatial resolution and 8, 16, 32, or 64 feature channels.
    2. Submanifold Sparse FCN: Employs an encoder that downsamples spatial scale by a factor of 2 at each level using SC(⋅,⋅,2,2)\text{SC}(\cdot, \cdot, 2, 2) convolutions, doubling the filter count per downsampling step. Convolutional blocks at each level contain 1 to 3 SSC layers or pre-activated residual blocks (each containing two SSC(⋅,⋅,3)\text{SSC}(\cdot, \cdot, 3) layers with identity shortcuts). Feature maps from downsampled scales are restored directly to original resolution via nearest-neighbor upsampling before being linearly combined for final point classification.
    3. Submanifold Sparse U-Net: Follows a symmetric contracting and expanding path with skip connections. The encoder downsamples using SC(⋅,⋅,2,2)\text{SC}(\cdot, \cdot, 2, 2), while the decoder upsamples using sparse deconvolutions DC(⋅,⋅,2,2)\text{DC}(\cdot, \cdot, 2, 2) to match and merge multi-scale feature maps.
  6. Knowl 6 — ShapeNet Part-Segmentation Benchmark Results

    data/table

    Performance of Submanifold SparseConvNet (SSCN) evaluated on the ShapeNet 3D part-segmentation challenge test set consisting of 2,874 point clouds across 16 object categories and 50 part classes, using average Intersection-over-Union (IoU) as the metric:

    Method Average IoU
    NN matching with Chamfer distance 77.57%
    Synchronized Spectral CNN 84.74%
    Pd-Network (extension of Kd-Network) 85.49%
    Densely Connected PointNet 84.32%
    PointCNN 82.29%
    Submanifold SparseConvNet 85.98%

    The SSCN model uses an FCN architecture with scale S=24S = 24, 64 initial filters, three downsampling levels, two residual blocks per resolution, random affine transform data augmentation, and 10-view test averaging, outperforming prior state-of-the-art methods by ≥0.49%\ge 0.49\% average IoU.

  7. Knowl 7 — Segmentation Accuracy vs Computational Cost Trade-Off

    empirical result

    When evaluated on randomly rotated and translated ShapeNet point clouds across varying computational budgets (measured in FLOPs):

    • Submanifold SparseConvNets (SSCNs) consistently outperform Shape Contexts, Dense 3D ConvNets, and 2D Multi-View ConvNets. At a compute budget of 10810^8 FLOPs, SSCNs achieve an average IoU that is 6% to 8% higher than all three baseline families.
    • Among SSCN architectures, multi-scale FCN and U-Net variants outperform single-scale C3 architectures across all FLOP levels because strided downsampling operations expand the effective receptive field and allow information to flow between disconnected components of the input shape.
    • Varying the voxel grid scaling parameter S∈{16,32,48}S \in \{16, 32, 48\} yields comparable IoU at lower FLOP budgets, while larger scales (S=48S = 48, ≈99%\approx 99\% voxel sparsity) achieve slightly superior accuracy at high FLOP budgets.
  8. Knowl 8 — Semantic Scene Segmentation Performance on NYU Depth v2

    empirical result

    On the NYU Depth (v2) dataset (1,449 RGB-D images evaluated across 40 semantic classes), RGB-D frames are converted into 3D point clouds with 4 input channels (normalized RGB in [−1,1][-1, 1] plus an indicator feature set to 1) and processed by SSCN-FCN architectures with 8 downsampling levels:

    Network Multi-view (kk) Accuracy FLOPs Memory
    2D FCN 1 61.5% 28.50G 135.7M
    SSCN-FCN A 1 64.1% 1.09G 5.2M
    SSCN-FCN A 4 66.9% 4.36G 20.7M
    SSCN-FCN B 1 66.4% 4.50G 11.6M
    SSCN-FCN B 4 68.5% 17.90G 46.4M

    SSCN-FCN A (16 initial filters, 1 SSC per level) and SSCN-FCN B (24 initial filters, 2 SSC per level) outperform the 2D FCN baseline by up to 7.0% in pixel accuracy while using fewer FLOPs and less memory.

    Setting all depth coordinates to zero in SSCN-FCN A drops classification accuracy from 64.1% to 50.8% and reduces FLOPs by 60% (due to fewer distinct active voxels), confirming that SSCNs directly exploit 3D spatial geometry.

  9. Knowl 9 — Point Cloud Voxelization and Training Setup for SSCNs

    experimental setup

    Point clouds are preprocessed and trained under the following pipeline:

    1. Voxelization: Each point cloud is centered and rescaled into a bounding sphere of diameter S∈{16,32,48}S \in \{16, 32, 48\}. The sphere is placed randomly within a discrete grid of size 4S4S for sparse networks (and size SS for dense baselines). Voxel values are set by counting the number of points per voxel and normalizing so that occupied voxels have a mean density of 1.
    2. Data Augmentation: Random 3D translations and 3D rotations are applied during training.
    3. Optimization: Trained jointly on all classes using SGD with Nesterov momentum 0.9, weight decay 10−410^{-4}, initial learning rate 0.1 decayed by a factor of e−0.04e^{-0.04} after each epoch, batch size 16, and multi-class negative log-likelihood loss for 100 epochs.
    4. Inference: Test predictions compute softmax probabilities restricted to valid category part classes. In multi-view testing, predictions are averaged over KK random 3D rotations of the input point cloud.

Coverage note — No substantial contributed material was omitted. All core definitions, algorithms, complexity analyses, architectures, and empirical findings on ShapeNet and NYU Depth v2 are fully covered.

References

  1. 1.S. Belongie, J. Malik, and J. Puzicha. Shape Matching and Object Recognition using Shape Contexts. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2002. 6, 8
  2. 2.Ö. Çiçek, A. Abdulkadir, S. S. Lienkamp, T. Brox, and O. Ronneberger. 3D U-Net: Learning Dense Volumetric Segmentation from Sparse Annotation. International Conference on Medical Image Computing and Computer-Assisted Intervention, 2016. 2, 6
  3. 3.M. Engelcke, D. Rao, D. Z. Wang, C. H. Tong, and I. Posner. Vote3Deep: Fast Object Detection in 3D Point Clouds using Efficient Convolutional Neural Networks. IEEE International Conference on Robotics and Automation, 2017. 1, 2, 5
  4. 4.B. Graham. Sparse 3D Convolutional Neural Networks. British Machine Vision Conference, 2015. 1, 2, 4, 5
  5. 5.B. Graham and L. van der Maaten. Submanifold Sparse Convolutional Networks. 2017. https://arxiv.org/abs/1706.01307. 2
  6. 6.S. Gupta, P. Arbelaez, and J. Malik. Perceptual Organization and Recognition of Indoor Scenes from RGB-D Images. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2013. 9
  7. 7.K. He, X. Zhang, S. Ren, and J. Sun. Deep Residual Learning for Image Recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2016. 2
  8. 8.K. He, X. Zhang, S. Ren, and J. Sun. Identity Mappings in Deep Residual Networks. European Conference on Computer Vision, 2016. 3, 6
  9. 9.G. Huang, Z. Liu, K. Q. Weinberger, and L. van der Maaten. Densely Connected Convolutional Networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2016. 3
  10. 10.R. Klokov and V. Lempitsky. Escape from Cells: Deep Kd-Networks for The Recognition of 3D Point Cloud Models. arXiv preprint arXiv:1704.01222, 2017. 2, 3
  11. 11.X. G. L. Yi, H. Su and L. Guibas. SyncSpecCNN: Synchronized Spectral CNN for 3D Shape Segmentation. arXiv preprint arXiv:1612.00606, 2016. 2
  12. 12.Y. LeCun, J. Denker, and S. Solla. Optimal Brain Damage. In Advances in Neural Information Processing Systems, 1990. 4
  13. 13.B. Liu, M. Wang, H. Foroosh, M. Tappen, , and M. Penksy. Sparse Convolutional Neural Networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2015. 4
  14. 14.J. Long, E. Shelhamer, and T. Darrell. Fully Convolutional Networks for Semantic Segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2015. 2, 3, 5, 6, 9, 10
  15. 15.D. Maturana and S. Scherer. VoxNet: A 3D Convolutional Neural Network for Real-Time Object Recognition. IEEE International Conference on Intelligent Robots and Systems, 2015. 2
  16. 16.P. K. Nathan Silberman, Derek Hoiem and R. Fergus. Indoor Segmentation and Support Inference from RGBD Images. European Conference on Computer Vision, 2012. 6, 9, 10
  17. 17.C. R. Qi, H. Su, K. Mo, and L. J. Guibas. PointNet: Deep Learning on Point Sets for 3D Classification and Segmentation. arXiv preprint arXiv:1612.00593, 2016. 2, 3
  18. 18.G. Riegler, A. O. Ulusoys, and A. Geiger. Octnet: Learning Deep 3D Representations at High Resolutions. arXiv preprint arXiv:1611.05009, 2016. 1, 2, 4
  19. 19.O. Ronneberger, P. Fischer, and T. Brox. U-Net: Convolutional Networks for Biomedical Image Segmentation. International Conference on Medical Image Computing and Computer-Assisted Intervention, 2015. 2, 3, 5
  20. 20.K. Simonyan and A. Zisserman. Very Deep Convolutional Networks for Large-Scale Image Recognition. arXiv preprint arXiv:1409.1556, 2014. 2, 3
  21. 21.H. Su, S. Maji, E. Kalogerakis, and E. G. Learned-Miller. Multi-View Convolutional Neural Networks for 3D Shape Recognition. International Conference on Computer Vision, 2015. 6, 8
  22. 22.D. Z. Wang and I. Posner. Voting for voting in online point cloud object detection. Robotics: Science and Systems, 2015. 1, 5
  23. 23.L. Yi, H. Su, L. Shao, M. Savva, et al. Large-Scale 3D Shape Reconstruction and Segmentation from ShapeNet Core55. arXiv preprint arXiv:1710.06104, 2017. 2, 5, 6, 8, 9
  24. 24.F. Yu and V. Koltun. Multi-Scale Context Aggregation by Dilated Convolutions. arXiv preprint arXiv:1511.07122, 2015. 2, 5
  25. 25.M. D. Zeiler, D. Krishnan, G. W. Taylor, and R. Fergus. Deconvolutional Networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2010. 4

Citation

MLA
Graham, B., et al. “3D Semantic Segmentation with Submanifold Sparse Convolutional Networks”. 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2018, pp. 9224–32, https://doi.org/10.1109/CVPR.2018.00961.
APA
Graham, B., Engelcke, M., & Maaten, L. van . der . (2018). 3D Semantic Segmentation with Submanifold Sparse Convolutional Networks. 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, 9224–9232. https://doi.org/10.1109/CVPR.2018.00961
Chicago
Graham, B., M. Engelcke, and L. van . der . Maaten. 2018. “3D Semantic Segmentation with Submanifold Sparse Convolutional Networks”. 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, 9224–32. https://doi.org/10.1109/CVPR.2018.00961.
Harvard
Graham, B., Engelcke, M. and Maaten, L. van . der . (2018) “3D Semantic Segmentation with Submanifold Sparse Convolutional Networks”, 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. IEEE, pp. 9224–9232. Available at: https://doi.org/10.1109/CVPR.2018.00961.
Vancouver
1. Graham B, Engelcke M, Maaten L van der (2018) 3D Semantic Segmentation with Submanifold Sparse Convolutional Networks. In: 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. IEEE, pp 9224–9232

BibTeX

@inproceedings{Graham_2018, title={3D Semantic Segmentation with Submanifold Sparse Convolutional Networks}, url={http://dx.doi.org/10.1109/CVPR.2018.00961}, DOI={10.1109/cvpr.2018.00961}, booktitle={2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition}, publisher={IEEE}, author={Graham, Benjamin and Engelcke, Martin and Maaten, Laurens van der}, year={2018}, month=June, pages={9224–9232} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE