Learning Local Displacements for Point Cloud Completion

Yida WangDavid Joseph TanNassir NavabFederico Tombari

article2022CVPR51 citations

Proposes a point cloud completion framework that preserves fine geometric details across objects and complex indoor scenes through displacement-based feature extraction and an activation-driven neighbor-pooling operation.

Listen

Autonomous navigation, robotic interaction, and spatial computing rely heavily on accurate 3D scene understanding. However, sensor scans captured from single viewpoints frequently miss substantial geometric data due to self-occlusion and surrounding obstacles. While existing geometric completion techniques attempt to predict missing 3D structures, they struggle with maintaining high surface resolution, preserving critical local geometric details, or expanding effectively from isolated objects to complex, semantically labeled scenes.

The article develops and evaluates a deep learning framework designed to reconstruct complete 3D shapes and indoor scenes from partial point cloud scans while predicting semantic labels. It introduces three core algorithmic operators—local displacement-based feature extraction, neighbor pooling, and progressive upsampling—integrated into both a direct encoder-decoder model and a transformer-based model.

To establish credibility across multiple domains, the researchers conducted benchmark evaluations on three standard datasets: single-object completion on ShapeNet (covering eight object categories) and full indoor semantic scene completion on the real-world Kinect-acquired NYU dataset and CompleteScanNet (with over 45,000 paired scans). The models process initial partial inputs of 2,048 3D points and generate detailed reconstructions scaled up to 16,384 points, utilizing a specialized ordering loss that progressively guides reconstruction from visible surfaces to occluded regions.

The experimental findings demonstrate state-of-the-art performance across all benchmarks. First, on object completion, the proposed transformer-based architecture outperformed prior methods, achieving a top F-Score of 0.816 and reducing Chamfer reconstruction error to 6.64. Second, the direct encoder-decoder model alone surpassed most existing methods without requiring complex attention layers, validating the raw strength of the underlying operators. Third, the system demonstrated the first dedicated point cloud completion for complex indoor scenes, reaching an average Chamfer distance of 3.04 on CompleteScanNet (outperforming prior baselines by over 25%) and achieving a competitive 42.4% intersection-over-union on NYU semantic completion. Finally, ablation analyses confirmed that combining the proposed neighbor-pooling tokenization with progressive coarse-to-fine upsampling consistently generated superior structural outlines compared to standard farthest point sampling.

These results provide a pathway to significantly enhance spatial awareness in automated robotics and computer vision systems. By operating directly on point clouds rather than computationally heavy volumetric voxel grids, this approach avoids rigid resolution caps and lowers memory consumption while capturing fine structural details like thin edges and small components. This balance reduces operational collision risks in robotics and accelerates spatial inference pipelines.

Stakeholders in automated systems and 3D computer vision should consider integrating displacement-based point cloud operations and neighbor pooling into their 3D perception pipelines. For deployment, teams should leverage the transformer variant where maximum geometric precision is essential, or adopt the lightweight direct encoder-decoder configuration to balance computational efficiency on resource-constrained platforms. Future engineering work should focus on validating the approach in real-time embedded environments and testing performance against dynamic outdoor environments with moving obstacles.

Confidence in these findings is high for standard indoor geometries and structured synthetic objects, supported by extensive cross-dataset comparisons. However, limitations remain when inputs present severely sparse data or highly irregular structures (such as vehicles missing foundational parts), where reconstruction errors can still occur. Additionally, evaluation on standard volumetric benchmarks introduces slight performance trade-offs due to point-to-voxel format conversions.

  • Paper: PointNet: Deep Learning on Point Sets for 3D Classification and Segmentation, Charles R. Qi et al. (2017). PointNet provides the foundational permutation-invariant deep learning architecture for directly processing unordered 3D point sets that underlies modern point-based neural networks.
  • Paper: A Point Set Generation Network for 3D Object Reconstruction from a Single Image, Haoqiang Fan et al. (2017). This paper establishes the foundational paradigm and loss formulations (such as Chamfer and Earth Mover's distances) for directly generating and reconstructing 3D point clouds using deep neural networks.
  • Paper: Semantic Scene Completion from a Single Depth Image, Shuran Song et al. (2016). This work introduces the task of 3D semantic scene completion from partial depth observations, establishing the benchmark problem setting and evaluation metrics extended to point clouds by the source.
  • Paper: Learning Representations and Generative Models for 3D Point Clouds, Panos Achlioptas et al. (2017). It introduces deep autoencoder architectures and Chamfer-distance-based reconstruction metrics for 3D point cloud generation, which directly inform the encoder-decoder design of the source.
  • Paper: PCT: Point cloud transformer, Meng-Hao Guo et al. (2020). PCT develops transformer architectures and offset-attention mechanisms for irregular point clouds, providing the direct architectural foundation for transformer-based point cloud completion.
  • Paper: Point Transformer, Nico Engel et al. (2020). Point Transformer demonstrates local-to-global attention mechanisms and neighborhood aggregation on unordered point sets, establishing key concepts for point transformer architectures.
  • Paper: PointConv: Deep Convolutional Networks on 3D Point Clouds, Wenxuan Wu et al. (2018). PointConv introduces continuous convolutions and density handling on raw point sets, paving the way for local displacement-based point operations.
  • Paper: RandLA-Net: Efficient Semantic Segmentation of Large-Scale Point Clouds, Qingyong Hu et al. (2019). RandLA-Net develops efficient local spatial encoding and attentive pooling modules for point clouds, which directly relate to the neighbor-pooling tokenization used in the source.
  • Paper: Deep Learning for 3D Point Clouds: A Survey, Yulan Guo et al. (2019). This comprehensive survey categorizes point-based versus volumetric representations, highlighting the efficiency advantages of native point cloud processing over voxel grids.
Cover for Learning Local Displacements for Point Cloud Completion

Abstract

We propose a novel approach aimed at object and semantic scene completion from a partial scan represented as a 3D point cloud. Our architecture relies on three novel layers that are used successively within an encoder-decoder structure and specifically developed for the task at hand. The first one carries out feature extraction by matching the point features to a set of pre-trained local descriptors. Then, to avoid losing individual descriptors as part of standard operations such as max-pooling, we propose an alternative neighbor-pooling operation that relies on adopting the feature vectors with the highest activations. Finally, upsampling in the decoder modifies our feature extraction in order to increase the output dimension. While this model is already able to achieve competitive results with the state of the art, we further propose a way to increase the versatility of our approach to process point clouds. To this aim, we introduce a second model that assembles our layers within a transformer architecture. We evaluate both architectures on object and indoor scene completion tasks, achieving state-of-the-art performance.

Table of Contents

  • 1. Introduction
  • 2. Related works
  • 3. Operators
  • 3.1. Down-sampling operation
  • 3.2. Up-sampling operation
  • 4. Encoder-decoder architectures
  • 4.1. Direct application
  • 4.2. Transformers
  • 5. Loss functions
  • 6. Experiments
  • 6.1. Object completion
  • 6.2. Semantic scene completion
  • 7. Ablation study
  • 8. Conclusion
  • References

Knowls

  1. Knowl 1 — Displacement-based feature extraction

    model/method

    The paper introduces a learned feature-extraction operator that encodes local metric geometry around every input feature. Let Fin={fa}a=1N⊂RDin\mathcal{F}_{\mathrm{in}}=\{f_a\}_{a=1}^{N}\subset\mathbb{R}^{D_{\mathrm{in}}} be a set of NN feature vectors; in the first layer, the faf_a are 3D point coordinates. For each output channel, the operator learns displacement vectors δi∈RDin\delta_i\in\mathbb{R}^{D_{\mathrm{in}}} and scalar weights σi\sigma_i, with i=1,…,si=1,\ldots,s. The distance from an anchor faf_a to the closest input feature near the displaced location fa+δif_a+\delta_i is

    d(fa,δi)=min⁡f~∈Fin∥(fa+δi)−f~∥2.d(f_a,\delta_i)=\min_{\tilde f\in\mathcal{F}_{\mathrm{in}}}\left\|(f_a+\delta_i)-\tilde f\right\|_2.

    For output channel bb, the local response and anchor projection are

    gb(fa)=∑i=1sσb,itanh⁡(αd(fa,δb,i)+β),hb(fa)=ρbTfa,g_b(f_a)=\sum_{i=1}^{s}\sigma_{b,i}\tanh\left(\frac{\alpha}{d(f_a,\delta_{b,i})+\beta}\right),\qquad h_b(f_a)=\rho_b^{\mathsf T}f_a,

    where ρb∈RDin\rho_b\in\mathbb{R}^{D_{\mathrm{in}}} is trainable and α\alpha and β\beta are numerical-stability constants. The resulting DoutD_{\mathrm{out}}-dimensional feature is

    Fout={[gb(fa)+hb(fa)]b=1Dout}a=1N.\mathcal{F}_{\mathrm{out}}=\left\{\left[g_b(f_a)+h_b(f_a)\right]_{b=1}^{D_{\mathrm{out}}}\right\}_{a=1}^{N}.

    Each output channel has its own displacement pattern and projection. The hyperbolic-tangent response is larger when a learned displaced location is close to an input feature, allowing the operator to represent local geometry using metric distances rather than only feature similarity. Nearest-neighbor search can accelerate the distance computation.

  2. Knowl 2 — Neighbor pooling preserves feature descriptors

    model/method

    Instead of applying element-wise max pooling to the feature set Fout\mathcal{F}_{\mathrm{out}}, the paper selects whole feature vectors with the strongest activations. For a feature vector indexed by aa, its activation is

    Aa=∑b=1Douttanh⁡(∣gb(fa)∣).\mathcal{A}_a=\sum_{b=1}^{D_{\mathrm{out}}}\tanh\left(\left|g_b(f_a)\right|\right).

    Given a pooling factor τ\tau, neighbor pooling retains the N/τN/\tau vectors with the largest values of Aa\mathcal{A}_a, rather than taking a separate maximum for each feature dimension. Consequently, the selected local descriptors remain intact instead of being assembled from different spatial neighbors. This provides permutation-insensitive down-sampling while preserving meaningful local structures for point-cloud completion.

  3. Knowl 3 — Learned up-sampling by replicated feature extraction

    model/method

    The decoder uses an up-sampling operator that applies multiple independently parameterized copies of the displacement-based feature extractor to each input feature. If Fin={fa}a=1N\mathcal{F}_{\mathrm{in}}=\{f_a\}_{a=1}^{N} and the up-sampling factor is NupN_{\mathrm{up}}, the output is

    Fup={[gbu(fa)+hbu(fa)]b=1Dout}a=1, u=1N, Nup,\mathcal{F}_{\mathrm{up}}=\left\{\left[g_b^{u}(f_a)+h_b^{u}(f_a)\right]_{b=1}^{D_{\mathrm{out}}}\right\}_{a=1,\,u=1}^{N,\,N_{\mathrm{up}}},

    where uu indexes the independently parameterized branches, DoutD_{\mathrm{out}} is the output feature dimension, and gbug_b^{u} and hbuh_b^{u} are the displacement response and anchor projection for branch uu. The output contains NupNN_{\mathrm{up}}N feature vectors, so the point or feature count increases iteratively through the decoder rather than collapsing to one vector when the latent representation contains a single feature.

  4. Knowl 4 — Direct encoder-decoder completion architecture

    model/method

    The first proposed completion model is built almost entirely from the three learned operators. Its encoder alternates four feature-extraction layers with neighbor-pooling layers, reducing the number of input points by a factor of 128128. A final conventional max-pooling operation converts the remaining features into one latent vector. The decoder applies a sequence of learned up-sampling operators to expand the latent representation into a fine point-cloud reconstruction containing 16,38416{,}384 points. This architecture demonstrates that the displacement-based feature extraction, descriptor-preserving neighbor pooling, and learned up-sampling are sufficient to form a competitive object- and scene-completion network without relying on a volumetric subnetwork.

  5. Knowl 5 — Geometry-aware transformer architecture

    model/method

    The second model incorporates the proposed operators into a transformer completion pipeline. Before attention, the partial scan is reduced to fit GPU memory. A points-to-tokens block replaces the farthest-point-sampling and MLP preprocessing used by the baseline transformer: it alternates displacement-based feature extraction with neighbor pooling and then adds positional coding derived from the 3D coordinates. The resulting tokens are processed by a geometry-aware transformer encoder and decoder to produce a coarse point cloud. The baseline coarse-to-fine reconstruction is replaced by alternating feature-extraction and learned up-sampling operators, which expand the coarse prediction into the final dense completion. This design uses the proposed operators both to select structurally meaningful tokens and to generate fine geometry.

  6. Knowl 6 — Bidirectional completion and ordering losses

    model/method

    Let Pin\mathcal{P}_{\mathrm{in}} be the observed input point cloud, Pout\mathcal{P}_{\mathrm{out}} the predicted completion, and Pgt\mathcal{P}_{\mathrm{gt}} the ground-truth point cloud. The geometric training objective uses bidirectional nearest-point distances:

    Lout→gt=∑p∈Pout∥p−ϕgt(p)∥2,\mathcal{L}_{\mathrm{out}\rightarrow\mathrm{gt}}=\sum_{p\in\mathcal{P}_{\mathrm{out}}}\left\|p-\phi_{\mathrm{gt}}(p)\right\|_2, Lgt→out=∑p∈Pgt∥p−ϕout(p)∥2,\mathcal{L}_{\mathrm{gt}\rightarrow\mathrm{out}}=\sum_{p\in\mathcal{P}_{\mathrm{gt}}}\left\|p-\phi_{\mathrm{out}}(p)\right\|_2,

    where ϕgt(p)\phi_{\mathrm{gt}}(p) is the closest point to pp in Pgt\mathcal{P}_{\mathrm{gt}} and ϕout(p)\phi_{\mathrm{out}}(p) is the closest point to pp in Pout\mathcal{P}_{\mathrm{out}}. The paper additionally exploits the observed behavior that output points tend to acquire an order from observed to occluded geometry. For an input point pp, let θout(p)\theta_{\mathrm{out}}(p) be the index in the ordered output of its closest predicted point. The ordering loss is

    Lorder=∑p∈PinS(θout(p))∥p−ϕout(p)∥2,\mathcal{L}_{\mathrm{order}}=\sum_{p\in\mathcal{P}_{\mathrm{in}}}\mathcal{S}(\theta_{\mathrm{out}}(p))\left\|p-\phi_{\mathrm{out}}(p)\right\|_2,

    with

    S(θ)={1,θ≤∣Pin∣,0,θ>∣Pin∣.\mathcal{S}(\theta)=\begin{cases}1,&\theta\leq |\mathcal{P}_{\mathrm{in}}|,\\0,&\theta>|\mathcal{P}_{\mathrm{in}}|.\end{cases}

    The ordering term encourages the first ∣Pin∣|\mathcal{P}_{\mathrm{in}}| predicted points to reproduce the observed input, leaving later output points to represent occluded regions.

  7. Knowl 7 — Adaptive semantic loss for scene completion

    model/method

    For semantic scene completion, every predicted point has a probability vector li=[li,c]c=1Ncl_i=[l_{i,c}]_{c=1}^{N_c} over NcN_c semantic classes, and l^i\hat l_i is the ground-truth one-hot label assigned through the predicted-to-ground-truth point correspondence. The per-point binary cross-entropy used by the paper is

    ϵi=−1Nc∑c=1Nc[l^i,clog⁡li,c+(1−l^i,c)log⁡(1−li,c)].\epsilon_i=-\frac{1}{N_c}\sum_{c=1}^{N_c}\left[\hat l_{i,c}\log l_{i,c}+(1-\hat l_{i,c})\log(1-l_{i,c})\right].

    The semantic loss is averaged over the points matched to the input:

    Lsemantic=γ∣Pin∣∑i=1∣Pin∣ϵi,\mathcal{L}_{\mathrm{semantic}}=\frac{\gamma}{|\mathcal{P}_{\mathrm{in}}|}\sum_{i=1}^{|\mathcal{P}_{\mathrm{in}}|}\epsilon_i,

    where its adaptive weight is

    γ=0.01Lout→gt+Lgt→out.\gamma=\frac{0.01}{\mathcal{L}_{\mathrm{out}\rightarrow\mathrm{gt}}+\mathcal{L}_{\mathrm{gt}\rightarrow\mathrm{out}}}.

    Thus, the semantic term has relatively little influence while the predicted geometry is unstable and becomes more influential as the geometric completion loss decreases.

  8. Knowl 8 — Object-completion performance on ShapeNet

    data/table

    The object-completion experiment uses ShapeNet partial scans with 2,0482{,}048 input points and evaluates reconstructions with 16,38416{,}384 points over eight categories: plane, cabinet, car, chair, lamp, sofa, table, and vessel. Coordinates are normalized to [−1,1][-1,1]. The reported L2- and L1-Chamfer values are multiplied by 10310^3, while higher F-Score@1% is better.

    Could not parse LaTeX table

    The direct model obtains the best F-Score among the compared methods, while the transformer model has lower Chamfer error and higher F-Score than the direct model. Adding Lorder\mathcal{L}_{\mathrm{order}} improves the direct model from (8.47,8.59,0.788)(8.47,8.59,0.788) to (8.35,8.46,0.801)(8.35,8.46,0.801) and the transformer from (6.74,8.09,0.795)(6.74,8.09,0.795) to (6.64,7.96,0.816)(6.64,7.96,0.816) in the three respective metrics. Qualitative comparisons show that the proposed models preserve both smooth surfaces and fine structures better than grid-deformation methods, although unusual structures such as cars without wheels and scans with too few points, including some chair examples, remain failure cases.

  9. Knowl 9 — Semantic scene-completion performance

    data/table

    The paper evaluates semantic scene completion on NYU and CompleteScanNet. For NYU voxel comparison, the predicted point clouds are converted to the common 60×36×6060\times36\times60 voxel resolution, and performance is mean intersection-over-union (IoU, in percent). For point-cloud scene completion, the input contains 2,0482{,}048 points and the output contains 16,38416{,}384 points; the metric is average L2-Chamfer distance multiplied by 10310^3.

    Could not parse LaTeX table
    Could not parse LaTeX table

    The transformer model achieves the best point-cloud scene-completion Chamfer values on both datasets and reaches 42.4%42.4\% NYU IoU after voxel conversion, exceeding all listed voxel methods except SISNet. Setting the semantic-loss weight to the fixed value γ=1\gamma=1 reduces IoU from 40.0%40.0\% to 37.2%37.2\% for the direct model and from 42.4%42.4\% to 38.9%38.9\% for the transformer, demonstrating the value of adaptive semantic weighting. The voxel results are affected by conversion artifacts and by NYU furniture ground truth represented as solid volumes, whereas the proposed models reconstruct surfaces as point clouds.

  10. Knowl 10 — Component ablation through mix-and-match evaluation

    empirical result

    The paper separately evaluates the transformer backbone, which maps partial points to a coarse point cloud, and the coarse-to-fine module, which expands that coarse prediction. Coarse-to-fine alternatives are grouped as deforming 3D grids, deconvolution/MLP expansion, edge-aware feature expansion (EFE), and the proposed learned feature-extraction plus up-sampling design. Lower values are better; the entries are average Chamfer distances under the corresponding object- and scene-completion evaluations.

    Could not parse LaTeX table
    Could not parse LaTeX table

    For every tested backbone, the proposed coarse-to-fine module gives the lowest value in its row. For every tested coarse-to-fine strategy, the proposed backbone gives the lowest value in its column. The best combinations are the complete proposed architecture, with values 3.043.04 for object completion and 7.967.96 for scene completion, supporting the claim that both the points-to-tokens backbone and the learned coarse-to-fine module contribute independently to the performance.

Coverage note — No substantial contributed method or quantitative result was omitted; qualitative visual comparisons and per-category breakdowns were not made separate knowls because they support the quantitative findings already summarized.

References

  1. 1.Paul J Besl and Neil D McKay. Method for registration of 3-d shapes. In Sensor fusion IV: control paradigms and data structures, volume 1611, pages 586–606. International Society for Optics and Photonics, 1992. 3
  2. 2.Yingjie Cai, Xuesong Chen, Chao Zhang, Kwan-Yee Lin, Xiaogang Wang, and Hongsheng Li. Semantic scene completion via integrating instances and scene in-the-loop. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 324–333, 2021. 1, 7
  3. 3.Angel X Chang, Thomas Funkhouser, Leonidas Guibas, Pat Hanrahan, Qixing Huang, Zimo Li, Silvio Savarese, Manolis Savva, Shuran Song, Hao Su, et al. Shapenet: An information-rich 3d model repository. arXiv preprint arXiv:1512.03012, 2015. 1, 2, 6
  4. 4.Xiaokang Chen, Kwan-Yee Lin, Chen Qian, Gang Zeng, and Hongsheng Li. 3d sketch-aware semantic scene completion via semi-supervised structure prior. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4193–4202, 2020. 2, 7
  5. 5.Yang Chen and Gerard Medioni. Object modelling by registration of multiple range images. Image and vision computing, 10(3):145–155, 1992. 3
  6. 6.Angela Dai, Angel X Chang, Manolis Savva, Maciej Halber, Thomas Funkhouser, and Matthias Nießner. Scannet: Richly-annotated 3d reconstructions of indoor scenes. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 5828–5839, 2017. 1, 7, 8
  7. 7.Angela Dai, Charles Ruizhongtai Qi, and Matthias Nießner. Shape completion using 3d-encoder-predictor cnns and shape synthesis. In Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), volume 3, 2017. 1, 2
  8. 8.Angela Dai, Daniel Ritchie, Martin Bokeloh, Scott Reed, Jurgen Sturm, and Matthias Nießner. Scancomplete: Large-scale scene completion and semantic segmentation for 3d scans. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 4578–4587, 2018. 1
  9. 9.Haoqiang Fan, Hao Su, and Leonidas J Guibas. A point set generation network for 3d object reconstruction from a single image. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 605–613, 2017. 5
  10. 10.Andreas Geiger and Chaohui Wang. Joint 3d object and layout inference from a single rgb-d image. In German Conference on Pattern Recognition, pages 183–195. Springer, 2015. 7
  11. 11.Thibault Groueix, Matthew Fisher, Vladimir G. Kim, Bryan C. Russell, and Mathieu Aubry. A papier-maché approach to learning 3d surface generation. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), June 2018. 1, 2, 8
  12. 12.Yuxiao Guo and Xin Tong. View-volume network for semantic scene completion from a single depth image. In Proceedings of the International Joint Conference on Artificial Intelligence (IJCAI). AAAI Press, 2018. 2, 7
  13. 13.Yangyan Li, Rui Bu, Mingchao Sun, Wei Wu, Xinhan Di, and Baoquan Chen. Pointcnn: Convolution on x-transformed points. In Advances in Neural Information Processing Systems, pages 820–830, 2018. 1
  14. 14.Dahua Lin, Sanja Fidler, and Raquel Urtasun. Holistic scene understanding for 3d object detection with rgbd cameras. In Proceedings of the IEEE international conference on computer vision, pages 1417–1424, 2013. 7
  15. 15.Zhi-Hao Lin, Sheng-Yu Huang, and Yu-Chiang Frank Wang. Convolution in the cloud: Learning deformable kernels in 3d graph convolution networks for point cloud analysis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1800–1809, 2020. 3, 5
  16. 16.Minghua Liu, Lu Sheng, Sheng Yang, Jing Shao, and Shi-Min Hu. Morphing and sampling network for dense point cloud completion. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, pages 11596–11603, 2020. 1, 2, 3, 6, 7, 8
  17. 17.Shice Liu, YU HU, Yiming Zeng, Qiankun Tang, Beibei Jin, Yinhe Han, and Xiaowei Li. See and think: Disentangling semantic scene completion. In S. Bengio, H. Wallach, H. Larochelle, K. Grauman, N. Cesa-Bianchi, and R. Garnett, editors, Advances in Neural Information Processing Systems 31, pages 263–274. Curran Associates, Inc., 2018. 7
  18. 18.Lars Mescheder, Michael Oechsle, Michael Niemeyer, Sebastian Nowozin, and Andreas Geiger. Occupancy networks: Learning 3d reconstruction in function space. In Proceedings IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), 2019. 1
  19. 19.Liang Pan. Ecg: Edge-aware point cloud completion with graph convolution. IEEE Robotics and Automation Letters, 5(3):4392–4398, 2020. 7, 8
  20. 20.Liang Pan, Xinyi Chen, Zhongang Cai, Junzhe Zhang, Haiyu Zhao, Shuai Yi, and Ziwei Liu. Variational relational point completion network. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8524–8533, 2021. 1, 2, 3, 6, 7, 8
  21. 21.Jeong Joon Park, Peter Florence, Julian Straub, Richard Newcombe, and Steven Lovegrove. Deepsdf: Learning continuous signed distance functions for shape representation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 165–174, 2019. 1
  22. 22.Charles R Qi, Hao Su, Kaichun Mo, and Leonidas J Guibas. Pointnet: Deep learning on point sets for 3d classification and segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 652–660, 2017. 1, 2
  23. 23.Charles Ruizhongtai Qi, Li Yi, Hao Su, and Leonidas J Guibas. Pointnet++: Deep hierarchical feature learning on point sets in a metric space. In Advances in Neural Information Processing Systems (NIPS), 2017. 1, 2
  24. 24.Dong Wook Shu, Sung Woo Park, and Junseok Kwon. 3d point cloud generative adversarial network based on tree structured graph convolutions. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 3859–3868, 2019. 8
  25. 25.Nathan Silberman, Derek Hoiem, Pushmeet Kohli, and Rob Fergus. Indoor segmentation and support inference from rgbd images. In European conference on computer vision, pages 746–760. Springer, 2012. 2, 7, 8
  26. 26.Vincent Sitzmann, Julien Martel, Alexander Bergman, David Lindell, and Gordon Wetzstein. Implicit neural representations with periodic activation functions. Advances in Neural Information Processing Systems, 33, 2020. 4
  27. 27.Shuran Song, Fisher Yu, Andy Zeng, Angel X Chang, Manolis Savva, and Thomas Funkhouser. Semantic scene completion from a single depth image. In Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR). IEEE, 2017. 1, 2, 7
  28. 28.Matthew Tancik, Pratul P. Srinivasan, Ben Mildenhall, Sara Fridovich-Keil, Nithin Raghavan, Utkarsh Singhal, Ravi Ramamoorthi, Jonathan T. Barron, and Ren Ng. Fourier features let networks learn high frequency functions in low dimensional domains. NeurIPS, 2020. 4
  29. 29.Lyne P Tchapmi, Vineet Kosaraju, Hamid Rezatofighi, Ian Reid, and Silvio Savarese. Topnet: Structural point cloud decoder. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 383–392, 2019. 6, 7
  30. 30.Xiaogang Wang, Marcelo H. Ang Jr. , and Gim Hee Lee. Cascaded refinement network for point cloud completion. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2020. 7
  31. 31.Xiaogang Wang, Marcelo H Ang Jr, and Gim Hee Lee. A self-supervised cascaded refinement network for point cloud completion. arXiv preprint arXiv:2010.08719, 2020. 7
  32. 32.Yida Wang, David Joseph Tan, Nassir Navab, and Federico Tombari. Forknet: Multi-branch volumetric semantic completion from a single depth image. In Proceedings of the IEEE International Conference on Computer Vision, pages 8608–8617, 2019. 1, 2, 7
  33. 33.Yida Wang, David Joseph Tan, Nassir Navab, and Federico Tombari. Softpoolnet: Shape descriptor for point cloud completion and classification. In Andrea Vedaldi, Horst Bischof, Thomas Brox, and Jan-Michael Frahm, editors, Computer Vision – ECCV 2020, pages 70–85, Cham, 2020. Springer International Publishing. 1, 2, 3, 6, 7, 8
  34. 34.Xin Wen, Tianyang Li, Zhizhong Han, and Yu-Shen Liu. Point cloud completion by skip-attention network with hierarchical folding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2020. 2
  35. 35.Xin Wen, Peng Xiang, Zhizhong Han, Yan-Pei Cao, Pengfei Wan, Wen Zheng, and Yu-Shen Liu. Pmp-net: Point cloud completion by learning multi-step point moving paths. arXiv preprint arXiv:2012.03408, 2020. 1, 2
  36. 36.Shun-Cheng Wu, Keisuke Tateno, Nassir Navab, and Federico Tombari. Scfusion: Real-time incremental scene reconstruction with semantic completion. arXiv preprint arXiv:2010.13662, 2020. 2, 7, 8
  37. 37.Wenxuan Wu, Zhongang Qi, and Li Fuxin. Pointconv: Deep convolutional networks on 3d point clouds. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 9621–9630, 2019. 1
  38. 38.Yaqi Xia, Yan Xia, Wei Li, Rui Song, Kailang Cao, and Uwe Stilla. Asfm-net: Asymmetrical siamese feature matching network for point completion. arXiv preprint arXiv:2104.09587, 2021. 2, 7
  39. 39.Haozhe Xie, Hongxun Yao, Shangchen Zhou, Jiageng Mao, Shengping Zhang, and Wenxiu Sun. Grnet: Gridding residual network for dense point cloud completion. In Andrea Vedaldi, Horst Bischof, Thomas Brox, and Jan-Michael Frahm, editors, Computer Vision – ECCV 2020, pages 365–381, Cham, 2020. Springer International Publishing. 1, 2, 7, 8
  40. 40.Qiangeng Xu, Weiyue Wang, Duygu Ceylan, Radomir Mech, and Ulrich Neumann. Disn: Deep implicit surface network for high-quality single-view 3d reconstruction. In H. Wallach, H. Larochelle, A. Beygelzimer, F. d'Alche-Buc, E. Fox, and R. Garnett, editors, Advances in Neural Information Processing Systems 32, pages 492–502. Curran Associates, Inc., 2019. 1
  41. 41.Bo Yang, Stefano Rosa, Andrew Markham, Niki Trigoni, and Hongkai Wen. Dense 3d object reconstruction from a single depth view. IEEE transactions on pattern analysis and machine intelligence, 2018. 1, 2
  42. 42.Bo Yang, Hongkai Wen, Sen Wang, Ronald Clark, Andrew Markham, and Niki Trigoni. 3d object reconstruction from a single depth view with adversarial learning. In Proceedings of the IEEE International Conference on Computer Vision Workshops, pages 679–688, 2017. 1, 2
  43. 43.Yaoqing Yang, Chen Feng, Yiru Shen, and Dong Tian. Foldingnet: Point cloud auto-encoder via deep grid deformation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 206–215, 2018. 1, 2, 3, 6, 7, 8
  44. 44.Xumin Yu, Yongming Rao, Ziyi Wang, Zuyan Liu, Jiwen Lu, and Jie Zhou. Pointr: Diverse point cloud completion with geometry-aware transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 12498–12507, 2021. 1, 2, 4, 5, 6, 7, 8
  45. 45.Wentao Yuan, Tejas Khot, David Held, Christoph Mertz, and Martial Hebert. Pcn: Point completion network. In 2018 International Conference on 3D Vision (3DV), pages 728–737. IEEE, 2018. 1, 2, 3, 6, 7, 8
  46. 46.Pingping Zhang, Wei Liu, Yinjie Lei, Huchuan Lu, and Xiaoyun Yang. Cascaded context pyramid for full-resolution 3d semantic scene completion. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 7801–7810, 2019. 7
  47. 47.Wenxiao Zhang, Qingan Yan, and Chunxia Xiao. Detail preserved point cloud completion via separated feature aggregation. arXiv preprint arXiv:2007.02374, 2020. 7

Citation

MLA
Wang, Y., et al. “Learning Local Displacements for Point Cloud Completion”. arXiv, 2022, http://arxiv.org/abs/2203.16600v1.
APA
Wang, Y., Tan, D. J., Navab, N., & Tombari, F. (2022). Learning Local Displacements for Point Cloud Completion. arXiv. http://arxiv.org/abs/2203.16600v1
Chicago
Wang, Y., D. J. Tan, N. Navab, and F. Tombari. 2022. “Learning Local Displacements for Point Cloud Completion”. arXiv. http://arxiv.org/abs/2203.16600v1.
Harvard
Wang, Y. et al. (2022) “Learning Local Displacements for Point Cloud Completion”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2203.16600v1.
Vancouver
1. Wang Y, Tan DJ, Navab N, Tombari F (2022) Learning Local Displacements for Point Cloud Completion. arXiv

BibTeX

@article{wang2022learning,
  title = {Learning Local Displacements for Point Cloud Completion},
  author = {Wang, Yida and Tan, David Joseph and Navab, Nassir and Tombari, Federico},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2203.16600v1},
  eprint = {2203.16600}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE