Learning Local Displacements for Point Cloud Completion
Yida WangDavid Joseph TanNassir NavabFederico Tombari
Proposes a point cloud completion framework that preserves fine geometric details across objects and complex indoor scenes through displacement-based feature extraction and an activation-driven neighbor-pooling operation.
Autonomous navigation, robotic interaction, and spatial computing rely heavily on accurate 3D scene understanding. However, sensor scans captured from single viewpoints frequently miss substantial geometric data due to self-occlusion and surrounding obstacles. While existing geometric completion techniques attempt to predict missing 3D structures, they struggle with maintaining high surface resolution, preserving critical local geometric details, or expanding effectively from isolated objects to complex, semantically labeled scenes.
The article develops and evaluates a deep learning framework designed to reconstruct complete 3D shapes and indoor scenes from partial point cloud scans while predicting semantic labels. It introduces three core algorithmic operators—local displacement-based feature extraction, neighbor pooling, and progressive upsampling—integrated into both a direct encoder-decoder model and a transformer-based model.
To establish credibility across multiple domains, the researchers conducted benchmark evaluations on three standard datasets: single-object completion on ShapeNet (covering eight object categories) and full indoor semantic scene completion on the real-world Kinect-acquired NYU dataset and CompleteScanNet (with over 45,000 paired scans). The models process initial partial inputs of 2,048 3D points and generate detailed reconstructions scaled up to 16,384 points, utilizing a specialized ordering loss that progressively guides reconstruction from visible surfaces to occluded regions.
The experimental findings demonstrate state-of-the-art performance across all benchmarks. First, on object completion, the proposed transformer-based architecture outperformed prior methods, achieving a top F-Score of 0.816 and reducing Chamfer reconstruction error to 6.64. Second, the direct encoder-decoder model alone surpassed most existing methods without requiring complex attention layers, validating the raw strength of the underlying operators. Third, the system demonstrated the first dedicated point cloud completion for complex indoor scenes, reaching an average Chamfer distance of 3.04 on CompleteScanNet (outperforming prior baselines by over 25%) and achieving a competitive 42.4% intersection-over-union on NYU semantic completion. Finally, ablation analyses confirmed that combining the proposed neighbor-pooling tokenization with progressive coarse-to-fine upsampling consistently generated superior structural outlines compared to standard farthest point sampling.
These results provide a pathway to significantly enhance spatial awareness in automated robotics and computer vision systems. By operating directly on point clouds rather than computationally heavy volumetric voxel grids, this approach avoids rigid resolution caps and lowers memory consumption while capturing fine structural details like thin edges and small components. This balance reduces operational collision risks in robotics and accelerates spatial inference pipelines.
Stakeholders in automated systems and 3D computer vision should consider integrating displacement-based point cloud operations and neighbor pooling into their 3D perception pipelines. For deployment, teams should leverage the transformer variant where maximum geometric precision is essential, or adopt the lightweight direct encoder-decoder configuration to balance computational efficiency on resource-constrained platforms. Future engineering work should focus on validating the approach in real-time embedded environments and testing performance against dynamic outdoor environments with moving obstacles.
Confidence in these findings is high for standard indoor geometries and structured synthetic objects, supported by extensive cross-dataset comparisons. However, limitations remain when inputs present severely sparse data or highly irregular structures (such as vehicles missing foundational parts), where reconstruction errors can still occur. Additionally, evaluation on standard volumetric benchmarks introduces slight performance trade-offs due to point-to-voxel format conversions.
- Paper: PointNet: Deep Learning on Point Sets for 3D Classification and Segmentation, Charles R. Qi et al. (2017). PointNet provides the foundational permutation-invariant deep learning architecture for directly processing unordered 3D point sets that underlies modern point-based neural networks.
- Paper: A Point Set Generation Network for 3D Object Reconstruction from a Single Image, Haoqiang Fan et al. (2017). This paper establishes the foundational paradigm and loss formulations (such as Chamfer and Earth Mover's distances) for directly generating and reconstructing 3D point clouds using deep neural networks.
- Paper: Semantic Scene Completion from a Single Depth Image, Shuran Song et al. (2016). This work introduces the task of 3D semantic scene completion from partial depth observations, establishing the benchmark problem setting and evaluation metrics extended to point clouds by the source.
- Paper: Learning Representations and Generative Models for 3D Point Clouds, Panos Achlioptas et al. (2017). It introduces deep autoencoder architectures and Chamfer-distance-based reconstruction metrics for 3D point cloud generation, which directly inform the encoder-decoder design of the source.
- Paper: PCT: Point cloud transformer, Meng-Hao Guo et al. (2020). PCT develops transformer architectures and offset-attention mechanisms for irregular point clouds, providing the direct architectural foundation for transformer-based point cloud completion.
- Paper: Point Transformer, Nico Engel et al. (2020). Point Transformer demonstrates local-to-global attention mechanisms and neighborhood aggregation on unordered point sets, establishing key concepts for point transformer architectures.
- Paper: PointConv: Deep Convolutional Networks on 3D Point Clouds, Wenxuan Wu et al. (2018). PointConv introduces continuous convolutions and density handling on raw point sets, paving the way for local displacement-based point operations.
- Paper: RandLA-Net: Efficient Semantic Segmentation of Large-Scale Point Clouds, Qingyong Hu et al. (2019). RandLA-Net develops efficient local spatial encoding and attentive pooling modules for point clouds, which directly relate to the neighbor-pooling tokenization used in the source.
- Paper: Deep Learning for 3D Point Clouds: A Survey, Yulan Guo et al. (2019). This comprehensive survey categorizes point-based versus volumetric representations, highlighting the efficiency advantages of native point cloud processing over voxel grids.
- Paper: Surface Reconstruction from Point Clouds by Learning Predictive Context Priors, Baorui Ma et al. (2022). This paper advances point cloud surface reconstruction by learning predictive context priors and query displacements to reconstruct continuous signed distance fields.
- Paper: Geometry-Consistent Neural Shape Representation with Implicit Displacement Fields, Yifan Wang et al. (2022). It extends displacement-based geometric representations to neural implicit surfaces by decoupling low-frequency base shapes from high-frequency normal displacement fields.
- Paper: PointNeXt: Revisiting PointNet++ with Improved Training and Scaling Strategies, Guocheng Qian et al. (2022). PointNeXt investigates modernized scaling, residual architectures, and inverted bottleneck designs for point cloud processing pipelines.
- Paper: Panoptic Lifting for 3D Scene Understanding with Neural Fields, Yawar Siddiqui et al. (2023). This work explores complete 3D scene understanding by lifting multi-view segmentations into unified, view-consistent 3D volumetric panoptic representations.
- Paper: VolRecon: Volume Rendering of Signed Ray Distance Functions for Generalizable Multi-View Reconstruction, Yufan Ren et al. (2023). VolRecon expands 3D reconstruction from partial multi-view inputs by integrating local ray transformers with global geometric priors.
- Paper: DUSt3R: Geometric 3D Vision Made Easy, Shuzhe Wang et al. (2023). DUSt3R demonstrates uncalibrated dense 3D pointmap regression across multi-view imagery using unified transformer architectures.
- Paper: VGGT: Visual Geometry Grounded Transformer, Jianyuan Wang et al. (2025). VGGT scales transformer-based feed-forward 3D geometry prediction across complex indoor and outdoor environments directly from input images.
