Neural RGB-D Surface Reconstruction
Dejan AzinovicRicardo Martin-BruallaDan B. GoldmanMatthias NießnerJustus Thies
Presents an implicit surface reconstruction framework that incorporates commodity RGB-D sensor depth into neural radiance fields via truncated signed distance functions and joint camera pose refinement to recover accurate, metric room-scale 3D meshes.
High-quality 3D digital reconstructions of indoor, room-scale environments are increasingly critical for augmented and virtual reality, virtual room planning, teleconferencing, and robotics. However, standard approaches face significant trade-offs: classical 3D scanning relies purely on depth sensors that leave holes in reflective or thin surfaces, while modern view synthesis models excel at generating realistic images but produce noisy, semi-transparent geometric meshes filled with floating artifacts.
The article demonstrates a novel neural surface reconstruction method that combines dense color images and consumer-grade depth measurements to produce clean, highly accurate 3D geometry. It evaluates how replacing density-based neural volumetric representations with an implicit signed distance function, combined with joint camera pose and distortion refinement, improves reconstructed surface quality.
The authors develop a hybrid architecture composed of two neural networks representing scene shape and surface appearance. The method renders images using a differentiable integration technique that concentrates rendering weights directly at the physical surface boundary. The entire pipeline is trained by jointly optimizing depth alignment, free-space emptiness, and color matching, while simultaneously correcting for imperfect camera positions, exposure variations, and lens distortions. The approach was evaluated on both real-world indoor scans and synthetic benchmark datasets.
The key findings show substantial performance improvements over existing classical and learning-based techniques. On benchmark tests, the proposed method achieved the highest overall reconstruction accuracy, reducing geometric surface error to 0.044 Chamfer distance and raising the completeness F-score to 0.924, outperforming traditional depth fusion and depth-augmented neural baselines. Jointly refining camera trajectories cut positional tracking error from 0.033 meters to 0.021 meters and rotational error from 0.571 degrees to 0.144 degrees. Furthermore, the photometric color loss successfully recovered missing structures—such as thin table legs and wire baskets—where physical depth sensors failed, achieving an average geometric accuracy of 11 millimeters in regions lacking depth data compared to 8 millimeters in sensor-covered areas.
These results demonstrate that commodity depth sensors found in standard smartphones and consumer hardware can produce high-fidelity, metric 3D models when properly paired with dense color optimization. By effectively handling sensor noise, tracking drift, and missing depth data, the method significantly reduces the need for expensive, specialized laser scanning hardware in virtual mapping and robotic simulation pipelines.
Organizations aiming to deploy high-precision 3D scanning should adopt implicit surface representations over density fields and incorporate automated pose and lens refinement to resolve alignment errors. However, because the current implementation operates offline—requiring approximately nine hours per scene on an advanced graphics processor—practitioners should implement this method where accuracy is prioritized over speed. Future development should focus on integrating fast voxel-grid structures to accelerate optimization time and adopting localized neural networks to capture fine geometric details across very large facilities.
- Paper: NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis, Ben Mildenhall et al. (2020). It introduces Neural Radiance Fields (NeRF) and volume rendering, which provide the foundational framework that the source adapts for surface reconstruction.
- Paper: NeuS: Learning Neural Implicit Surfaces by Volume Rendering for Multi-view Reconstruction, Peng Wang et al. (2021). It establishes the core methodology for training neural signed distance functions (SDFs) via volume rendering to extract accurate surfaces.
- Paper: Volume Rendering of Neural Implicit Surfaces, Lior Yariv et al. (2021). It defines a theoretical bridge connecting learnable signed distance fields with volume density rendering, which directly informs the source's implicit surface representation.
- Paper: DeepSDF: Learning Continuous Signed Distance Functions for Shape Representation, Jeong Joon Park et al. (2019). It introduces continuous neural signed distance functions for high-quality implicit 3D shape representation used across neural surface reconstruction pipelines.
- Paper: A volumetric method for building complex models from range images, Brian Curless et al. (1996). It establishes the foundational volumetric truncated signed distance function (TSDF) integration technique for reconstructing 3D surfaces from range and depth measurements.
- Paper: HF-NeuS: Improved Surface Reconstruction Using High-Frequency Details, Yiqun Wang et al. (2022). It extends neural SDF surface reconstruction by explicitly decomposing and recovering high-frequency geometric details to prevent over-smoothing.
- Paper: PermutoSDF: Fast Multi-View Reconstruction with Implicit Surfaces Using Permutohedral Lattices, Radu Alexandru Rosu et al. (2023). It builds upon implicit SDF reconstruction by introducing permutohedral lattice multi-resolution hash encodings for drastically faster training and high-fidelity surface extraction.
- Paper: VolRecon: Volume Rendering of Signed Ray Distance Functions for Generalizable Multi-View Reconstruction, Yufan Ren et al. (2023). It advances neural implicit surface volume rendering by predicting signed ray distance functions that generalize across novel, unseen scenes without per-scene retraining.
- Paper: NeuralUDF: Learning Unsigned Distance Fields for Multi-View Reconstruction of Surfaces with Arbitrary Topologies, Xiaoxiao Long et al. (2023). It generalizes neural implicit surface reconstruction from closed, watertight signed distance fields to unsigned distance fields capable of handling open topologies.
- Paper: 2D Gaussian Splatting for Geometrically Accurate Radiance Fields, Binbin Huang et al. (2024). It introduces oriented 2D surface primitives to achieve accurate geometric surface extraction alongside real-time rendering, offering an alternative paradigm to neural SDF rendering.
