Neural RGB-D Surface Reconstruction

Dejan AzinovicRicardo Martin-BruallaDan B. GoldmanMatthias NießnerJustus Thies

article2022CVPR435 citations

Presents an implicit surface reconstruction framework that incorporates commodity RGB-D sensor depth into neural radiance fields via truncated signed distance functions and joint camera pose refinement to recover accurate, metric room-scale 3D meshes.

Listen

High-quality 3D digital reconstructions of indoor, room-scale environments are increasingly critical for augmented and virtual reality, virtual room planning, teleconferencing, and robotics. However, standard approaches face significant trade-offs: classical 3D scanning relies purely on depth sensors that leave holes in reflective or thin surfaces, while modern view synthesis models excel at generating realistic images but produce noisy, semi-transparent geometric meshes filled with floating artifacts.

The article demonstrates a novel neural surface reconstruction method that combines dense color images and consumer-grade depth measurements to produce clean, highly accurate 3D geometry. It evaluates how replacing density-based neural volumetric representations with an implicit signed distance function, combined with joint camera pose and distortion refinement, improves reconstructed surface quality.

The authors develop a hybrid architecture composed of two neural networks representing scene shape and surface appearance. The method renders images using a differentiable integration technique that concentrates rendering weights directly at the physical surface boundary. The entire pipeline is trained by jointly optimizing depth alignment, free-space emptiness, and color matching, while simultaneously correcting for imperfect camera positions, exposure variations, and lens distortions. The approach was evaluated on both real-world indoor scans and synthetic benchmark datasets.

The key findings show substantial performance improvements over existing classical and learning-based techniques. On benchmark tests, the proposed method achieved the highest overall reconstruction accuracy, reducing geometric surface error to 0.044 Chamfer distance and raising the completeness F-score to 0.924, outperforming traditional depth fusion and depth-augmented neural baselines. Jointly refining camera trajectories cut positional tracking error from 0.033 meters to 0.021 meters and rotational error from 0.571 degrees to 0.144 degrees. Furthermore, the photometric color loss successfully recovered missing structures—such as thin table legs and wire baskets—where physical depth sensors failed, achieving an average geometric accuracy of 11 millimeters in regions lacking depth data compared to 8 millimeters in sensor-covered areas.

These results demonstrate that commodity depth sensors found in standard smartphones and consumer hardware can produce high-fidelity, metric 3D models when properly paired with dense color optimization. By effectively handling sensor noise, tracking drift, and missing depth data, the method significantly reduces the need for expensive, specialized laser scanning hardware in virtual mapping and robotic simulation pipelines.

Organizations aiming to deploy high-precision 3D scanning should adopt implicit surface representations over density fields and incorporate automated pose and lens refinement to resolve alignment errors. However, because the current implementation operates offline—requiring approximately nine hours per scene on an advanced graphics processor—practitioners should implement this method where accuracy is prioritized over speed. Future development should focus on integrating fast voxel-grid structures to accelerate optimization time and adopting localized neural networks to capture fine geometric details across very large facilities.

Cover for Neural RGB-D Surface Reconstruction

Abstract

Obtaining high-quality 3D reconstructions of room-scale scenes is of paramount importance for upcoming applications in AR or VR. These range from mixed reality applications for teleconferencing, virtual measuring, virtual room planing, to robotic applications. While current volume-based view synthesis methods that use neural radiance fields (NeRFs) show promising results in reproducing the appearance of an object or scene, they do not reconstruct an actual surface. The volumetric representation of the surface based on densities leads to artifacts when a surface is extracted using Marching Cubes, since during optimization, densities are accumulated along the ray and are not used at a single sample point in isolation. Instead of this volumetric representation of the surface, we propose to represent the surface using an implicit function (truncated signed distance function). We show how to incorporate this representation in the NeRF framework, and extend it to use depth measurements from a commodity RGB-D sensor, such as a Kinect. In addition, we propose a pose and camera refinement technique which improves the overall reconstruction quality. In contrast to concurrent work on integrating depth priors in NeRF which concentrates on novel view synthesis, our approach is able to reconstruct high-quality, metrical 3D reconstructions.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Method
  • 3.1. Hybrid Scene Representation
  • 3.2. Optimization
  • 4. Results
  • 5. Conclusion
  • Acknowledgements
  • References

Knowls

  1. Knowl 1 — Hybrid implicit-surface and radiance-field reconstruction

    model/method

    The method reconstructs a metrically scaled scene from an RGB-D frame sequence by representing geometry as a continuous truncated signed distance function (TSDF) and appearance as a volumetric radiance field. A coordinate-based neural representation is queried at arbitrary 3D locations, and differentiable ray integration is used to reproduce the observed RGB images. The optimization jointly estimates the scene representation, camera parameters, and per-frame appearance corrections from aligned color and depth observations. After optimization, a triangle mesh is extracted from the zero level set of the learned TSDF using Marching Cubes, rather than from an arbitrary density isosurface.

  2. Knowl 2 — Signed-distance-based differentiable volumetric rendering

    equation

    For a ray sampled at KK points, let Di∈[−tr,tr]D_i\in[-t_r,t_r] be the predicted signed distance at sample ii, let ci∈R3\mathbf{c}_i\in\mathbb{R}^3 be its predicted RGB radiance, and let tr>0t_r>0 be the TSDF truncation distance. With σ(x)=1/(1+e−x)\sigma(x)=1/(1+e^{-x}) denoting the logistic sigmoid, the rendering weight is

    wi=σ ⁣(Ditr)σ ⁣(−Ditr).w_i=\sigma\!\left(\frac{D_i}{t_r}\right)\sigma\!\left(-\frac{D_i}{t_r}\right).

    This bell-shaped weight is maximal at the surface zero crossing (Di=0D_i=0) and decreases as the sample moves away from the surface. Samples beyond the first truncation region along a ray are assigned zero weight to avoid integrating across multiple surface intersections. The rendered color C∈R3\mathbf{C}\in\mathbb{R}^3 is the normalized weighted average

    C=1∑i=0K−1wi∑i=0K−1wici.\mathbf{C}=\frac{1}{\sum_{i=0}^{K-1}w_i}\sum_{i=0}^{K-1}w_i\mathbf{c}_i.

    The formulation is differentiable but is not intended as a physically based density-rendering model; its purpose is to concentrate color evidence around an explicit surface and preserve a hard occupied/free-space boundary.

  3. Knowl 3 — Two-MLP scene representation with per-frame appearance correction

    model/method

    The learned scene representation consists of a shape MLP and a radiance MLP. For a 3D query point p∈R3\mathbf{p}\in\mathbb{R}^3, the shape MLP receives a sinusoidal positional encoding γ(p)\gamma(\mathbf{p}) and predicts both a truncated signed distance D(p)D(\mathbf{p}) and an intermediate feature vector. The radiance MLP receives that feature vector, the positional encoding γ(d)\gamma(\mathbf{d}) of a viewing direction d∈R3\mathbf{d}\in\mathbb{R}^3, and a learned latent code specific to the input frame; it outputs view-dependent RGB radiance. View-direction conditioning models effects such as specularities without deforming the geometry, while the per-frame latent code compensates for exposure and white-balance changes across the capture sequence.

  4. Knowl 4 — Joint camera pose and image-plane refinement

    model/method

    Camera poses are initialized with BundleFusion and represented by a rotation parameterized by Euler angles together with a 3D translation. The poses are refined jointly with the neural scene representation. To absorb image distortions and errors in the camera intrinsics, the method also learns a shared image-plane deformation field: a six-layer ReLU MLP maps an input pixel location to a 2D residual displacement. For every frame, the pixel is first shifted by this residual, then unprojected using the camera model and transformed into world space using the refined pose. The deformation field is shared across frames, whereas the pose parameters are frame-specific.

  5. Knowl 5 — Depth, free-space, and photometric optimization objective

    equation

    Let P\mathcal{P} denote all optimized parameters, including the MLP parameters, camera parameters, image-plane deformation parameters, and per-frame appearance codes. For BB training batches, with Pb\mathcal{P}_b the set of sampled pixels/rays in batch bb, the objective is

    L(P)=∑b=0B−1[λ1Lrgbb(P)+λ2Lfsb(P)+λ3Ltrb(P)].\mathcal{L}(\mathcal{P})=\sum_{b=0}^{B-1}\left[\lambda_1\mathcal{L}_{rgb}^{b}(\mathcal{P})+\lambda_2\mathcal{L}_{fs}^{b}(\mathcal{P})+\lambda_3\mathcal{L}_{tr}^{b}(\mathcal{P})\right].

    For an observed RGB value C^p∈R3\widehat{\mathbf{C}}_p\in\mathbb{R}^3 and rendered value Cp∈R3\mathbf{C}_p\in\mathbb{R}^3 at pixel pp, the photometric loss is

    Lrgbb=1∣Pb∣∑p∈Pb(Cp−C^p)2.\mathcal{L}_{rgb}^{b}=\frac{1}{|\mathcal{P}_b|}\sum_{p\in\mathcal{P}_b}\left(\mathbf{C}_p-\widehat{\mathbf{C}}_p\right)^2.

    For each ray, Spfs\mathcal{S}^{fs}_p is the set of samples between the camera and the truncation region of the observed surface, and Sptr\mathcal{S}^{tr}_p is the set of samples within that truncation region. If DsD_s is the predicted signed distance and D^s\widehat{D}_s is the signed distance implied by the depth measurement at sample ss, the free-space and truncated-region losses are

    Lfsb=1∣Pb∣∑p∈Pb1∣Spfs∣∑s∈Spfs(Ds−tr)2,\mathcal{L}_{fs}^{b}=\frac{1}{|\mathcal{P}_b|}\sum_{p\in\mathcal{P}_b}\frac{1}{|\mathcal{S}^{fs}_p|}\sum_{s\in\mathcal{S}^{fs}_p}(D_s-t_r)^2, Ltrb=1∣Pb∣∑p∈Pb1∣Sptr∣∑s∈Sptr(Ds−D^s)2.\mathcal{L}_{tr}^{b}=\frac{1}{|\mathcal{P}_b|}\sum_{p\in\mathcal{P}_b}\frac{1}{|\mathcal{S}^{tr}_p|}\sum_{s\in\mathcal{S}^{tr}_p}(D_s-\widehat{D}_s)^2.

    The experiments use tr=5 cmt_r=5\,\mathrm{cm}, scale the truncation interval to [−1,1][-1,1] with positive values in front of the surface and negative values behind it, and set (λ1,λ2,λ3)=(0.1,10,6×103)(\lambda_1,\lambda_2,\lambda_3)=(0.1,10,6\times10^3).

  6. Knowl 6 — Coarse-to-fine ray sampling and training procedure

    algorithm

    The reconstruction is optimized from aligned RGB-D frames, BundleFusion-initialized camera poses, and camera intrinsics. Each training iteration samples 1,024 random rays and uses a two-stage sampling procedure to concentrate computation near a predicted surface.

    Input: RGB-D frames, initial camera poses, camera intrinsics, truncation distance tr=5t_r=5 cm
    Initialize: shape MLP, radiance MLP, per-frame corrective codes, camera poses, and shared image-plane deformation MLP
    For 200,000 iterations:
        Randomly sample a batch of 1,024 pixels from the input frames
        For each sampled pixel:
            Apply the image-plane deformation residual
            Unproject the corrected pixel and transform its ray using the current camera pose
            Stratify-sample coarse points along a 4–8 m ray, with about one point every 1.5 cm
            Evaluate the shape MLP at all coarse points
            Search the predicted signed distances for the first zero crossing
            Sample 16 additional fine points around the zero crossing
            Evaluate both MLPs at the coarse and fine points
            Integrate all samples with the signed-distance weights
            Compute RGB, free-space, and truncated-region depth losses
        Sum the weighted losses over the batch
        Update all trainable parameters with ADAM at learning rate 5×10−45\times10^{-4}
    Return: optimized scene MLPs, camera parameters, deformation field, and corrective codes
    Extract: the zero level set of the optimized TSDF with Marching Cubes

    The coarse sampling density must be high enough to place samples inside the truncation region; otherwise the first surface zero crossing can be missed. The extracted meshes use a spatial resolution of 1 cm1\,\mathrm{cm}.

  7. Knowl 7 — Synthetic reconstruction benchmark

    data/table

    The method was evaluated against depth-fusion, learned occupancy, SIREN, and density-based NeRF baselines on 10 synthetic scenes. Each scene has known ground-truth geometry and camera trajectory; the trajectory was used only to render and evaluate the data, not to reconstruct it. Photorealistic RGB images were rendered with Blender, and depth images were corrupted with artifacts modeled after consumer depth sensors. Chamfer ℓ1\ell_1 distance (CC-ℓ1\ell_1) is lower when predicted and ground-truth point clouds are closer; IoU, normal consistency (NC), and F-score are higher when reconstruction quality is better. Point clouds were sampled at 1 point per cm2\mathrm{cm}^2, F-score used a 5 cm threshold, and IoU was computed after voxelizing the mesh.

    Method CC-ℓ1\ell_1 ↓\downarrow IoU ↑\uparrow NC ↑\uparrow F-score ↑\uparrow
    BundleFusion 0.062 0.594 0.892 0.805
    RoutedFusion 0.057 0.615 0.864 0.838
    COLMAP + Poisson 0.057 0.619 0.901 0.839
    Conv. Occ. Nets 0.077 0.461 0.849 0.643
    SIREN 0.060 0.603 0.893 0.816
    NeRF + Depth 0.065 0.550 0.768 0.782
    Ours (w/o pose) 0.049 0.655 0.908 0.868
    Ours 0.044 0.747 0.918 0.924

    The full method obtains the best value for every metric. In particular, it outperforms the density-based NeRF with an additional depth loss, supporting the use of an explicit signed-distance surface representation. Joint camera refinement improves the method from 0.0490.049 to 0.0440.044 in Chamfer ℓ1\ell_1, from 0.6550.655 to 0.7470.747 in IoU, from 0.9080.908 to 0.9180.918 in NC, and from 0.8680.868 to 0.9240.924 in F-score.

  8. Knowl 8 — Color supervision recovers geometry missing from depth

    empirical result

    Depth measurements from the real ScanNet RGB-D sequences are noisy and frequently omit thin structures such as chair legs. The method's photometric loss supplies geometric evidence in these regions: in a synthetic test with missing depth for table legs and a meshed basket, the full model reconstructs the missing geometry from RGB observations, whereas removing the photometric term reduces completeness. In the full model, regions supported only by color have an average accuracy error of 11 mm11\,\mathrm{mm}, compared with 8 mm8\,\mathrm{mm} for regions that also have depth measurements.

    On ScanNet room-scale scenes captured with a StructureIO camera, the model without pose refinement produces smoother geometry than density-based NeRF but retains camera-misalignment artifacts. The complete model, which jointly refines poses and the image-plane deformation field, substantially reduces those artifacts and produces cleanly aligned reconstructions in the evaluated scenes.

  9. Knowl 9 — Camera refinement and image-plane deformation improve accuracy

    empirical result

    On all synthetic scenes, the estimated camera poses were compared with ground truth using mean translational error in meters and rotational error in degrees. The proposed joint refinement improves the BundleFusion initialization and also outperforms COLMAP.

    Method Pos. error (meters) ↓\downarrow Rot. error (degrees) ↓\downarrow
    BundleFusion 0.033 0.571
    COLMAP 0.038 0.692
    Ours 0.021 0.144

    The image-plane deformation field was separately tested on a synthetic scene with an incorrect focal length of 570 instead of the ground-truth value 554.26. Without the deformation field, the reconstruction scores were CC-ℓ1=0.061\ell_1=0.061, IoU =0.266=0.266, NC =0.886=0.886, and F-score =0.406=0.406. With the deformation field, the scores improved to CC-ℓ1=0.031\ell_1=0.031, IoU =0.609=0.609, NC =0.911=0.911, and F-score =0.904=0.904. Thus, the learned pixel-space correction compensates for intrinsic-parameter error and improves geometric alignment.

  10. Knowl 10 — Computational and representation limitations

    limitation

    The method is an offline, scene-specific optimization: approximately 9 hours on an NVIDIA RTX 3090 are required for 2×1052\times10^5 iterations. Its global MLP stores the entire scene in one network, which can omit high-frequency local detail in very large scenes. The authors identify locally conditioned MLPs or voxel-grid representations as possible ways to improve scalability and convergence. The method is designed for opaque surfaces and does not address transparent or otherwise non-opaque geometry.

Coverage note — Detailed baseline implementation descriptions, per-scene visual comparisons, and related-work discussion were omitted because they do not add independent contributed knowledge beyond the reported method, measurements, and ablations.

References

  1. 1.Jonathan T. Barron and Jitendra Malik. Intrinsic scene properties from a single rgb-d image. CVPR, 2013. 5
  2. 2.Michael Bloesch, Jan Czarnowski, Ronald Clark, Stefan Leutenegger, and Andrew J. Davison. Codeslam — learning a compact, optimisable representation for dense visual slam. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), June 2018. 2
  3. 3.Jeannette Bohg, Javier Romero, Alexander Herzog, and Stefan Schaal. Robot arm pose estimation through pixel-wise part classification. ICRA, 2014. 5
  4. 4.Aljaz Božičić, Pablo Palafox, Justus Thies, Angela Dai, and Matthias Nießner. Transformerfusion: Monocular rgb scene reconstruction using transformers. Proc. Neural Information Processing Systems (NeurIPS), 2021. 2
  5. 5.Eric Chan, Marco Monteiro, Petr Kellnhofer, Jiajun Wu, and Gordon Wetzstein. pi-gan: Periodic implicit generative adversarial networks for 3d-aware image synthesis. In arXiv, 2020. 3
  6. 6.Angel Chang, Angela Dai, Thomas Funkhouser, Maciej Halber, Matthias Niessner, Manolis Savva, Shuran Song, Andy Zeng, and Yinda Zhang. Matterport3d: Learning from rgb-d data in indoor environments. International Conference on 3D Vision (3DV), 2017. 5
  7. 7.Julian Chibane, Thiemo Alldieck, and Gerard Pons-Moll. Implicit functions in feature space for 3d shape reconstruction and completion. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, jun 2020. 2, 8
  8. 8.Blender Online Community. Blender - a 3D modelling and rendering package. Blender Foundation, Stichting Blender Foundation, Amsterdam, 2018. 5
  9. 9.Brian Curless and Marc Levoy. A volumetric method for building complex models from range images. In Proceedings of the 23rd Annual Conference on Computer Graphics and Interactive Techniques, SIGGRAPH ’96, page 303–312, New York, NY, USA, 1996. Association for Computing Machinery. 1, 2, 5
  10. 10.J Czarnowski, T Laidlow, R Clark, and AJ Davison. Deepfactors: Real-time probabilistic dense monocular slam. IEEE Robotics and Automation Letters, 5:721–728, 2020. 2
  11. 11.Manuel Dahnert, Ji Hou, , Matthias Nießner, and Angela Dai. Panoptic 3D scene reconstruction from a single RGB image. 2021. 2
  12. 12.Angela Dai, Angel X. Chang, Manolis Savva, Maciej Halber, Thomas Funkhouser, and Matthias Nießner. Scannet: Richly-annotated 3d reconstructions of indoor scenes. In Proc. Computer Vision and Pattern Recognition (CVPR), IEEE, 2017. 2, 5, 7
  13. 13.Angela Dai, Christian Diller, and Matthias Nießner. Sg-nn: Sparse generative neural networks for self-supervised scene completion of rgb-d scans. In Proc. Computer Vision and Pattern Recognition (CVPR), IEEE, 2020. 2
  14. 14.Angela Dai, Matthias Nießner, Michael Zollhofer, Shahram Izadi, and Christian Theobalt. Bundlefusion: Real-time globally consistent 3d reconstruction using on-the-fly surface reintegration. ACM Transactions on Graphics (TOG), 36(4):76a, 2017. 2, 3, 4, 5, 7, 8
  15. 15.Angela Dai, Yawar Siddiqui, Justus Thies, Julien Valentin, and Matthias Nießner. Spsg: Self-supervised photometric scene generation from rgb-d scans. In Proc. Computer Vision and Pattern Recognition (CVPR), IEEE, 2021. 2
  16. 16.Kangle Deng, Andrew Liu, Jun-Yan Zhu, and Deva Ramanan. Depth-supervised nerf: Fewer views and faster training for free. arXiv preprint arXiv:2107.02791, 2021. 2, 3
  17. 17.Maximilian Denninger, Martin Sundermeyer, Dominik Winkelbauer, Youssef Zidan, Dmitry Olefir, Mohamad Elbadrawy, Ahsan Lodhi, and Harinandan Katam. Blenderproc. arXiv preprint arXiv:1911.01911, 2019. 5
  18. 18.Maximilian Denninger and Rudolph Triebel. 3d scene reconstruction from a single viewport. In Proceedings of the European Conference on Computer Vision (ECCV), 2020. 2
  19. 19.Wei Dong, Qiuyuan Wang, Xin Wang, and Hongbin Zha. PSDF fusion: Probabilistic signed distance function for on-the-fly 3d data fusion and scene reconstruction. CoRR, abs/1807.11034, 2018. 2
  20. 20.J. Engel, T. Schops, and D. Cremers. LSD-SLAM: Large-scale direct monocular SLAM. In European Conference on Computer Vision (ECCV), September 2014. 2
  21. 21.J. Engel, J. Sturm, and D. Cremers. Semi-dense visual odometry for a monocular camera. In 2013 IEEE International Conference on Computer Vision, pages 1449–1456, 2013. 2
  22. 22.John Flynn, Ivan Neulander, James Philbin, and Noah Snavely. Deepstereo: Learning to predict new views from the world’s imagery. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 5515–5524, 2016. 2
  23. 23.C. Forster, M. Pizzoli, and D. Scaramuzza. Svo: Fast semi-direct monocular visual odometry. In 2014 IEEE International Conference on Robotics and Automation (ICRA), pages 15–22, 2014. 2
  24. 24.Huan Fu, Mingming Gong, Chaohui Wang, Kayhan Batmanghelich, and Dacheng Tao. Deep ordinal regression network for monocular depth estimation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 2002–2011, 2018. 2
  25. 25.Y. Furukawa and J. Ponce. Accurate, dense, and robust multi-view stereopsis. IEEE Transactions on Pattern Analysis and Machine Intelligence, 32(8):1362–1376, 2010. 2
  26. 26.Guy Gafni, Justus Thies, Michael Zollhofer, and Matthias Nießner. Dynamic neural radiance fields for monocular 4d facial avatar reconstruction. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2021. 3
  27. 27.Georgia Gkioxari, Jitendra Malik, and Justin Johnson. Mesh R-CNN. CoRR, abs/1906.02739, 2019. 2
  28. 28.Clement Godard, Oisin Mac Aodha, and Gabriel J Brostow. Unsupervised monocular depth estimation with left-right consistency. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 270–279, 2017. 2
  29. 29.M. Goesele, N. Snavely, B. Curless, H. Hoppe, and S. M. Seitz. Multi-view stereo for community photo collections. In 2007 IEEE 11th International Conference on Computer Vision, pages 1–8, 2007. 2
  30. 30.Ankur Handa. Simulating kinect noise for the icl-nuim dataset. https://github.com/ankurhanda/simkinect. Accessed: 2021-11-15. 5
  31. 31.Ankur Handa, Thomas Whelan, John McDonald, and Andrew J Davison. A benchmark for rgb-d visual odometry, 3d reconstruction and slam. ICRA, 2014. 5, 7
  32. 32.Vu Hoang Hiep, Renaud Keriven, Patrick Labatut, and Jean-Philippe Pons. Towards high-resolution large-scale multi-view stereo. In 2009 IEEE Conference on Computer Vision and Pattern Recognition, pages 1430–1437. IEEE, 2009. 2
  33. 33.Shahram Izadi, David Kim, Otmar Hilliges, David Molyneaux, Richard Newcombe, Pushmeet Kohli, Jamie Shotton, Steve Hodges, Dustin Freeman, Andrew Davison, et al. Kinectfusion: real-time 3d reconstruction and interaction using a moving depth camera. In Proceedings of the 24th annual ACM symposium on User interface software and technology, pages 559–568. ACM, 2011. 1
  34. 34.Abhishek Kar, Christian Hane, and Jitendra Malik. Learning a multi-view stereo machine. arXiv preprint arXiv:1708.05375, 2017. 2
  35. 35.Michael Kazhdan and Hugues Hoppe. Screened poisson surface reconstruction. ACM Trans. Graph., 32(3), July 2013. 5
  36. 36.Diederik P. Kingma and Jimmy Ba. Adam: A method for stochastic optimization. CoRR, abs/1412.6980, 2014. 5
  37. 37.Arno Knapitsch, Jaesik Park, Qian-Yi Zhou, and Vladlen Koltun. Tanks and temples: Benchmarking large-scale scene reconstruction. ACM Trans. Graph., 36(4), July 2017. 5, 7
  38. 38.Kiriakos N Kutulakos and Steven M Seitz. A theory of shape by space carving. International journal of computer vision, 38(3):199–218, 2000. 2
  39. 39.Iro Laina, Christian Rupprecht, Vasileios Belagiannis, Federico Tombari, and Nassir Navab. Deeper depth prediction with fully convolutional residual networks. In 2016 Fourth international conference on 3D vision (3DV), pages 239–248. IEEE, 2016. 2
  40. 40.Zhengqi Li, Simon Niklaus, Noah Snavely, and Oliver Wang. Neural scene flow fields for space-time view synthesis of dynamic scenes. arXiv preprint arXiv:2011.13084, 2020. 3
  41. 41.Chen-Hsuan Lin, Wei-Chiu Ma, Antonio Torralba, and Simon Lucey. Barf: Bundle-adjusting neural radiance fields. In IEEE International Conference on Computer Vision (ICCV), 2021. 3
  42. 42.Stephen Lombardi, Tomas Simon, Jason Saragih, Gabriel Schwartz, Andreas Lehrmann, and Yaser Sheikh. Neural volumes. ACM Transactions on Graphics, 38(4):1–14, Jul 2019. 1
  43. 43.William E. Lorensen and Harvey E. Cline. Marching cubes: A high resolution 3d surface construction algorithm. In Proceedings of the 14th Annual Conference on Computer Graphics and Interactive Techniques, SIGGRAPH ’87, page 163–169, New York, NY, USA, 1987. Association for Computing Machinery. 3, 5
  44. 44.D. G. Lowe. Object recognition from local scale-invariant features. In Proceedings of the Seventh IEEE International Conference on Computer Vision, volume 2, pages 1150–1157 vol.2, 1999. 2
  45. 45.Ricardo Martin-Brualla, Noha Radwan, Mehdi S. M. Sajjadi, Jonathan T. Barron, Alexey Dosovitskiy, and Daniel Duckworth. NeRF in the Wild: Neural Radiance Fields for Unconstrained Photo Collections. In CVPR, 2021. 3, 4, 5
  46. 46.N. Max. Optical models for direct volume rendering. IEEE Transactions on Visualization and Computer Graphics, 1(2):99–108, 1995. 3
  47. 47.Lars Mescheder, Michael Oechsle, Michael Niemeyer, Sebastian Nowozin, and Andreas Geiger. Occupancy networks: Learning 3d reconstruction in function space. In Proceedings IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), 2019. 2
  48. 48.Ben Mildenhall, Pratul P. Srinivasan, Matthew Tancik, Jonathan T. Barron, Ravi Ramamoorthi, and Ren Ng. Nerf: Representing scenes as neural radiance fields for view synthesis. In ECCV, 2020. 1, 2, 4, 7
  49. 49.T. Neff, P. Stadlbauer, M. Parger, A. Kurz, J. H. Mueller, C. R.A. Chaitanya, A. Kaplanyan, and M. Steinberger. DONeRF: Towards Real-Time Rendering of Compact Neural Radiance Fields using Depth Oracle Networks. Computer Graphics Forum, 40(4):45–59, 2021. 2, 3
  50. 50.Richard A Newcombe, Shahram Izadi, Otmar Hilliges, David Molyneaux, David Kim, Andrew J Davison, Pushmeet Kohi, Jamie Shotton, Steve Hodges, and Andrew Fitzgibbon. Kinectfusion: Real-time dense surface mapping and tracking. In Mixed and augmented reality (ISMAR), 2011 10th IEEE international symposium on, pages 127–136. IEEE, 2011. 2
  51. 51.Yinyu Nie, Xiaoguang Han, Shihui Guo, Yujian Zheng, Jian Chang, and Jian Jun Zhang. Total3DUnderstanding: Joint layout, object pose and mesh reconstruction for indoor scenes from a single image. pages 55–64, 2020. 2
  52. 52.Michael Niemeyer, Lars Mescheder, Michael Oechsle, and Andreas Geiger. Differentiable volumetric rendering: Learning implicit 3d representations without 3d supervision. In Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), 2020. 2
  53. 53.M. Nießner, M. Zollhofer, S. Izadi, and M. Stamminger. Real-time 3d reconstruction at scale using voxel hashing. ACM Transactions on Graphics (TOG), 2013. 1, 2
  54. 54.Michael Oechsle, Lars Mescheder, Michael Niemeyer, Thilo Strauss, and Andreas Geiger. Texture fields: Learning texture representations in function space. In International Conference on Computer Vision, Oct. 2019. 2
  55. 55.Michael Oechsle, Songyou Peng, and Andreas Geiger. Unisurf: Unifying neural implicit surfaces and radiance fields for multi-view reconstruction. In International Conference on Computer Vision (ICCV), 2021. 2
  56. 56.Jeong Joon Park, Peter Florence, Julian Straub, Richard Newcombe, and Steven Lovegrove. Deepsdf: Learning continuous signed distance functions for shape representation. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), June 2019. 2
  57. 57.Keunhong Park, Utkarsh Sinha, Jonathan T Barron, Sofien Bouaziz, Dan B Goldman, Steven M Seitz, and Ricardo Martin Brualla. Deformable neural radiance fields. arXiv preprint arXiv:2011.12948, 2020. 3
  58. 58.Songyou Peng, Michael Niemeyer, Lars Mescheder, Marc Pollefeys, and Andreas Geiger. Convolutional occupancy networks. In European Conference on Computer Vision (ECCV), Cham, Aug. 2020. Springer International Publishing. 2, 5, 8
  59. 59.Shunsuke Saito, , Zeng Huang, Ryota Natsume, Shigeo Morishima, Angjoo Kanazawa, and Hao Li. Pifu: Pixel-aligned implicit function for high-resolution clothed human digitization. arXiv preprint arXiv:1905.05172, 2019. 2
  60. 60.Shunsuke Saito, Tomas Simon, Jason Saragih, and Hanbyul Joo. Pifuhd: Multi-level pixel-aligned implicit function for high-resolution 3d human digitization. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, June 2020. 2
  61. 61.Nikolay Savinov, Christian Hane, L’ubor Ladický, and Marc Pollefeys. Semantic 3d reconstruction with continuous regularization and ray potentials using a visibility consistency constraint. In 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 5460–5469, 2016. 2
  62. 62.D. Scharstein, R. Szeliski, and R. Zabih. A taxonomy and evaluation of dense two-frame stereo correspondence algorithms. In Proceedings IEEE Workshop on Stereo and Multi-Baseline Vision (SMBV 2001), pages 131–140, 2001. 2
  63. 63.Johannes L. Schonberger and Jan-Michael Frahm. Structure-from-motion revisited. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), June 2016. 2, 5
  64. 64.Johannes Lutz Schönberger, True Price, Torsten Sattler, Jan-Michael Frahm, and Marc Pollefeys. A vote-and-verify strategy for fast spatial verification in image retrieval. In Asian Conference on Computer Vision (ACCV), 2016. 5
  65. 65.Johannes Lutz Schönberger, Enliang Zheng, Marc Pollefeys, and Jan-Michael Frahm. Pixelwise view selection for unstructured multi-view stereo. In European Conference on Computer Vision (ECCV), 2016. 5
  66. 66.Katja Schwarz, Yiyi Liao, Michael Niemeyer, and Andreas Geiger. Graf: Generative radiance fields for 3d-aware image synthesis. In Advances in Neural Information Processing Systems (NeurIPS), 2020. 3
  67. 67.Vincent Sitzmann, Julien N.P. Martel, Alexander W. Bergman, David B. Lindell, and Gordon Wetzstein. Implicit neural representations with periodic activation functions. In arXiv, 2020. 2, 3, 7
  68. 68.Vincent Sitzmann, Justus Thies, Felix Heide, Matthias Nießner, Gordon Wetzstein, and Michael Zollhofer. Deepvoxels: Learning persistent 3d feature embeddings. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2437–2446, 2019. 2
  69. 69.Vincent Sitzmann, Michael Zollhofer, and Gordon Wetzstein. Scene representation networks: Continuous 3d-structure-aware neural scene representations. In Advances in Neural Information Processing Systems, 2019. 2
  70. 70.Cheng Sun, Min Sun, and Hwann-Tzong Chen. Direct voxel grid optimization: Super-fast convergence for radiance fields reconstruction. arXiv preprint arXiv:2111.11215, 2021. 8
  71. 71.Jiaming Sun, Yiming Xie, Linghao Chen, Xiaowei Zhou, and Hujun Bao. NeuralRecon: Real-time coherent 3D reconstruction from monocular video. CVPR, 2021. 2
  72. 72.Matthew Tancik, Pratul P. Srinivasan, Ben Mildenhall, Sara Fridovich-Keil, Nithin Raghavan, Utkarsh Singhal, Ravi Ramamoorthi, Jonathan T. Barron, and Ren Ng. Fourier features let networks learn high frequency functions in low dimensional domains. NeurIPS, 2020. 3
  73. 73.Ayush Tewari, Ohad Fried, Justus Thies, Vincent Sitzmann, Stephen Lombardi, Kalyan Sunkavalli, Ricardo Martin-Brualla, Tomas Simon, Jason Saragih, Matthias Nießner, et al. State of the art on neural rendering. In Computer Graphics Forum, volume 39, pages 701–727. Wiley Online Library, 2020. 1, 2
  74. 74.Ayush Tewari, Justus Thies, Ben Mildenhall, Pratul Srinivasan, Edgar Tretschk, Yifan Wang, Christoph Lassner, Vincent Sitzmann, Ricardo Martin-Brualla, Stephen Lombardi, Tomas Simon, Christian Theobalt, Matthias Niessner, Jonathan T. Barron, Gordon Wetzstein, Michael Zollhoefer, and Vladislav Golyanik. Advances in neural rendering, 2021. 2
  75. 75.Nanyang Wang, Yinda Zhang, Zhuwen Li, Yanwei Fu, Wei Liu, and Yu-Gang Jiang. Pixel2mesh: Generating 3d mesh models from single rgb images. In ECCV, 2018. 2
  76. 76.Peng Wang, Lingjie Liu, Yuan Liu, Christian Theobalt, Taku Komura, and Wenping Wang. Neus: Learning neural implicit surfaces by volume rendering for multi-view reconstruction. arXiv preprint arXiv:2106.10689, 2021. 2, 3, 4
  77. 77.Qianqian Wang, Zhicheng Wang, Kyle Genova, Pratul Srinivasan, Howard Zhou, Jonathan T. Barron, Ricardo Martin-Brualla, Noah Snavely, and Thomas Funkhouser. Ibrnet: Learning multi-view image-based rendering. CVPR, 2021. 3
  78. 78.Zirui Wang, Shangzhe Wu, Weidi Xie, Min Chen, and Victor Adrian Prisacariu. Nerf–: Neural radiance fields without known camera parameters, 2021. 3
  79. 79.Silvan Weder, Johannes L. Schönberger, Marc Pollefeys, and Martin R. Oswald. Routedfusion: Learning real-time depth map fusion. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2020. 2, 5
  80. 80.Silvan Weder, Johannes L. Schönberger, Marc Pollefeys, and Martin R. Oswald. Neuralfusion: Online depth fusion in latent space. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 3162–3172, June 2021. 2
  81. 81.Yi Wei, Shaohui Liu, Yongming Rao, Wang Zhao, Jiwen Lu, and Jie Zhou. Nerfingmvs: Guided optimization of neural radiance fields for indoor multi-view stereo. In ICCV, 2021. 2, 3
  82. 82.Yao Yao, Zixin Luo, Shiwei Li, Tian Fang, and Long Quan. Mvsnet: Depth inference for unstructured multi-view stereo. In Proceedings of the European Conference on Computer Vision (ECCV), pages 767–783, 2018. 2
  83. 83.Lior Yariv, Jiatao Gu, Yoni Kasten, and Yaron Lipman. Volume rendering of neural implicit surfaces. Advances in Neural Information Processing Systems, 34, 2021. 2
  84. 84.Lior Yariv, Yoni Kasten, Dror Moran, Meirav Galun, Matan Atzmon, Basri Ronen, and Yaron Lipman. Multiview neural surface reconstruction by disentangling geometry and appearance. Advances in Neural Information Processing Systems, 33, 2020. 2
  85. 85.Lin Yen-Chen, Pete Florence, Jonathan T. Barron, Alberto Rodriguez, Phillip Isola, and Tsung-Yi Lin. inerf: Inverting neural radiance fields for pose estimation. 2020. 3
  86. 86.Alex Yu, Sara Fridovich-Keil, Matthew Tancik, Qinhong Chen, Benjamin Recht, and Angjoo Kanazawa. Plenoxels: Radiance fields without neural networks. arXiv preprint arXiv:2112.05131, 2021. 8
  87. 87.Alex Yu, Vickie Ye, Matthew Tancik, and Angjoo Kanazawa. pixelnerf: Neural radiance fields from one or few images. In CVPR, 2021. 3
  88. 88.Christopher Zach, Thomas Pock, and Horst Bischof. A globally optimal algorithm for robust tv-l1 range image integration. In 2007 IEEE 11th International Conference on Computer Vision, pages 1–8, 2007. 2
  89. 89.Yinda Zhang and Thomas Funkhouser. Deep depth completion of a single rgb-d image. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 175–185, 2018. 2
  90. 90.Shuaifeng Zhi, Michael Bloesch, Stefan Leutenegger, and Andrew J. Davison. Scenecode: Monocular dense semantic reconstruction using learned encoded scene representations. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2019. 2
  91. 91.Qian-Yi Zhou and Vladlen Koltun. Color map optimization for 3d reconstruction with consumer depth cameras. 33(4), July 2014. 4
  92. 92.M. Zollhöfer, P. Stotko, A. Görlitz, C. Theobalt, M. Nießner, R. Klein, and A. Kolb. State of the Art on 3D Reconstruction with RGB-D Cameras. Computer Graphics Forum (Eurographics State of the Art Reports 2018), 37(2), 2018. 2

Citation

MLA
Azinović, D., et al. “Neural RGB-D Surface Reconstruction”. arXiv, 2021, http://arxiv.org/abs/2104.04532v3.
APA
Azinović, D., Martin-Brualla, R., Goldman, D. B., Nießner, M., & Thies, J. (2021). Neural RGB-D Surface Reconstruction. arXiv. http://arxiv.org/abs/2104.04532v3
Chicago
Azinović, D., R. Martin-Brualla, D. B. Goldman, M. Nießner, and J. Thies. 2021. “Neural RGB-D Surface Reconstruction”. arXiv. http://arxiv.org/abs/2104.04532v3.
Harvard
Azinović, D. et al. (2021) “Neural RGB-D Surface Reconstruction”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2104.04532v3.
Vancouver
1. Azinović D, Martin-Brualla R, Goldman DB, Nießner M, Thies J (2021) Neural RGB-D Surface Reconstruction. arXiv

BibTeX

@article{azinovic2021neural,
  title = {Neural RGB-D Surface Reconstruction},
  author = {Azinović, Dejan and Martin-Brualla, Ricardo and Goldman, Dan B and Nießner, Matthias and Thies, Justus},
  year = {2021},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2104.04532v3},
  eprint = {2104.04532}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE