Volume Rendering of Neural Implicit Surfaces

Lior YarivJiatao GuYoni KastenYaron Lipman

article2021NeurIPS1,412 citations

Proposes a neural volume rendering approach that derives volume density directly from a signed distance function, enabling high-fidelity 3D surface reconstruction and unsupervised disentanglement of shape and appearance from multiview images.

Listen

Synthesizing 3D scenes and novel views from sparse sets of 2D images is a central challenge in computer vision and graphics. While recent neural volume rendering techniques excel at novel view synthesis, they represent scene density using generic density fields that typically result in noisy, low-fidelity 3D geometry. Concurrently, methods that directly optimize neural implicit surfaces require accurate object masks to separate foreground from background and often generate erroneous surface artifacts due to optimization traps. The article introduces and evaluates VolSDF, a volume rendering framework designed to recover high-quality 3D surface geometry from multi-view images without requiring foreground masks, while retaining high visual rendering quality.

The authors develop an approach that models volumetric density directly as a mathematical transformation of a learnable signed distance function, which defines the distance to the scene's boundary surface. This formulation provides a strong geometric inductive bias and enables a theoretical bound on the error of numerical opacity approximations along camera viewing rays. Leveraging this bound, the method employs an adaptive sampling algorithm that tightly couples geometry and radiance during numerical integration. The overall pipeline uses two neural networks trained end-to-end on image color and geometric regularization losses, and is tested across benchmark multi-view datasets including DTU and BlendedMVS.

The evaluation reveals three key findings. First, VolSDF reconstructs significantly more accurate 3D geometry than standard neural volume rendering baselines, lowering the reconstruction error metric by more than 50% relative to standard baselines on DTU and achieving an average 51.8% geometric improvement on BlendedMVS scenes. Second, the method achieves geometric accuracy comparable to state-of-the-art implicit surface methods while operating fully unsupervised without object segmentation masks. Third, the framework successfully disentangles underlying 3D geometry from appearance, enabling the direct transfer of surface materials between distinct scenes without the visual failures seen in existing approaches.

These findings demonstrate that embedding a principled geometric surface representation into volumetric rendering overcomes the trade-off between visual rendering fidelity and geometric reconstruction accuracy. In practical terms, this lowers data capture and annotation costs by removing the need for manual foreground masking, while simultaneously reducing the risk of reconstruction artifacts in automated 3D modeling workflows. The ability to separate shape and radiance also opens viable avenues for modular asset editing and digital content creation across simulation, gaming, and visualization pipelines.

Organizations seeking automated, mask-free 3D reconstruction from multi-view imagery should consider adopting signed-distance-based neural volume rendering frameworks. Future engineering and research efforts should focus on formalizing theoretical convergence guarantees for the sampling algorithm, expanding the formulation to handle non-watertight surfaces and open boundaries, and extending the model to support dynamic, moving scenes and broader shape spaces.

Confidence in these results is high across closed, opaque foreground objects under standard multi-view capture conditions. However, performance remains constrained in unobserved or weakly observed regions, where the model may complete missing geometry arbitrarily, as well as across large, textureless surfaces. Stakeholders should account for these limitations when applying the technique to scenes with significant occlusions or lack of visual texture.

Cover for Volume Rendering of Neural Implicit Surfaces

Abstract

Neural volume rendering became increasingly popular recently due to its success in synthesizing novel views of a scene from a sparse set of input images. So far, the geometry learned by neural volume rendering techniques was modeled using a generic density function. Furthermore, the geometry itself was extracted using an arbitrary level set of the density function leading to a noisy, often low fidelity reconstruction. The goal of this paper is to improve geometry representation and reconstruction in neural volume rendering. We achieve that by modeling the volume density as a function of the geometry. This is in contrast to previous work modeling the geometry as a function of the volume density. In more detail, we define the volume density function as Laplace's cumulative distribution function (CDF) applied to a signed distance function (SDF) representation. This simple density representation has three benefits: (i) it provides a useful inductive bias to the geometry learned in the neural volume rendering process; (ii) it facilitates a bound on the opacity approximation error, leading to an accurate sampling of the viewing ray. Accurate sampling is important to provide a precise coupling of geometry and radiance; and (iii) it allows efficient unsupervised disentanglement of shape and appearance in volume rendering. Applying this new density representation to challenging scene multiview datasets produced high quality geometry reconstructions, outperforming relevant baselines. Furthermore, switching shape and appearance between scenes is possible due to the disentanglement of the two.

Citation

MLA
Yariv, L., et al. “Volume Rendering of Neural Implicit Surfaces”. arXiv, 2021, http://arxiv.org/abs/2106.12052v2.
APA
Yariv, L., Gu, J., Kasten, Y., & Lipman, Y. (2021). Volume Rendering of Neural Implicit Surfaces. arXiv. http://arxiv.org/abs/2106.12052v2
Chicago
Yariv, L., J. Gu, Y. Kasten, and Y. Lipman. 2021. “Volume Rendering of Neural Implicit Surfaces”. arXiv. http://arxiv.org/abs/2106.12052v2.
Harvard
Yariv, L. et al. (2021) “Volume Rendering of Neural Implicit Surfaces”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2106.12052v2.
Vancouver
1. Yariv L, Gu J, Kasten Y, Lipman Y (2021) Volume Rendering of Neural Implicit Surfaces. arXiv

BibTeX

@article{yariv2021volume,
  title = {Volume Rendering of Neural Implicit Surfaces},
  author = {Yariv, Lior and Gu, Jiatao and Kasten, Yoni and Lipman, Yaron},
  year = {2021},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2106.12052v2},
  eprint = {2106.12052}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors