VolRecon: Volume Rendering of Signed Ray Distance Functions for Generalizable Multi-View Reconstruction
Yufan RenFangjinhua WangTong ZhangMarc PollefeysSabine Süsstrunk
Proposes a generalizable neural implicit reconstruction framework that combines local multi-view projection features with coarse global volume features using signed ray distance functions to recover high-detail surfaces across unseen scenes without per-scene optimization.
Three-dimensional surface reconstruction from two-dimensional images is a foundational capability for robotics, autonomous navigation, and augmented reality. Recent advances in neural implicit representations have enabled photorealistic image rendering, but existing methods typically require lengthy, per-scene optimization that cannot generalize to unseen environments. Early generalizable implicit reconstruction frameworks often rely on coarse global feature representations that constrain geometric resolution, leading to over-smoothed surfaces and loss of critical spatial detail, especially when only a few input views are available.
The article introduces VolRecon, a neural implicit reconstruction framework designed to reconstruct detailed 3D surfaces across new, unseen scenes without per-scene retraining. The primary objective is to demonstrate that combining directional ray distance formulations with both local and global image features achieves high-fidelity geometry and robust generalization across different scene scales.
The evaluated approach uses a Signed Ray Distance Function (SRDF)—which calculates the shortest distance to a surface along a specific viewing ray—integrated into a volume rendering pipeline. The architecture uses a view transformer to aggregate local multi-view projection features and a ray transformer to process points along a ray alongside global shape priors from a coarse 3D feature volume. The system was trained on the standard indoor DTU dataset using color and depth supervision, and subsequently evaluated on 15 test DTU scenes as well as the larger-scale ETH3D benchmark to measure zero-shot generalization across varying camera viewpoints.
The experimental findings demonstrate significant performance gains over existing approaches. In sparse three-view reconstructions on DTU, VolRecon improved reconstruction accuracy by approximately 30% over the leading generalizable implicit baseline, SparseNeuS, and outperformed the traditional multi-view stereo pipeline COLMAP by about 10%. In full-view reconstructions, VolRecon reduced Chamfer distance errors by roughly 22% compared to SparseNeuS while achieving geometric accuracy comparable to dedicated deep multi-view stereo networks like MVSNet. Depth estimation evaluations showed superior absolute and relative depth accuracy across all benchmark thresholds, while tests on the ETH3D dataset confirmed that the model generalizes directly to large-scale scenes without fine-tuning.
These results indicate that generalizable neural reconstruction can deliver sharp boundaries and fine geometric features without requiring separate optimization cycles for every new target environment. For operational deployments, this capability substantially reduces the computational overhead and latency associated with scene-specific model training, offering a more scalable path for automated mapping and perception pipelines.
For future development, the authors recommend adopting progressive local bounding volume strategies to overcome memory constraints when scaling to very large scenes. Operational users should note that the current rendering speed of approximately 30 seconds per standard-resolution view remains a bottleneck for real-time applications, meaning near-term deployment is best suited for offline or batch reconstruction workflows until rendering efficiency is improved.
- Paper: NeuS: Learning Neural Implicit Surfaces by Volume Rendering for Multi-view Reconstruction, Peng Wang et al. (2021). NeuS establishes the foundational framework for learning neural implicit surfaces via volume rendering of signed distance functions, which VolRecon directly adapts and generalises.
- Paper: Volume Rendering of Neural Implicit Surfaces, Lior Yariv et al. (2021). VolSDF provides key theoretical and practical principles for transforming signed distance functions into volumetric densities for neural rendering, underpinning VolRecon's formulation.
- Paper: MVSNet: Depth Inference for Unstructured Multi-view Stereo, Yao Yao et al. (2018). MVSNet introduces the global cost volume and multi-view feature aggregation techniques that VolRecon relies upon to construct its coarse global feature volume.
- Paper: NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis, Ben Mildenhall et al. (2020). NeRF introduces the fundamental coordinate-based neural radiance field and differentiable volume rendering formulation that neural implicit reconstruction methods build upon.
- Paper: pixelNeRF: Neural Radiance Fields from One or Few Images, Alex Yu et al. (2021). pixelNeRF pioneered generalisable, feed-forward neural radiance fields by conditioning continuous representations on extracted multi-view 2D image features.
- Paper: DeepSDF: Learning Continuous Signed Distance Functions for Shape Representation, Jeong Joon Park et al. (2019). DeepSDF introduces continuous neural implicit signed distance functions as a powerful representation for high-fidelity 3D shape reconstruction.
- Paper: PIFu: Pixel-Aligned Implicit Function for High-Resolution Clothed Human Digitization, Shunsuke Saito et al. (2019). PIFu demonstrates the mechanism of pixel-aligned feature projection to predict continuous implicit 3D geometry from 2D images.
- Paper: 2D Gaussian Splatting for Geometrically Accurate Radiance Fields, Binbin Huang et al. (2024). 2D Gaussian Splatting advances geometrically accurate surface reconstruction and fast rendering, offering an alternative explicit primitive formulation compared to generalizable implicit ray marching.
- Paper: DUSt3R: Geometric 3D Vision Made Easy, Shuzhe Wang et al. (2023). DUSt3R extends generalizable multi-view 3D reconstruction by regressing dense geometric 3D pointmaps directly from uncalibrated image pairs using transformers.
- Paper: VGGT: Visual Geometry Grounded Transformer, Jianyuan Wang et al. (2025). VGGT scales up transformer-based feed-forward multi-view 3D vision to simultaneously predict camera poses, depth maps, and point clouds across many images.
- Paper: Depth Pro: Sharp Monocular Metric Depth in Less Than a Second, Alexey Bochkovskiy et al. (2025). Depth Pro builds on modern transformer architectures to deliver zero-shot, sharp metric depth estimation applicable to downstream novel view synthesis and reconstruction workflows.
- Paper: Free3D: Consistent Novel View Synthesis Without 3D Representation, Chuanxia Zheng et al. (2024). Free3D explores an alternative paradigm by synthesizing consistent multi-view perspectives via ray-conditioned 2D diffusion without building explicit or implicit 3D volumes.
