Direct Voxel Grid Optimization: Super-fast Convergence for Radiance Fields Reconstruction
Cheng SunMin SunHwann-Tzong Chen
Introduces direct voxel grid optimization with post-activation interpolation to train high-quality radiance fields for novel view synthesis in under 15 minutes on a single GPU.
Synthesizing novel, free-viewpoint 3D views from a set of 2D images is critical for immersive consumer experiences, such as digital product showcases and virtual navigation. While Neural Radiance Fields (NeRF) have revolutionized visual quality in 3D scene reconstruction, their practical adoption is severely hindered by computational bottlenecks. Standard NeRF models rely on deep multilayer perceptron networks that require between 10 to 20 hours—and sometimes multiple days—of training per scene on high-end hardware, alongside sluggish rendering speeds. Existing acceleration techniques either demand extensive multi-day cross-scene pre-training, depend on external depth data, or require converting pre-trained implicit models into explicit data structures, leaving per-scene training times bottlenecked.
The article demonstrates an explicit direct voxel grid optimization method that reconstructs high-fidelity 3D radiance fields directly from scratch in under 15 minutes on a single consumer-grade graphics processing unit (GPU). The main objective is to match or exceed NeRF’s visual fidelity while reducing per-scene optimization time by roughly two orders of magnitude, eliminating the need for pre-training or auxiliary depth inputs.
To achieve this, the approach replaces deep continuous coordinate networks with explicit, discretized 3D voxel grids for scene geometry (volume density) paired with a shallow two-layer network for view-dependent color effects. Training occurs in two successive stages: coarse geometry search followed by fine-detail reconstruction. The method incorporates two core innovations: post-activated interpolation—applying mathematical activation functions after grid interpolation rather than before—and two optimization priors (near-zero density initialization and view-count-based learning rate scaling) to prevent geometry errors. Evaluation was conducted across five standard inward-facing benchmark datasets (Synthetic-NeRF, Synthetic-NSVF, BlendedMVS, Tanks&Temples, and DeepVoxels) using a single NVIDIA RTX 2080 Ti GPU.
Key findings show that the proposed method achieves convergence speeds about 100 times faster than NeRF, dropping per-scene training time to roughly 14 to 18 minutes while simultaneously accelerating test-time rendering by approximately 45 times (0.64 seconds versus 29 seconds per image). Across all five benchmark datasets, the reconstructed visual quality consistently matches or surpasses the original NeRF baseline in standard image metrics (such as Peak Signal-to-Noise Ratio and Structural Similarity Index). Furthermore, mathematical proofs and experiments verify that post-activated interpolation allows compact grid resolutions of only 160 cubed voxels to capture sharp surface boundaries, whereas prior explicit grid methods require resolutions ranging from 512 cubed to 1300 cubed voxels. Ablation studies confirmed that initializing densities close to zero is essential to prevent false semi-transparent artifacts near the camera.
These results demonstrate that high-quality 3D view synthesis does not inherently require deep implicit networks or days of compute. For engineering and business stakeholders, this transition dramatically reduces cloud compute costs, shortens development timelines from days to minutes per asset, and enables scalable, on-demand 3D asset generation workflows without complex multi-stage training pipelines.
Organizations developing 3D reconstruction and rendering pipelines should adopt direct voxel grid optimization for bounded, object-centric reconstruction tasks to realize immediate time and cost efficiencies. When evaluating implementation trade-offs, standard grid configurations (160 cubed voxels) offer the optimal speed-to-quality balance, while larger grids (256 cubed voxels) can be selected when maximum visual precision is required at a minor time cost (approximately 22 minutes).
A primary limitation of this work is that it targets bounded, inward-facing scenes and does not yet support forward-facing or unbounded 360-degree environments. Additionally, real-world captures with inconsistent lighting or camera calibration errors yield slightly smaller sharpness gains due to multi-view data uncertainties. Nevertheless, because the findings are mathematically validated and demonstrated across multiple standardized datasets, stakeholders can have high confidence in adopting these techniques for bounded 3D reconstruction applications.
- Paper: NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis, Ben Mildenhall et al. (2020). It establishes the core neural radiance field paradigm and volumetric rendering formulation that DVGO accelerates by replacing deep coordinate MLPs with explicit voxel grids.
- Paper: Neural Sparse Voxel Fields, Lingjie Liu et al. (2020). It introduces hybrid neural voxel representations to bound and accelerate radiance field rendering, paving the way for DVGO's purely explicit grid optimization.
- Paper: OctNet: Learning Deep 3D Representations at High Resolutions, Gernot Riegler et al. (2016). It provides foundational principles for handling spatial sparsity and efficient computation across volumetric grid representations in 3D deep learning.
- Paper: TensoRF: Tensorial Radiance Fields, Anpei Chen et al. (2022). It directly advances beyond dense voxel grid optimization by introducing low-rank tensor decompositions to dramatically reduce the memory footprint while maintaining fast convergence.
- Paper: Plenoxels: Radiance Fields without Neural Networks, Alex Yu et al. (2022). It further develops explicit radiance field reconstruction by optimizing spherical harmonic voxel grids entirely without neural networks, running on similar principles of fast grid-based optimization.
- Paper: Instant neural graphics primitives with a multiresolution hash encoding, Thomas Müller et al. (2022). It achieves even faster convergence and real-time evaluation using multiresolution hash encodings paired with tiny MLPs as an alternative to explicit voxel grids.
- Paper: 3D Gaussian Splatting for Real-Time Radiance Field Rendering, Bernhard Kerbl et al. (2023). It represents the next architectural evolution beyond explicit voxel grids by optimizing explicit 3D Gaussian primitives for rapid training and real-time rendering.
- Paper: Mip-NeRF 360: Unbounded Anti-Aliased Neural Radiance Fields, Jonathan T. Barron et al. (2022). It develops advanced spatial parameterization and contraction techniques to handle large unbounded 360-degree scenes, addressing a fundamental limitation of standard bounded voxel volumes.
- Paper: Ref-NeRF: Structured View-Dependent Appearance for Neural Radiance Fields, Dor Verbin et al. (2022). It introduces structured reflection parameterizations to accurately capture complex specular highlights and surface normals that shallow view-dependent networks in voxel methods struggle to resolve.
- Paper: 4D Gaussian Splatting for Real-Time Dynamic Scene Rendering, Guanjun Wu et al. (2023). It extends fast explicit radiance field representations to dynamic, time-varying scenes via temporal deformation fields.
- Paper: 2D Gaussian Splatting for Geometrically Accurate Radiance Fields, Binbin Huang et al. (2024). It builds upon explicit radiance representations by utilizing 2D planar disks to enforce geometrically accurate surface reconstruction while maintaining fast training speeds.
