Direct Voxel Grid Optimization: Super-fast Convergence for Radiance Fields Reconstruction

Cheng SunMin SunHwann-Tzong Chen

article2021CVPR1,398 citations

Introduces direct voxel grid optimization with post-activation interpolation to train high-quality radiance fields for novel view synthesis in under 15 minutes on a single GPU.

Listen

Synthesizing novel, free-viewpoint 3D views from a set of 2D images is critical for immersive consumer experiences, such as digital product showcases and virtual navigation. While Neural Radiance Fields (NeRF) have revolutionized visual quality in 3D scene reconstruction, their practical adoption is severely hindered by computational bottlenecks. Standard NeRF models rely on deep multilayer perceptron networks that require between 10 to 20 hours—and sometimes multiple days—of training per scene on high-end hardware, alongside sluggish rendering speeds. Existing acceleration techniques either demand extensive multi-day cross-scene pre-training, depend on external depth data, or require converting pre-trained implicit models into explicit data structures, leaving per-scene training times bottlenecked.

The article demonstrates an explicit direct voxel grid optimization method that reconstructs high-fidelity 3D radiance fields directly from scratch in under 15 minutes on a single consumer-grade graphics processing unit (GPU). The main objective is to match or exceed NeRF’s visual fidelity while reducing per-scene optimization time by roughly two orders of magnitude, eliminating the need for pre-training or auxiliary depth inputs.

To achieve this, the approach replaces deep continuous coordinate networks with explicit, discretized 3D voxel grids for scene geometry (volume density) paired with a shallow two-layer network for view-dependent color effects. Training occurs in two successive stages: coarse geometry search followed by fine-detail reconstruction. The method incorporates two core innovations: post-activated interpolation—applying mathematical activation functions after grid interpolation rather than before—and two optimization priors (near-zero density initialization and view-count-based learning rate scaling) to prevent geometry errors. Evaluation was conducted across five standard inward-facing benchmark datasets (Synthetic-NeRF, Synthetic-NSVF, BlendedMVS, Tanks&Temples, and DeepVoxels) using a single NVIDIA RTX 2080 Ti GPU.

Key findings show that the proposed method achieves convergence speeds about 100 times faster than NeRF, dropping per-scene training time to roughly 14 to 18 minutes while simultaneously accelerating test-time rendering by approximately 45 times (0.64 seconds versus 29 seconds per image). Across all five benchmark datasets, the reconstructed visual quality consistently matches or surpasses the original NeRF baseline in standard image metrics (such as Peak Signal-to-Noise Ratio and Structural Similarity Index). Furthermore, mathematical proofs and experiments verify that post-activated interpolation allows compact grid resolutions of only 160 cubed voxels to capture sharp surface boundaries, whereas prior explicit grid methods require resolutions ranging from 512 cubed to 1300 cubed voxels. Ablation studies confirmed that initializing densities close to zero is essential to prevent false semi-transparent artifacts near the camera.

These results demonstrate that high-quality 3D view synthesis does not inherently require deep implicit networks or days of compute. For engineering and business stakeholders, this transition dramatically reduces cloud compute costs, shortens development timelines from days to minutes per asset, and enables scalable, on-demand 3D asset generation workflows without complex multi-stage training pipelines.

Organizations developing 3D reconstruction and rendering pipelines should adopt direct voxel grid optimization for bounded, object-centric reconstruction tasks to realize immediate time and cost efficiencies. When evaluating implementation trade-offs, standard grid configurations (160 cubed voxels) offer the optimal speed-to-quality balance, while larger grids (256 cubed voxels) can be selected when maximum visual precision is required at a minor time cost (approximately 22 minutes).

A primary limitation of this work is that it targets bounded, inward-facing scenes and does not yet support forward-facing or unbounded 360-degree environments. Additionally, real-world captures with inconsistent lighting or camera calibration errors yield slightly smaller sharpness gains due to multi-view data uncertainties. Nevertheless, because the findings are mathematically validated and demonstrated across multiple standardized datasets, stakeholders can have high confidence in adopting these techniques for bounded 3D reconstruction applications.

Cover for Direct Voxel Grid Optimization: Super-fast Convergence for Radiance Fields Reconstruction

Abstract

We present a super-fast convergence approach to reconstructing the per-scene radiance field from a set of images that capture the scene with known poses. This task, which is often applied to novel view synthesis, is recently revolutionized by Neural Radiance Field (NeRF) for its state-of-the-art quality and flexibility. However, NeRF and its variants require a lengthy training time ranging from hours to days for a single scene. In contrast, our approach achieves NeRF-comparable quality and converges rapidly from scratch in less than 15 minutes with a single GPU. We adopt a representation consisting of a density voxel grid for scene geometry and a feature voxel grid with a shallow network for complex view-dependent appearance. Modeling with explicit and discretized volume representations is not new, but we propose two simple yet non-trivial techniques that contribute to fast convergence speed and high-quality output. First, we introduce the post-activation interpolation on voxel density, which is capable of producing sharp surfaces in lower grid resolution. Second, direct voxel density optimization is prone to suboptimal geometry solutions, so we robustify the optimization process by imposing several priors. Finally, evaluation on five inward-facing benchmarks shows that our method matches, if not surpasses, NeRF's quality, yet it only takes about 15 minutes to train from scratch for a new scene.

Table of Contents

  • 1 Introduction
  • 2 Related work
  • 3 Preliminaries
  • 4 Post-activated density voxel grid
  • 5 Fast and direct voxel grid optimization
  • 5.1 Coarse geometry searching
  • 5.2 Fine detail reconstruction
  • 6 Experiments
  • 6.1 Datasets
  • 6.2 Implementation details
  • 6.3 Comparisons
  • 6.4 Ablation studies
  • 7 Conclusion
  • H Additional implementation details
  • I Additional ablation experiments
  • J Main ablation studies details
  • K Additional training time details
  • K.1 Training time comparisons
  • K.2 The detailed training time of our model
  • L Per-scene analysis
  • L.1 Per-scene quantitative results
  • L.2 Per-scene qualitative results
  • M Derivation of low-density initialization
  • N Derivation of post-activation
  • N.1 Derivation for a 1D grid cell
  • N.2 Derivation for a 2D grid cell
  • N.3 Derivation for a 3D grid cell
  • N.4 Future extensions
  • References

Knowls

  1. Knowl 1 — Direct Voxel Grid Optimization Framework

    model/method

    Direct Voxel Grid Optimization (DVGO) is an approach for per-scene radiance field reconstruction from calibrated multi-view images that optimizes explicit discretized voxel grids directly via gradient descent. Unlike coordinate-based implicit models (such as NeRF) or post-hoc baked representations (such as PlenOctrees or FastNeRF), DVGO requires neither deep coordinate multilayer perceptrons (MLPs), cross-scene pretraining, nor multi-view stereo depth priors, training directly from scratch on a single scene in under 15 minutes.

    The reconstruction is structured into two sequential stages:

    1. Coarse Geometry Searching: Identifies the coarse occupied volume using a low-resolution density voxel grid V(density)(c)∈R1×Nx(c)×Ny(c)×Nz(c)V^{(\text{density})(c)} \in \mathbb{R}^{1 \times N_x^{(c)} \times N_y^{(c)} \times N_z^{(c)}} and a diffuse color grid V(rgb)(c)∈R3×Nx(c)×Ny(c)×Nz(c)V^{(\text{rgb})(c)} \in \mathbb{R}^{3 \times N_x^{(c)} \times N_y^{(c)} \times N_z^{(c)}} aligned with a bounding box (BBox) enclosing the camera frustums.
    2. Fine Detail Reconstruction: Freezes the coarse density grid to determine known free space, bounds the remaining unknown region with a tighter BBox, and jointly optimizes a high-resolution density grid V(density)(f)V^{(\text{density})(f)} with a hybrid explicit-implicit appearance model comprising a feature voxel grid V(feat)(f)V^{(\text{feat})(f)} and a shallow MLP for view-dependent color emission.

    Along any camera ray r(t)=o+tdr(t) = o + t d with origin oo and direction dd, KK query points are sampled and accumulated into an estimated pixel color C^(r)\hat{C}(r) via numerical quadrature: C^(r)=∑i=1KTiαici+TK+1cbg\hat{C}(r) = \sum_{i=1}^K T_i \alpha_i c_i + T_{K+1} c_{\text{bg}} αi=1−exp⁡(−σiδi)\alpha_i = 1 - \exp(-\sigma_i \delta_i) Ti=∏j=1i−1(1−αj)T_i = \prod_{j=1}^{i-1} (1 - \alpha_j) where σi≥0\sigma_i \ge 0 is the queried volume density at sample ii, ci∈[0,1]3c_i \in [0, 1]^3 is the queried color, δi\delta_i is the distance between adjacent query points, TiT_i is the accumulated transmittance from the near plane to point ii, and cbgc_{\text{bg}} is a pre-defined background color.

  2. Knowl 2 — Post-Activated Density Voxel Grid Interpolation

    model/method

    To enable explicit voxel grids to represent sharp object boundaries at moderate resolutions without over-smoothing, post-activation applies non-linear density activation functions strictly after trilinear spatial interpolation of the raw density grid.

    Let V(density)∈R1×Nx×Ny×NzV^{(\text{density})} \in \mathbb{R}^{1 \times N_x \times N_y \times N_z} store unconstrained raw density values σ¨∈R\ddot{\sigma} \in \mathbb{R}. For any continuous 3D coordinate x∈R3x \in \mathbb{R}^3, the density activation uses shifted Softplus: σ(x)=softplus(σ¨(x))=log⁡(1+exp⁡(σ¨(x)+b))\sigma(x) = \text{softplus}(\ddot{\sigma}(x)) = \log(1 + \exp(\ddot{\sigma}(x) + b)) where b∈Rb \in \mathbb{R} is a bias parameter. The termination opacity α(post)(x)\alpha^{(\text{post})}(x) for a ray step size δ\delta is evaluated as: α(post)(x)=1−exp⁡(−σ(x)δ)=1−(1+exp⁡(interp(x,V(density))+b))−δ\alpha^{(\text{post})}(x) = 1 - \exp(-\sigma(x)\delta) = 1 - \left(1 + \exp(\text{interp}(x, V^{(\text{density})}) + b)\right)^{-\delta} where interp(x,V(density))\text{interp}(x, V^{(\text{density})}) denotes trilinear interpolation of the raw voxel grid at position xx.

    In contrast to pre-activation (interpolating pre-computed opacities) and in-activation (interpolating softplus-activated densities before computing alpha): α(pre)(x)=interp(x,1−(1+exp⁡(V(density)+b))−δ)\alpha^{(\text{pre})}(x) = \text{interp}\left(x, 1 - (1 + \exp(V^{(\text{density})} + b))^{-\delta}\right) α(in)(x)=1−exp⁡(−interp(x,softplus(V(density)))δ)\alpha^{(\text{in})}(x) = 1 - \exp\left(-\text{interp}\left(x, \text{softplus}(V^{(\text{density})})\right)\delta\right) post-activation avoids smooth blending of activated densities and allows a single voxel grid cell to form a sharp step-like decision boundary. Using shifted Softplus instead of ReLU prevents dying zero-gradient voxels if density parameters become negative.

  3. Knowl 3 — Exact Linear Surface Representation in Single Post-Activated Grid Cells

    theoretical result

    Post-activated interpolation allows a single grid cell to represent a sharp planar decision boundary arbitrarily closely.

    For a 1D grid cell over x∈[0,1]x \in [0, 1] storing raw boundary densities a,b∈Ra, b \in \mathbb{R} at x=0x=0 and x=1x=1, the post-activated opacity function with volume rendering step size δ>0\delta > 0 is: S(x;a,b)=1−(1+exp⁡(a(1−x)+bx))−δS(x; a, b) = 1 - \left(1 + \exp(a(1 - x) + bx)\right)^{-\delta} Let T(x;c)=1(x>c)T(x; c) = \mathbf{1}(x > c) be a step function defining an exact surface boundary at c∈(0,1)c \in (0, 1). For any error threshold ϵ∈(0,1)\epsilon \in (0, 1) and boundary transition width Δ∈(0,min⁡(c,1−c))\Delta \in (0, \min(c, 1-c)), there exist finite grid values aa and bb such that: ∣S(x;a,b)−T(x;c)∣≤ϵfor all ∣x−c∣≥Δ|S(x; a, b) - T(x; c)| \le \epsilon \quad \text{for all } |x - c| \ge \Delta Under the symmetry constraint S(c;a,b)=0.5S(c; a, b) = 0.5, the grid values satisfy the exact linear relation: b=ac−1c+log⁡(21/δ−1)cb = a \frac{c - 1}{c} + \frac{\log(2^{1/\delta} - 1)}{c} with the upper bound on aa given for δ<1\delta < 1 by: a≤log⁡(21/δ−1)c+ΔΔ−log⁡(ϵ−1/δ−1)cΔa \le \log(2^{1/\delta} - 1)\frac{c + \Delta}{\Delta} - \log\left(\epsilon^{-1/\delta} - 1\right)\frac{c}{\Delta} and for δ≥1\delta \ge 1 by: a≤log⁡((1−ϵ)−1/δ−1)cΔ−log⁡(21/δ−1)c−ΔΔa \le \log\left((1 - \epsilon)^{-1/\delta} - 1\right)\frac{c}{\Delta} - \log(2^{1/\delta} - 1)\frac{c - \Delta}{\Delta} Because aa and bb are linear functions of cc, this result generalizes directly to 2D bilinear and 3D trilinear cells, guaranteeing that post-activation can model arbitrary linear boundaries crossing a grid cell without requiring cell subdivisions.

  4. Knowl 4 — Low-Density Transmittance Initialization

    equation

    To prevent direct voxel density optimization from becoming trapped in suboptimal "cloudy" local minima where high densities near camera near planes absorb all incoming ray transmittance and block deeper scene geometry from receiving gradient updates, voxel density grids are initialized such that the accumulated transmittance across free space remains near 1 at the start of training.

    All raw density values σ¨\ddot{\sigma} in the coarse grid V(density)(c)V^{(\text{density})(c)} are set to 0. Given a target single-voxel alpha value α(init)≪1\alpha^{(\text{init})} \ll 1 (default α(init)(c)=10−6\alpha^{(\text{init})(c)} = 10^{-6} for coarse search and 10−210^{-2} for fine reconstruction) and physical voxel dimension ss, the softplus bias shift bb is set to: b=log⁡((1−α(init))−1s−1)b = \log\left(\left(1 - \alpha^{(\text{init})}\right)^{-\frac{1}{s}} - 1\right) When a camera ray traverses a distance equal to one voxel length ss, the accumulated transmittance TT decays by exactly (1−α(init))≈1(1 - \alpha^{(\text{init})}) \approx 1, allowing optical rays to penetrate the entire volume and backpropagate valid photometric gradients across all sampled depth points at step zero.

  5. Knowl 5 — View-Count-Based Learning Rate Modulation

    model/method

    In direct density optimization, voxels visible to only a small subset of camera viewpoints are susceptible to generating spurious floating density artifacts that overfit specific camera views at the expense of global multi-view consistency. To suppress these artifacts, the optimization modifies the learning rate of each voxel based on its training view visibility.

    For each grid vertex jj in the coarse density grid V(density)(c)V^{(\text{density})(c)}, the algorithm counts the number of training camera viewpoints njn_j from which vertex jj is unoccluded. The base learning rate ηbase\eta_{\text{base}} (set to 0.10.1 for voxel grids) is scaled per vertex as: ηj=ηbase⋅njnmax⁡\eta_j = \eta_{\text{base}} \cdot \frac{n_j}{n_{\max}} where nmax⁡=max⁡knkn_{\max} = \max_k n_k is the maximum view count across all grid points in the bounding box. Density points seen by fewer training views receive proportionally smaller gradient updates, preventing the allocation of redundant private voxels.

  6. Knowl 6 — Hybrid Explicit-Implicit View-Dependent Color Representation

    model/method

    To model complex non-Lambertian view-dependent color emissions without incurring the slow training speed of deep MLPs or the memory bloat and quality degradation of purely explicit color grids, DVGO combines an explicit feature voxel grid with a compact multilayer perceptron in the fine reconstruction stage.

    The fine stage appearance model comprises:

    1. A feature voxel grid V(feat)(f)∈RD×Nx(f)×Ny(f)×Nz(f)V^{(\text{feat})(f)} \in \mathbb{R}^{D \times N_x^{(f)} \times N_y^{(f)} \times N_z^{(f)}}, where D=12D = 12 is the feature channel dimension.
    2. A shallow MLP parameterized by Θ\Theta, consisting of 2 hidden layers with 128 channels and Sigmoid output activation.

    For a 3D query point x∈R3x \in \mathbb{R}^3 and viewing direction unit vector d∈S2d \in \mathbb{S}^2, the view-dependent RGB color emission c(f)(x,d)∈[0,1]3c^{(f)}(x, d) \in [0, 1]^3 is queried as: c(f)(x,d)=MLPΘ(rgb)([interp(x,V(feat)(f)),γx(x),γd(d)])c^{(f)}(x, d) = \text{MLP}^{(\text{rgb})}_\Theta\left(\left[\text{interp}\left(x, V^{(\text{feat})(f)}\right), \gamma_x(x), \gamma_d(d)\right]\right) where interp(x,V(feat)(f))∈R12\text{interp}\left(x, V^{(\text{feat})(f)}\right) \in \mathbb{R}^{12} is the trilinearly interpolated feature vector, γx(x)\gamma_x(x) is positional encoding with kx=5k_x = 5 frequency bands, and γd(d)\gamma_d(d) is positional encoding with kd=4k_d = 4 frequency bands.

  7. Knowl 7 — Two-Stage Coarse-to-Fine Reconstruction Algorithm

    algorithm

    The complete DVGO radiance field reconstruction pipeline operates through two optimization stages with progressive grid scaling and two-tier free space skipping:

    Input: Calibrated multi-view RGB images with camera poses, coarse voxel budget M(c)=1003M^{(c)} = 100^3, fine voxel budget M(f)=1603M^{(f)} = 160^3, checkpoints pg_ckpt={1000,2000,3000}\text{pg\_ckpt} = \{1000, 2000, 3000\}, thresholds τ(c)=10−3\tau^{(c)} = 10^{-3}, τ(f)=10−4\tau^{(f)} = 10^{-4}.
    Output: Optimized density grid V(density)(f)V^{(\text{density})(f)}, feature grid V(feat)(f)V^{(\text{feat})(f)}, and color network MLPΘ(rgb)\text{MLP}^{(\text{rgb})}_\Theta.
    1. Coarse Geometry Search:
       Compute BBox tightly enclosing training camera frustums.
       Initialize V(density)(c)V^{(\text{density})(c)} to 0 with bias bb from low-density initialization (α(init)(c)=10−6\alpha^{(\text{init})(c)} = 10^{-6}).
       Initialize diffuse color grid V(rgb)(c)V^{(\text{rgb})(c)}.
       Compute per-voxel visibility view counts njn_j and scale learning rates ηj=0.1⋅(nj/nmax⁡)\eta_j = 0.1 \cdot (n_j / n_{\max}).
       for step = 1 to 10000 do
           Sample batch of 8192 rays with step size δ(c)=0.5s(c)\delta^{(c)} = 0.5 s^{(c)}.
           Render colors C^(r)\hat{C}(r) via volume rendering quadrature.
           Update V(density)(c)V^{(\text{density})(c)} and V(rgb)(c)V^{(\text{rgb})(c)} by minimizing photometric loss plus background entropy loss.
       end for
       Freeze V(density)(c)V^{(\text{density})(c)}.
    2. Fine Detail Reconstruction:
       Query frozen V(density)(c)V^{(\text{density})(c)} over a dense grid to identify unknown space where post-activated α≥τ(c)\alpha \ge \tau^{(c)}.
       Compute tightened BBox enclosing unknown space.
       Initialize V(density)(f)V^{(\text{density})(f)} and V(feat)(f)V^{(\text{feat})(f)} with initial voxel count ⌊M(f)/2∣pg_ckpt∣⌋=⌊1603/8⌋\lfloor M^{(f)} / 2^{|\text{pg\_ckpt}|} \rfloor = \lfloor 160^3 / 8 \rfloor.
       for step = 1 to 20000 do
           if step in pg_ckpt\text{pg\_ckpt} then
               Double grid resolutions along each axis via trilinear interpolation until reaching M(f)M^{(f)}.
           end if
           Sample batch of 8192 rays with step size δ(f)=0.5s(f)\delta^{(f)} = 0.5 s^{(f)}.
           Clip ray segments to tightened BBox.
           Skip query points if coarse alpha <τ(c)< \tau^{(c)} (known free space) or fine alpha <τ(f)< \tau^{(f)}.
           Query remaining points via V(density)(f)V^{(\text{density})(f)} and hybrid feature grid with MLPΘ(rgb)\text{MLP}^{(\text{rgb})}_\Theta.
           Update parameters using Adam optimizer with exponential learning rate decay.
       end for
       return V(density)(f)V^{(\text{density})(f)}, V(feat)(f)V^{(\text{feat})(f)}, MLPΘ(rgb)\text{MLP}^{(\text{rgb})}_\Theta
  8. Knowl 8 — Novel View Synthesis Quality and Speed Benchmarks

    data/table

    DVGO was evaluated across five inward-facing benchmarks: Synthetic-NeRF (8 scenes), Synthetic-NSVF (8 scenes), BlendedMVS (4 scenes), Tanks & Temples (5 scenes), and DeepVoxels (4 scenes). On a single NVIDIA RTX 2080 Ti GPU, DVGO trains in roughly 15 minutes from scratch per scene, outperforming original NeRF (which requires 1–2 days on an NVIDIA V100 GPU) and achieving competitive quality with prior fast neural volumetric methods.

    Method Synthetic-NeRF Synthetic-NSVF BlendedMVS Tanks Temples
    PSNR↑\uparrow SSIM↑\uparrow LPIPS↓\downarrow PSNR↑\uparrow SSIM↑\uparrow LPIPS↓\downarrow PSNR↑\uparrow SSIM↑\uparrow LPIPS↓\downarrow PSNR↑\uparrow SSIM↑\uparrow LPIPS↓\downarrow
    NeRF 31.01 0.947 0.081 30.81 0.952 0.043 24.15 0.828 0.192 25.78 0.864 0.198
    JaxNeRF 31.69 0.953 0.068 - - - - - - 27.94 0.904 0.168
    NSVF 31.75 0.953 0.047 35.18 0.979 0.015 26.89 0.898 0.114 28.48 0.901 0.155
    PlenOctrees 31.71 0.958 0.053 - - - - - - 27.99 0.917 0.131
    KiloNeRF 31.00 0.950 0.030 33.37 0.970 0.020 27.39 0.920 0.060 28.41 0.910 0.090
    DVGO (M(f)=1603M^{(f)}=160^3) 31.95 0.957 0.035 35.08 0.975 0.019 28.02 0.922 0.075 28.41 0.911 0.148
    DVGO (M(f)=2563M^{(f)}=256^3) 32.80 0.961 0.027 36.21 0.980 0.012 28.64 0.933 0.052 28.82 0.920 0.124

    Training time and hardware comparisons on Synthetic-NeRF:

    • NeRF: 31.01 PSNR, no pre-training, 1–2 days per scene on an NVIDIA V100 GPU.
    • MVSNeRF: 27.21 PSNR, 30 hours pre-training, 15 minutes per-scene optimization on RTX 2080 Ti.
    • IBRNet: 28.14 PSNR, 1 day pre-training on 8×\timesV100, 6 hours per-scene optimization on V100.
    • NeuRay: 32.42 PSNR, 2 days pre-training, 23 hours per-scene optimization on RTX 2080 Ti.
    • Mip-NeRF (early stopped): 30.85 PSNR, no pre-training, 6 hours per scene on RTX 2080 Ti.
    • DVGO (1603160^3): 31.95 PSNR, no pre-training, 15 minutes per scene on RTX 2080 Ti.
    • DVGO (2563256^3): 32.80 PSNR, no pre-training, 22 minutes per scene on RTX 2080 Ti.

    In inference, DVGO renders an 800×800800 \times 800 image in 0.64 seconds ( ≈45×\,\approx 45\times faster than NeRF's 29 seconds per image).

  9. Knowl 9 — Ablation Study on Post-Activation and Optimization Priors

    data/table

    Ablation studies demonstrate the contribution of post-activated interpolation, low-density initialization, and view-count learning rate scaling across four datasets (Synthetic-NeRF, Synthetic-NSVF, BlendedMVS, Tanks & Temples):

    Interpolation Scheme Syn.-NeRF (PSNR) Syn.-NSVF (PSNR) BlendedMVS (PSNR) Tanks Temples (PSNR)
    Nearest-Neighbor 28.61 (-2.77) 28.86 (-6.22) 25.49 (-2.48) 26.39 (-1.27)
    Pre-activation 30.84 (-0.55) 32.66 (-2.41) 27.39 (-0.58) 27.44 (-0.21)
    In-activation 29.91 (-1.48) 32.42 (-2.66) 27.29 (-0.68) 27.52 (-0.13)
    Post-activation (Ours) 31.39 35.08 27.97 27.66
    α(init)(c)\alpha^{(\text{init})(c)} View-count LR Syn.-NeRF (PSNR) Syn.-NSVF (PSNR) BlendedMVS (PSNR) Tanks Temples (PSNR)
    None No 28.88 (-2.51) 25.12 (-9.96) 22.17 (-5.79) 25.33 (-2.33)
    10−310^{-3} No 30.96 (-0.42) 27.24 (-7.84) 23.17 (-4.79) 26.04 (-1.61)
    10−410^{-4} No 31.29 (-0.09) 31.05 (-4.03) 26.09 (-1.88) 27.60 (-0.05)
    10−510^{-5} No 31.41 (+0.02) 35.04 (-0.04) 27.36 (-0.61) 27.63 (-0.02)
    10−610^{-6} No 31.39 35.08 27.97 27.66
    10−610^{-6} Yes 31.40 (+0.01) 35.03 (-0.04) 27.37 (-0.60) 27.59 (-0.07)
    10−710^{-7} No 31.36 (-0.02) 35.03 (-0.05) 27.73 (-0.23) 27.59 (-0.06)

    Key takeaways from the ablation data:

    1. Post-activation improves reconstruction PSNR by up to 6.22 dB over nearest-neighbor and up to 2.41–2.66 dB over pre- and in-activation schemes at identical grid resolutions.
    2. Low-density initialization is necessary to avoid catastrophic failure: omitting initialization causes up to a 9.96 dB drop in PSNR.
    3. View-count learning rate modulation suppresses floating artifacts in regions observed by few cameras.
  10. Knowl 10 — Limitations in Unbounded and Forward-Facing Scenes

    limitation

    DVGO is designed specifically for bounded inward-facing capture setups where the entire scene of interest can be tightly enclosed within a finite 3D bounding box. It does not accommodate:

    1. Unbounded 360∘360^\circ outdoor environments: Discretizing large background volumes with a dense uniform voxel grid causes cubic memory scaling (O(N3)O(N^3)), leading to out-of-memory errors on standard GPUs.
    2. Forward-facing camera trajectories with infinite depth: The uniform voxel parametrization lacks the non-linear projective parameterizations (such as normalized device coordinates or inverted sphere representations) necessary to cover scenes extending to infinity without prohibitive memory allocations.

Coverage note — None omitted; all core technical contributions, mathematical formulations, optimization algorithms, theoretical properties, empirical benchmarks, and stated limitations are included.

References

  1. 1.Jonathan T. Barron, Ben Mildenhall, Matthew Tancik, Peter Hedman, Ricardo Martin-Brualla, and Pratul P. Srinivasan. Mip-nerf: A multiscale representation for anti-aliasing neural radiance fields. In ICCV, 2021. 1, 3, 4, 6, 7, 9, 11, 12, 13
  2. 2.Sai Bi, Zexiang Xu, Pratul P. Srinivasan, Ben Mildenhall, Kalyan Sunkavalli, Milos Hasan, Yannick Hold-Geoffroy, David J. Kriegman, and Ravi Ramamoorthi. Neural reflectance fields for appearance acquisition. arxiv CS.CV 2106.01970, 2020. 2
  3. 3.Mark Boss, Raphael Braun, Varun Jampani, Jonathan T. Barron, Ce Liu, and Hendrik P. A. Lensch. Nerd: Neural reflectance decomposition from image collections. In ICCV, 2021. 2
  4. 4.Chris Buehler, Michael Bosse, Leonard McMillan, Steven J. Gortler, and Michael F. Cohen. Unstructured lumigraph rendering. In SIGGRAPH, 2001. 2
  5. 5.Eric R. Chan, Marco Monteiro, Petr Kellnhofer, Jiajun Wu, and Gordon Wetzstein. Pi-gan: Periodic implicit generative adversarial networks for 3d-aware image synthesis. In CVPR, 2021. 2
  6. 6.Anpei Chen, Zexiang Xu, Fuqiang Zhao, Xiaoshuai Zhang, Fanbo Xiang, Jingyi Yu, and Hao Su. Mvsnerf: Fast generalizable radiance field reconstruction from multi-view stereo. In ICCV, 2021. 1, 3, 7
  7. 7.Abe Davis, Marc Levoy, and Fredo Durand. Unstructured light fields. Comput. Graph. Forum, 2012. 2
  8. 8.Paul E. Debevec, Camillo J. Taylor, and Jitendra Malik. Modeling and rendering architecture from photographs: A hybrid geometry- and image-based approach. In SIGGRAPH, 1996. 2
  9. 9.Boyang Deng, Jonathan T. Barron, and Pratul P. Srinivasan. JaxNeRF: an efficient JAX implementation of NeRF, 2020. 6, 7, 11, 13, 17
  10. 10.Kangle Deng, Andrew Liu, Jun-Yan Zhu, and Deva Ramanan. Depth-supervised nerf: Fewer views and faster training for free. arxiv CS.CV 2107.02791, 2021. 1, 3
  11. 11.Helisa Dhamo, Keisuke Tateno, Iro Laina, Nassir Navab, and Federico Tombari. Peeking behind objects: Layered depth prediction from a single image. Pattern Recognit. Lett., 2019. 2
  12. 12.John Flynn, Michael Broxton, Paul E. Debevec, Matthew DuVall, Graham Fyffe, Ryan S. Overbeck, Noah Snavely, and Richard Tucker. Deepview: View synthesis with learned gradient descent. In CVPR, 2019. 2
  13. 13.Guy Gafni, Justus Thies, Michael Zollhofer, and Matthias Nießner. Dynamic neural radiance fields for monocular 4d facial avatar reconstruction. In CVPR, 2021. 2
  14. 14.Chen Gao, Ayush Saraf, Johannes Kopf, and Jia-Bin Huang. Dynamic view synthesis from dynamic monocular video. In ICCV, 2021. 2
  15. 15.Stephan J. Garbin, Marek Kowalski, Matthew Johnson, Jamie Shotton, and Julien P. C. Valentin. Fastnerf: High-fidelity neural rendering at 200fps. In ICCV, 2021. 1, 2, 3, 7, 13
  16. 16.Steven J. Gortler, Radek Grzeszczuk, Richard Szeliski, and Michael F. Cohen. The lumigraph. In SIGGRAPH, 1996. 2
  17. 17.Tong He, John P. Collomosse, Hailin Jin, and Stefano Soatto. Deepvoxels++: Enhancing the fidelity of novel view synthesis from 3d voxel embeddings. In Hiroshi Ishikawa, Cheng-Lin Liu, Tomas Pajdla, and Jianbo Shi, editors, ACCV, 2020. 2, 18
  18. 18.Peter Hedman, Pratul P. Srinivasan, Ben Mildenhall, Jonathan T. Barron, and Paul E. Debevec. Baking neural radiance fields for real-time view synthesis. In ICCV, 2021. 1, 2, 3, 5, 7, 10, 13
  19. 19.Yoonwoo Jeong, Seokjun Ahn, Christopher Choy, Animashree Anandkumar, Minsu Cho, and Jaesik Park. Self-calibrating neural radiance fields. In ICCV, 2021. 2
  20. 20.Diederik P. Kingma and Jimmy Ba. Adam: A method for stochastic optimization. In ICLR, 2015. 6
  21. 21.Arno Knapitsch, Jaesik Park, Qian-Yi Zhou, and Vladlen Koltun. Tanks and temples: benchmarking large-scale scene reconstruction. ACM Trans. Graph., 2017. 6, 11, 12, 17
  22. 22.Adam R. Kosiorek, Heiko Strathmann, Daniel Zoran, Pol Moreno, Rosalia Schneider, Sona Mokra, and Danilo Jimenez Rezende. Nerf-vae: A geometry aware 3d scene generative model. In ICML, 2021. 2
  23. 23.Anat Levin and Fredo Durand. Linear view synthesis using a dimensionality gap light field prior. In CVPR, 2010. 2
  24. 24.Marc Levoy and Pat Hanrahan. Light field rendering. In SIGGRAPH, 1996. 2
  25. 25.Zhengqi Li, Simon Niklaus, Noah Snavely, and Oliver Wang. Neural scene flow fields for space-time view synthesis of dynamic scenes. In CVPR, 2021. 2
  26. 26.Zhengqi Li, Wenqi Xian, Abe Davis, and Noah Snavely. Crowdsampling the plenoptic function. In ECCV, 2020. 2
  27. 27.Chen-Hsuan Lin, Wei-Chiu Ma, Antonio Torralba, and Simon Lucey. BARF: bundle-adjusting neural radiance fields. In ICCV, 2021. 2
  28. 28.Yen-Chen Lin, Pete Florence, Jonathan T. Barron, Alberto Rodriguez, Phillip Isola, and Tsung-Yi Lin. inerf: Inverting neural radiance fields for pose estimation. In IROS, 2021. 2
  29. 29.David B. Lindell, Julien N. P. Martel, and Gordon Wetzstein. Autoint: Automatic integration for fast neural volume rendering. In CVPR, 2021. 1, 7, 13
  30. 30.Lingjie Liu, Jiatao Gu, Kyaw Zaw Lin, Tat-Seng Chua, and Christian Theobalt. Neural sparse voxel fields. In NeurIPS, 2020. 1, 2, 3, 5, 6, 7, 11, 12, 13, 15, 16, 17
  31. 31.Yuan Liu, Sida Peng, Lingjie Liu, Qianqian Wang, Peng Wang, Christian Theobalt, Xiaowei Zhou, and Wenping Wang. Neural rays for occlusion-aware image-based rendering. arxiv CS.CV 2107.13421, 2021. 1, 3, 7, 10
  32. 32.Stephen Lombardi, Tomas Simon, Jason M. Saragih, Gabriel Schwartz, Andreas M. Lehrmann, and Yaser Sheikh. Neural volumes: learning dynamic renderable volumes from images. ACM Trans. Graph., 2019. 2, 7, 13, 15, 16, 17, 18
  33. 33.Ricardo Martin-Brualla, Noha Radwan, Mehdi S. M. Sajjadi, Jonathan T. Barron, Alexey Dosovitskiy, and Daniel Duckworth. Nerf in the wild: Neural radiance fields for unconstrained photo collections. In CVPR, 2021. 2
  34. 34.Nelson L. Max. Optical models for direct volume rendering. IEEE Trans. Vis. Comput. Graph., 1995. 3
  35. 35.Quan Meng, Anpei Chen, Haimin Luo, Minye Wu, Hao Su, Lan Xu, Xuming He, and Jingyi Yu. Gnerf: Gan-based neural radiance field without posed camera. In ICCV, 2021. 2
  36. 36.Ben Mildenhall, Pratul P. Srinivasan, Rodrigo Ortiz Cayon, Nima Khademi Kalantari, Ravi Ramamoorthi, Ren Ng, and Abhishek Kar. Local light field fusion: practical view synthesis with prescriptive sampling guidelines. ACM Trans. Graph., 2019. 2, 18
  37. 37.Ben Mildenhall, Pratul P. Srinivasan, Matthew Tancik, Jonathan T. Barron, Ravi Ramamoorthi, and Ren Ng. Nerf: Representing scenes as neural radiance fields for view synthesis. In ECCV, 2020. 1, 2, 3, 5, 6, 7, 9, 11, 12, 13, 14, 15, 16, 17, 18
  38. 38.Atsuhiro Noguchi, Xiao Sun, Stephen Lin, and Tatsuya Harada. Neural articulated radiance field. In ICCV, 2021. 2
  39. 39.Michael Oechsle, Songyou Peng, and Andreas Geiger. UNISURF: unifying neural implicit surfaces and radiance fields for multi-view reconstruction. In ICCV, 2021. 9
  40. 40.Keunhong Park, Utkarsh Sinha, Jonathan T. Barron, Sofien Bouaziz, Dan B. Goldman, Steven M. Seitz, and Ricardo Martin-Brualla. Deformable neural radiance fields. In ICCV, 2021. 2
  41. 41.Keunhong Park, Utkarsh Sinha, Peter Hedman, Jonathan T. Barron, Sofien Bouaziz, Dan B. Goldman, Ricardo Martin-Brualla, and Steven M. Seitz. Hypernerf: A higher-dimensional representation for topologically varying neural radiance fields. arxiv CS.CV 2106.13228, 2021. 2
  42. 42.Albert Pumarola, Enric Corona, Gerard Pons-Moll, and Francesc Moreno-Noguer. D-nerf: Neural radiance fields for dynamic scenes. In CVPR, 2021. 2
  43. 43.Daniel Rebain, Wei Jiang, Soroosh Yazdani, Ke Li, Kwang Moo Yi, and Andrea Tagliasacchi. Derf: Decomposed radiance fields. In CVPR, 2021. 1
  44. 44.Christian Reiser, Songyou Peng, Yiyi Liao, and Andreas Geiger. Kilonerf: Speeding up neural radiance fields with thousands of tiny mlps. In ICCV, 2021. 1, 3, 7, 13, 15, 16, 17
  45. 45.Katja Schwarz, Yiyi Liao, Michael Niemeyer, and Andreas Geiger. GRAF: generative radiance fields for 3d-aware image synthesis. In NeurIPS, 2020. 2
  46. 46.Jonathan Shade, Steven J. Gortler, Li-wei He, and Richard Szeliski. Layered depth images. In SIGGRAPH, 1998. 2
  47. 47.Lixin Shi, Haitham Hassanieh, Abe Davis, Dina Katabi, and Fredo Durand. Light field reconstruction using sparsity in the continuous fourier domain. ACM Trans. Graph., 2014. 2
  48. 48.Meng-Li Shih, Shih-Yang Su, Johannes Kopf, and Jia-Bin Huang. 3d photography using context-aware layered depth inpainting. In CVPR, 2020. 2
  49. 49.Vincent Sitzmann, Justus Thies, Felix Heide, Matthias Nießner, Gordon Wetzstein, and Michael Zollhofer. Deep-voxels: Learning persistent 3d feature embeddings. In CVPR, 2019. 2, 6, 11, 12, 18
  50. 50.Vincent Sitzmann, Michael Zollhofer, and Gordon Wetzstein. Scene representation networks: Continuous 3d-structure-aware neural scene representations. In NeurIPS, 2019. 7, 13, 15, 16, 17, 18
  51. 51.Pratul P. Srinivasan, Boyang Deng, Xiuming Zhang, Matthew Tancik, Ben Mildenhall, and Jonathan T. Barron. Nerv: Neural reflectance and visibility fields for relighting and view synthesis. In CVPR, 2021. 2
  52. 52.Pratul P. Srinivasan, Richard Tucker, Jonathan T. Barron, Ravi Ramamoorthi, Ren Ng, and Noah Snavely. Pushing the boundaries of view extrapolation with multiplane images. In CVPR, 2019. 2
  53. 53.Matthew Tancik, Ben Mildenhall, Terrance Wang, Divi Schmidt, Pratul P. Srinivasan, Jonathan T. Barron, and Ren Ng. Learned initializations for optimizing coordinate-based neural representations. In CVPR, 2021. 2
  54. 54.Matthew Tancik, Pratul P. Srinivasan, Ben Mildenhall, Sara Fridovich-Keil, Nithin Raghavan, Utkarsh Singhal, Ravi Ramamoorthi, Jonathan T. Barron, and Ren Ng. Fourier features let networks learn high frequency functions in low dimensional domains. In NeurIPS, 2020. 3
  55. 55.Justus Thies, Michael Zollhofer, and Matthias Nießner. Deferred neural rendering: image synthesis using neural textures. ACM Trans. Graph., 2019. 2
  56. 56.Edgar Tretschk, Ayush Tewari, Vladislav Golyanik, Michael Zollhofer, Christoph Lassner, and Christian Theobalt. Non-rigid neural radiance fields: Reconstruction and novel view synthesis of a deforming scene from monocular video. In ICCV, 2021. 2
  57. 57.Richard Tucker and Noah Snavely. Single-view view synthesis with multiplane images. In CVPR, 2020. 2
  58. 58.Shubham Tulsiani, Richard Tucker, and Noah Snavely. Layer-structured 3d scene inference via view synthesis. In ECCV, 2018. 2
  59. 59.Michael Waechter, Nils Moehrle, and Michael Goesele. Let there be color! large-scale texturing of 3d reconstructions. In ECCV, 2014. 2
  60. 60.Qianqian Wang, Zhicheng Wang, Kyle Genova, Pratul P. Srinivasan, Howard Zhou, Jonathan T. Barron, Ricardo Martin-Brualla, Noah Snavely, and Thomas A. Funkhouser. Ibrnet: Learning multi-view image-based rendering. In CVPR, 2021. 1, 3, 7, 18
  61. 61.Zhou Wang, Alan C. Bovik, Hamid R. Sheikh, and Eero P. Simoncelli. Image quality assessment: from error visibility to structural similarity. IEEE TIP, 2004. 6
  62. 62.Zirui Wang, Shangzhe Wu, Weidi Xie, Min Chen, and Victor Adrian Prisacariu. Nerf-: Neural radiance fields without known camera parameters. arxiv CS.CV 2102.07064, 2021. 2
  63. 63.Suttisak Wizadwongsa, Pakkapon Phongthawee, Jiraphon Yenphraphai, and Supasorn Suwajanakorn. Nex: Real-time view synthesis with neural basis expansion. In CVPR, 2021. 2, 3, 10
  64. 64.Daniel N. Wood, Daniel I. Azuma, Ken Aldinger, Brian Curless, Tom Duchamp, David Salesin, and Werner Stuetzle. Surface light fields for 3d photography. In SIGGRAPH, 2000. 2
  65. 65.Wenqi Xian, Jia-Bin Huang, Johannes Kopf, and Changil Kim. Space-time neural irradiance fields for free-viewpoint video. In CVPR, 2021. 2
  66. 66.Yao Yao, Zixin Luo, Shiwei Li, Jingyang Zhang, Yufan Ren, Lei Zhou, Tian Fang, and Long Quan. Blendedmvs: A large-scale dataset for generalized multi-view stereo networks. In CVPR, 2020. 6, 11, 12, 16
  67. 67.Alex Yu, Ruilong Li, Matthew Tancik, Hao Li, Ren Ng, and Angjoo Kanazawa. Plenoctrees for real-time rendering of neural radiance fields. In ICCV, 2021. 1, 2, 3, 5, 7, 10, 11, 13, 17
  68. 68.Alex Yu, Vickie Ye, Matthew Tancik, and Angjoo Kanazawa. pixelnerf: Neural radiance fields from one or few images. In CVPR, 2021. 3
  69. 69.Kai Zhang, Gernot Riegler, Noah Snavely, and Vladlen Koltun. Nerf++: Analyzing and improving neural radiance fields. arxiv CS.CV 2010.07492, 2020. 3
  70. 70.Richard Zhang, Phillip Isola, Alexei A. Efros, Eli Shechtman, and Oliver Wang. The unreasonable effectiveness of deep features as a perceptual metric. In CVPR, 2018. 6
  71. 71.Xiuming Zhang, Pratul P. Srinivasan, Boyang Deng, Paul E. Debevec, William T. Freeman, and Jonathan T. Barron. Nerfactor: Neural factorization of shape and reflectance under an unknown illumination. arxiv CS.CV 2106.01970, 2021. 2
  72. 72.Tinghui Zhou, Richard Tucker, John Flynn, Graham Fyffe, and Noah Snavely. Stereo magnification: learning view synthesis using multiplane images. ACM Trans. Graph., 2018. 2

Citation

MLA
Sun, C., et al. “Direct Voxel Grid Optimization: Super-fast Convergence for Radiance Fields Reconstruction”. arXiv, 2021, http://arxiv.org/abs/2111.11215v2.
APA
Sun, C., Sun, M., & Chen, H.-T. (2021). Direct Voxel Grid Optimization: Super-fast Convergence for Radiance Fields Reconstruction. arXiv. http://arxiv.org/abs/2111.11215v2
Chicago
Sun, C., M. Sun, and H.-T. Chen. 2021. “Direct Voxel Grid Optimization: Super-fast Convergence for Radiance Fields Reconstruction”. arXiv. http://arxiv.org/abs/2111.11215v2.
Harvard
Sun, C., Sun, M. and Chen, H.-T. (2021) “Direct Voxel Grid Optimization: Super-fast Convergence for Radiance Fields Reconstruction”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2111.11215v2.
Vancouver
1. Sun C, Sun M, Chen H-T (2021) Direct Voxel Grid Optimization: Super-fast Convergence for Radiance Fields Reconstruction. arXiv

BibTeX

@article{sun2021direct,
  title = {Direct Voxel Grid Optimization: Super-fast Convergence for Radiance Fields Reconstruction},
  author = {Sun, Cheng and Sun, Min and Chen, Hwann-Tzong},
  year = {2021},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2111.11215v2},
  eprint = {2111.11215}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE