PlenOctrees for Real-time Rendering of Neural Radiance Fields
Alex YuRuilong LiMatthew TancikHao LiRen NgAngjoo Kanazawa
Develops PlenOctrees, an octree-based representation that pre-tabulates neural radiance fields using spherical harmonics to achieve real-time rendering at over 150 frames per second while preserving view-dependent visual quality.
Generating photorealistic, view-dependent 3D scenes from images using neural radiance fields has gained significant interest, but standard neural radiance approaches suffer from extremely slow rendering speeds that prevent real-time interactive use. The article evaluates an alternative representation called PlenOctrees (an octree-based radiance field representation) and investigates how spherical basis function choices, optimization, and compression techniques enable real-time rendering while maintaining visual fidelity.
The authors conducted experimental evaluations across standard synthetic and real-world benchmark datasets (such as NeRF-synthetic and Tanks&Temples). They analyzed rendering speeds, reconstruction quality, and memory storage using different spherical basis formulations, including spherical harmonics and learnable spherical Gaussians. The workflow involves training a modified neural radiance field model with spherical harmonics, converting it directly into an octree representation, and fine-tuning the structure using derived analytic rendering derivatives.
The findings demonstrate substantial performance gains. Converted and fine-tuned PlenOctrees achieve real-time rendering speeds averaging 167.7 frames per second on synthetic data and 42.2 frames per second on real-world datasets, compared to standard neural radiance baselines that render at less than 0.1 frames per second. Visual reconstruction quality remains on par with or slightly exceeds existing baselines after fine-tuning. Additionally, ablations show that lower-order spherical harmonics provide competitive visual quality while drastically reducing storage footprints and increasing framerates up to 261.7 frames per second. Applying standard quantization and data compression reduces overall file sizes by 20 to 30 times, making web-based transmission feasible.
These results indicate that volumetric neural rendering can transition from offline computational pipelines into interactive, consumer-facing applications, such as web and virtual reality environments, without compromising visual fidelity. The initial neural network training remains computationally intensive, requiring roughly 50 hours per scene on a single high-end processor, but the subsequent octree conversion and fine-tuning requires only about 10 minutes. Practitioners seeking to deploy real-time 3D viewers should adopt lower-order spherical harmonics with compression pipelines to balance visual accuracy against memory constraints, while future efforts should focus on reducing the upfront training duration.
- Paper: NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis, Ben Mildenhall et al. (2020). It introduces the fundamental Neural Radiance Fields (NeRF) representation and volumetric rendering formulation that PlenOctrees explicitly accelerates and tabulates into an octree.
- Paper: Neural Sparse Voxel Fields, Lingjie Liu et al. (2020). It pioneers neural sparse voxel fields for bounding implicit scene evaluations within voxels to speed up rendering, providing direct context for voxel- and octree-based radiance field acceleration.
- Paper: OctNet: Learning Deep 3D Representations at High Resolutions, Gernot Riegler et al. (2016). It introduces deep 3D representations utilizing hierarchical octree data structures to overcome volumetric cubic complexity, establishing the data structuring principles used in PlenOctrees.
- Paper: Scene Representation Networks: Continuous 3D-Structure-Aware Neural Scene Representations, V. Sitzmann et al. (2019). It lays foundational ground for coordinate-based neural representations and differentiable ray-marching for continuous 3D scene rendering from 2D images.
- Paper: Light field rendering, Marc Levoy et al. (1996). It establishes classic light field sampling and view-dependent radiance interpolation without heavy geometry, inspiring explicit plenoptic tabulation techniques.
- Paper: Plenoxels: Radiance Fields without Neural Networks, Alex Yu et al. (2022). It extends PlenOctrees' spherical-harmonic voxel concept by completely discarding neural networks and optimizing explicit sparse voxel grids directly from scratch.
- Paper: Direct Voxel Grid Optimization: Super-fast Convergence for Radiance Fields Reconstruction, Cheng Sun et al. (2021). It builds directly upon real-time explicit grid optimization by developing direct voxel grid optimization with post-activated interpolation for rapid convergence.
- Paper: TensoRF: Tensorial Radiance Fields, Anpei Chen et al. (2022). It generalizes explicit radiance field representations like PlenOctrees by factorizing 3D voxel grids into low-rank tensor components for compact memory and fast rendering.
- Paper: Instant neural graphics primitives with a multiresolution hash encoding, Thomas Müller et al. (2022). It advances real-time neural radiance field rendering and training by replacing explicit tree structures with a multiresolution hash table encoding.
- Paper: 3D Gaussian Splatting for Real-Time Radiance Field Rendering, Bernhard Kerbl et al. (2023). It represents a major evolution in real-time view-dependent rendering, moving from octree-based voxel grids to explicit 3D Gaussian primitives with spherical harmonics.
- Paper: Ref-NeRF: Structured View-Dependent Appearance for Neural Radiance Fields, Dor Verbin et al. (2022). It refines the modeling of view-dependent appearances and specularities in radiance fields, addressing limitations inherent to standard spherical harmonic parameterizations.
- Paper: 4D Gaussian Splatting for Real-Time Dynamic Scene Rendering, Guanjun Wu et al. (2023). It builds upon explicit real-time radiance field rendering techniques to extend high-framerate novel view synthesis to dynamic 4D scenes.
