Local light field fusion
Ben MildenhallPratul P. SrinivasanRodrigo Ortiz-CayonNima Khademi KalantariRavi RamamoorthiRen NgAbhishek Kar
Develops a practical view synthesis pipeline that blends local multiplane image representations and derives plenoptic sampling bounds, allowing users to reliably capture and render complex real-world scenes using up to 4000x fewer views.
Immersive virtual exploration of real-world environments requires synthesizing novel viewpoints from captured photographs. Standard light field sampling methods demand an impractically large number of images—often millions per square meter—to render scenes smoothly without visual distortion. While geometry-assisted image-based rendering techniques attempt to reconstruct views from sparser inputs, they typically rely on trial-and-error capture setups and struggle with complex geometries, occlusions, and reflective surfaces.
The article develops and validates a practical view synthesis framework that combines deep learning with sampling theory to establish concrete mathematical guidelines for capturing real-world scenes reliably with significantly fewer photographs.
The approach introduces a two-stage pipeline. First, a deep three-dimensional convolutional neural network expands each captured image into a layered scene representation known as a multiplane image, which models local light fields and transparencies across depth planes. Second, novel viewpoints are rendered continuously in real time by projecting and blending adjacent multiplane representations using accumulated opacity to resolve occlusions. To support real-world use, the authors also developed an augmented reality smartphone capture application and evaluated performance across synthetic benchmarks and over 60 real handheld captures.
The primary finding is that this framework reduces the required number of captured views by up to 4000 times compared to traditional light field sampling standards while matching full perceptual visual quality. The maximum allowable disparity between adjacent views was determined to be 64 pixels, establishing a practical rule that governs the relationship between capture spacing, scene depth, and camera field of view. Quantitative and perceptual metrics demonstrated that the method consistently outperformed existing global mesh reconstruction, heuristic volume blending, and per-view deep warping techniques, particularly when handling fine geometric structures and non-Lambertian reflections.
These findings indicate that high-fidelity, interactive virtual exploration can be achieved casually using commodity handheld devices without specialized camera rigs or multi-hour reconstruction pipelines. By providing exact capture rules, the method eliminates costly guesswork and trial-and-error capture failures. Furthermore, the light rendering compute requirements make real-time interaction feasible on both desktop and mobile platforms.
Organizations developing virtual reality, augmented reality, or interactive 3D media applications should adopt the prescriptive sampling formulas to configure automated capture interfaces. Practitioners should ensure capture workflows guide users so that pixel shifts between adjacent views remain within the 64-pixel threshold. Future engineering work should explore multiresolution neural architectures to better support ultra-high-resolution imagery and improve geometric disambiguation in scenes with highly repetitive textures or moving objects.
The findings are supported by theoretical proofs, extensive quantitative evaluations, and diverse real-world demonstrations. However, users should exercise caution in environments containing subject motion or highly repetitive, untextured patterns, where local depth ambiguities can still produce minor visual artifacts.
- Paper: Light field rendering, Marc Levoy et al. (1996). Introduces 4D light field rendering and sampling foundations that Local Light Field Fusion directly adapts and extends with multiplane image representations and sampling bounds.
- Paper: The lumigraph, Steven J. Gortler et al. (1996). Provides the foundational plenoptic representation and geometry-assisted light field blending principles that motivate irregular-grid local light field fusion.
- Paper: Layered depth images, Jonathan Shade et al. (1998). Pioneers layered depth representations for occlusion handling and novel view synthesis that serve as direct conceptual precursors to multiplane images.
- Paper: High-quality video view interpolation using a layered representation, C. Lawrence Zitnick et al. (2004). Establishes layered scene decompositions for view interpolation to eliminate boundary and disocclusion artifacts.
- Paper: MVSNet: Depth Inference for Unstructured Multi-view Stereo, Yao Yao et al. (2018). Demonstrates plane-sweep cost volumes for multi-view depth inference, underpinning modern deep learning pipelines for multiplane representations.
- Paper: NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis, Ben Mildenhall et al. (2020). Extends continuous view synthesis beyond local multiplane image blending by encoding 5D radiance fields into implicit coordinate neural networks (NeRF).
- Paper: Mip-NeRF: A Multiscale Representation for Anti-Aliasing Neural Radiance Fields, Jonathan T. Barron et al. (2021). Improves neural radiance fields with cone tracing to resolve scale and anti-aliasing issues inherent in discrete view sampling.
- Paper: Mip-NeRF 360: Unbounded Anti-Aliased Neural Radiance Fields, Jonathan T. Barron et al. (2022). Generalizes neural view synthesis from forward-facing captures to unbounded 360-degree scenes using scene contraction and proposal sampling.
- Paper: Ref-NeRF: Structured View-Dependent Appearance for Neural Radiance Fields, Dor Verbin et al. (2022). Explicitly addresses complex specular highlights and view-dependent appearance that challenge traditional multiplane and radiance field fusion.
- Paper: 3D Gaussian Splatting for Real-Time Radiance Field Rendering, Bernhard Kerbl et al. (2023). Replaces ray-marched implicit fields and blended MPIs with fast-rendering 3D Gaussian primitives to deliver real-time novel view synthesis.
- Paper: Plenoxels: Radiance Fields without Neural Networks, Alex Yu et al. (2022). Develops explicit sparse voxel grids with spherical harmonics, achieving fast radiance field reconstruction without heavy neural networks.
- Paper: PlenOctrees for Real-time Rendering of Neural Radiance Fields, Alex Yu et al. (2021). Converts continuous radiance fields into octree-based data structures to enable real-time view rendering on standard hardware.
- Paper: TensoRF: Tensorial Radiance Fields, Anpei Chen et al. (2022). Factorizes 4D scene radiance fields using low-rank tensor decompositions to accelerate reconstruction and reduce memory overhead.
