Built independently by an author, for readers. Read the story and support ChapterPal

keyword

MPI scene representation

A multiplane image scene representation is a three-dimensional visual model that represents a scene as a sequence of planar layers arranged at discrete depth levels from a reference viewpoint. Each parallel plane in the stack contains color and transparency values, enabling the model to capture complex scene geometry, thin structures, occlusions, and semi-transparent surfaces. Novel camera perspectives are rendered by warping each plane to a target viewpoint using planar homographies and blending the layers from back to front using standard alpha compositing. This representation is widely used in computer vision for novel view synthesis and local light field generation, providing a computationally efficient and differentiable framework for reconstructing realistic three-dimensional scenes from captured imagery.

1 item

Local light field fusion

Local light field fusion

Ben Mildenhall, Pratul P. Srinivasan, Rodrigo Ortiz-Cayon, Nima Khademi Kalantari, Ravi Ramamoorthi, Ren Ng, Abhishek Kar

OrganizationsFyusion Inc.Texas A&M UniversityUniversity of California BerkeleyUniversity of California, San Diego

Why you should read this

Develops a practical view synthesis pipeline that blends local multiplane image representations and derives plenoptic sampling bounds, allowing users to reliably capture and render complex real-world scenes using up to 4000x fewer views.

We present a practical and robust deep learning solution for capturing and rendering novel views of complex real world scenes for virtual exploration. Previous approaches either require intractably dense view sampling or provide little to no guidance for how users should sample views of a scene to reliably render high-quality novel views. Instead, we propose an algorithm for view synthesis from an irregular grid of sampled views that first expands each sampled view into a local light field via a multiplane image (MPI) scene representation, then renders novel views by blending adjacent local light fields. We extend traditional plenoptic sampling theory to derive a bound that specifies precisely how densely users should sample views of a given scene when using our algorithm. In practice, we apply this bound to capture and render views of real world scenes that achieve the perceptual quality of Nyquist rate view sampling while using up to 4000x fewer views. We demonstrate our approach's practicality with an augmented reality smartphone app that guides users to capture input images of a scene and viewers that enable realtime virtual exploration on desktop and mobile platforms.

Added

2026-09-25