Layered depth images
Jonathan ShadeSteven GortlerLi-wei HeRichard Szeliski
Introduces the Layered Depth Image representation to render complex novel views at interactive frame rates by storing multiple depth pixels along each line of sight and using a back-to-front warp ordering that avoids explicit depth sorting.
Generating new synthetic views of complex three-dimensional scenes typically demands substantial computing power, with processing times scaling sharply as geometric detail and visual richness increase. While image-based rendering techniques aim to accelerate this process by reprojecting existing two-dimensional images rather than recalculating full three-dimensional models, standard approaches suffer from visible rendering gaps, poor handling of occlusions, and severe visual artifacts when viewing angles shift.
The article evaluates and demonstrates two novel image-based rendering techniques designed to achieve interactive, multi-frame-per-second performance on standard personal computers without requiring specialized graphics hardware. The authors introduce Sprites with Depth to render smoothly varying surfaces and Layered Depth Images to handle complex scenes exhibiting high depth variation and occlusion.
To establish these solutions, the authors designed algorithmic pipelines and evaluated them across diverse test cases, including synthetic geometric models, dense ray-traced natural environments, and photographic image sets. For Sprites with Depth, the method forward-maps surface displacement values before applying backward color mapping and planar perspective warps. For Layered Depth Images, multiple depth pixels are stored along single lines of sight from a single camera view and rendered using an adapted back-to-front ordering algorithm combined with efficient pixel splatting, cache-aligned data structures, and view-frustum clipping.
The experimental findings show that these representations deliver substantial performance and efficiency gains. Rendering Sprites with Depth achieved frame rates between 16 and 47 frames per second on a 300 MHz Pentium II processor while eliminating surface rendering gaps. For complex scenes, Layered Depth Images maintained a very low average depth complexity—such as 1.24 layers per pixel in a multi-view test—resulting in only a 24 percent increase in rendering cost relative to single-layer images while sustaining interactive speeds of 8 to 10 frames per second. On a highly complex scene containing over 1.1 million depth pixels, frustum clipping accelerated rendering speeds by a factor of 2 to 4, sustaining 4 to 10 frames per second. Furthermore, packing depth pixel data into 8-byte structures to fit CPU cache lines yielded an immediate 25 percent improvement in rendering throughput.
These results indicate that complex visual scenes can be rendered interactively on commodity hardware without relying on expensive depth-sorting buffers or specialized rendering pipelines. By scaling storage linearly with depth complexity rather than multiplying it by the number of input viewpoints, the approach lowers hardware costs, decreases memory overhead, and broadens the deployment of interactive graphics in resource-constrained environments.
Organizations developing interactive visualization systems should consider adopting Layered Depth Images for complex geometries and Sprites with Depth for smoothly curved surfaces as intermediate primitives within a tiered rendering framework. Future efforts should focus on creating automated pipelines that classify scene components into the most appropriate rendering primitive and developing methods that account for dynamic lighting and view-dependent effects such as specular highlights.
The primary limitations of this approach include potential image degradation caused by multiple resampling stages and reduced spatial resolution for surfaces viewed at glancing angles from the base camera perspective. The sampling and aliasing behaviors under extreme view angles remain not fully formalized, so implementers should exercise caution when viewpoints diverge significantly from the original reference camera.
- Paper: View Interpolation for Image Synthesis, Shenchang Eric Chen et al. (1993). This foundational paper introduces image-based rendering via range-data warping and visibility-ordered compositing, which Layered Depth Images generalizes to handle complex multi-surface occlusions.
- Paper: Compositing digital images, Thomas K. Porter et al. (1984). This classic paper establishes the mathematical rules for alpha compositing and back-to-front blending that enable LDI rendering to correctly combine warped layers without a z-buffer.
- Paper: Light field rendering, Marc Levoy et al. (1996). This work establishes 4D light field representations for novel view synthesis, providing key conceptual motivation for using compact, depth-augmented image representations to bypass dense spatial sampling.
- Paper: The lumigraph, Steven J. Gortler et al. (1996). This paper demonstrates how incorporating geometric depth into image-based rendering improves reconstruction and resampling coherence, directly preceding the layered depth concept.
- Paper: High-quality video view interpolation using a layered representation, C. Lawrence Zitnick et al. (2004). This paper extends layered depth representations to interactive dynamic video view interpolation by adding explicit boundary matting and stereo layer segmentation.
