keyword
multiplane images
Multiplane images are a layered three-dimensional scene representation used in computer vision and graphics, consisting of a sequence of planar layers positioned at distinct depth levels within a reference camera view frustum. Each layer contains color and transparency channels that capture the visual appearance of surfaces at that specific depth, as well as background content that may be occluded in the original viewpoint. By projecting and alpha-compositing these fronto-parallel planes from back to front into new camera poses, this representation enables efficient novel view synthesis, realistic parallax effects, and depth reasoning from single-view or multi-view inputs.
2 items

Behind the Scenes: Density Fields for Single View Reconstruction
Felix Wimbauer, Nan Yang, Christian Rupprecht, Daniel Cremers
Why you should read this
Proposes a self-supervised framework that predicts continuous 3D density fields from a single image by decoupling geometry from color, enabling accurate reconstruction of occluded scene structures and novel view synthesis in complex outdoor environments.
Inferring a meaningful geometric scene representation from a single image is a fundamental problem in computer vision. Approaches based on traditional depth map prediction can only reason about areas that are visible in the image. Currently, neural radiance fields (NeRFs) can capture true 3D including color, but are too complex to be generated from a single image. As an alternative, we propose to predict an implicit density field from a single image. It maps every location in the frustum of the image to volumetric density. By directly sampling color from the available views instead of storing color in the density field, our scene representation becomes significantly less complex compared to NeRFs, and a neural network can predict it in a single forward pass. The network is trained through self-supervision from only video data. Our formulation allows volume rendering to perform both depth prediction and novel view synthesis. Through experiments, we show that our method is able to predict meaningful geometry for regions that are occluded in the input image. Additionally, we demonstrate the potential of our approach on three datasets for depth prediction and novel-view synthesis.
Added
2026-09-26

Look Outside the Room: Synthesizing A Consistent Long-Term 3D Scene Video from A Single Image
Xuanchi Ren, Xiaolong Wang
Why you should read this
Develops an autoregressive Transformer framework that generates geometrically consistent long-term 3D scene videos from a single input image and a large camera trajectory by using a camera-aware bias to guide space-time attention.
Novel view synthesis from a single image has recently attracted a lot of attention, and it has been primarily advanced by 3D deep learning and rendering techniques. However, most work is still limited by synthesizing new views within relatively small camera motions. In this paper, we propose a novel approach to synthesize a consistent long-term video given a single scene image and a trajectory of large camera motions. Our approach utilizes an autoregressive Transformer to perform sequential modeling of multiple frames, which reasons the relations between multiple frames and the corresponding cameras to predict the next frame. To facilitate learning and ensure consistency among generated frames, we introduce a locality constraint based on the input cameras to guide self-attention among a large number of patches across space and time. Our method outperforms state-of-the-art view synthesis approaches by a large margin, especially when synthesizing long-term future in indoor 3D scenes. Project page at https://xrenaa.github.io/look-outside-room/.
Added
2026-09-26
