Built independently by an author, for readers. Read the story and support ChapterPal

keyword

depth map

A depth map is an image or image channel in computer vision and computer graphics where each pixel value represents the spatial distance between a viewpoint and the corresponding physical surface in a three-dimensional scene. Instead of capturing visual properties such as color or brightness, a depth map records metric distance or optical disparity relative to a camera or sensor coordinate system. These representations can be acquired directly using specialized hardware such as structured-light or time-of-flight sensors, or computed algorithmically through techniques like stereo vision matching, focus cues, coded apertures, and neural network depth estimation. Depth maps are fundamental to various applications, including three-dimensional reconstruction, view interpolation, image-based rendering, scene segmentation, and robotic spatial perception.

9 items

GLIGEN: Open-Set Grounded Text-to-Image Generation

GLIGEN: Open-Set Grounded Text-to-Image Generation

Yuheng Li, Haotian Liu, Qingyang Wu, Fangzhou Mu, Jianwei Yang, Jianfeng Gao, Chunyuan Li, Yong Jae Lee

OrganizationsColumbia UniversityMicrosoftUniversity of Wisconsin Madison

Why you should read this

Proposes GLIGEN, a framework that injects spatial grounding inputs into frozen pre-trained diffusion models via gated layers, enabling precise layout-controlled image generation across open-world concepts without retraining the base model.

Large-scale text-to-image diffusion models have made amazing advances. However, the status quo is to use text input alone, which can impede controllability. In this work, we propose GLIGEN, Grounded-Language-to-Image Generation, a novel approach that builds upon and extends the functionality of existing pre-trained text-to-image diffusion models by enabling them to also be conditioned on grounding inputs. To preserve the vast concept knowledge of the pre-trained model, we freeze all of its weights and inject the grounding information into new trainable layers via a gated mechanism. Our model achieves open-world grounded text2img generation with caption and bounding box condition inputs, and the grounding ability generalizes well to novel spatial configurations and concepts. GLIGEN's zero-shot performance on COCO and LVIS outperforms that of existing supervised layout-to-image baselines by a large margin.

Added

2026-10-05

Image and depth from a conventional camera with a coded aperture

Image and depth from a conventional camera with a coded aperture

Anat Levin, R. Fergus, F. Durand, W. Freeman

OrganizationsMassachusetts Institute of Technology

Why you should read this

Proposes inserting an optimized coded aperture mask into a standard camera lens to simultaneously extract depth maps and recover all-focus images from a single exposure without sacrificing spatial resolution.

A conventional camera captures blurred versions of scene information away from the plane of focus. Camera systems have been proposed that allow for recording all-focus images, or for extracting depth, but to record both simultaneously has required more extensive hardware and reduced spatial resolution. We propose a simple modification to a conventional camera that allows for the simultaneous recovery of both (a) high resolution image information and (b) depth information adequate for semi-automatic extraction of a layered depth representation of the image. Our modification is to insert a patterned occluder within the aperture of the camera lens, creating a coded aperture. We introduce a criterion for depth discriminability which we use to design the preferred aperture pattern. Using a statistical model of images, we can recover both depth information and an all-focus image from single photographs taken with the modified camera. A layered depth map is then extracted, requiring user-drawn strokes to clarify layer assignments in some cases. The resulting sharp image and layered depth map can be combined for various photographic applications, including automatic scene segmentation, post-exposure refocusing, or re-rendering of the scene from an alternate viewpoint.

Added

2026-09-24

Unsupervised CNN for Single View Depth Estimation: Geometry to the Rescue

Unsupervised CNN for Single View Depth Estimation: Geometry to the Rescue

Ravi Garg, Vijay Kumar BG, Gustavo Carneiro, Ian Reid

OrganizationsUniversity of Adelaide

Why you should read this

Proposes an unsupervised framework for single-view depth estimation that trains a convolutional neural network using image reconstruction loss across stereo pairs, matching the accuracy of fully supervised models on KITTI without requiring ground-truth depth data.

A significant weakness of most current deep Convolutional Neural Networks is the need to train them using vast amounts of manu- ally labelled data. In this work we propose a unsupervised framework to learn a deep convolutional neural network for single view depth predic- tion, without requiring a pre-training stage or annotated ground truth depths. We achieve this by training the network in a manner analogous to an autoencoder. At training time we consider a pair of images, source and target, with small, known camera motion between the two such as a stereo pair. We train the convolutional encoder for the task of predicting the depth map for the source image. To do so, we explicitly generate an inverse warp of the target image using the predicted depth and known inter-view displacement, to reconstruct the source image; the photomet- ric error in the reconstruction is the reconstruction loss for the encoder. The acquisition of this training data is considerably simpler than for equivalent systems, requiring no manual annotation, nor calibration of depth sensor to camera. We show that our network trained on less than half of the KITTI dataset (without any further augmentation) gives com- parable performance to that of the state of art supervised methods for single view depth estimation.

Added

2026-09-24

High-quality video view interpolation using a layered representation

High-quality video view interpolation using a layered representation

C. Lawrence Zitnick, S. Kang, M. Uyttendaele, Simon A. J. Winder, R. Szeliski

OrganizationsMicrosoft

Why you should read this

Presents a novel two-layer depth and matting representation with a segmentation-based stereo algorithm that enables real-time, interactive free-viewpoint video synthesis from a sparse set of synchronized cameras.

The ability to interactively control viewpoint while watching a video is an exciting application of image-based rendering. The goal of our work is to render dynamic scenes with interactive viewpoint control using a relatively small number of video cameras. In this paper, we show how high-quality video-based rendering of dynamic scenes can be accomplished using multiple synchronized video streams combined with novel image-based modeling and rendering algorithms. Once these video streams have been processed, we can synthesize any intermediate view between cameras at any time, with the potential for space-time manipulation. In our approach, we first use a novel color segmentation-based stereo algorithm to generate high-quality photoconsistent correspondences across all camera views. Mattes for areas near depth discontinuities are then automatically extracted to reduce artifacts during view synthesis. Finally, a novel temporal two-layer compressed representation that handles matting is developed for rendering at interactive rates.

Added

2026-09-24