keyword
depth map
A depth map is an image or image channel in computer vision and computer graphics where each pixel value represents the spatial distance between a viewpoint and the corresponding physical surface in a three-dimensional scene. Instead of capturing visual properties such as color or brightness, a depth map records metric distance or optical disparity relative to a camera or sensor coordinate system. These representations can be acquired directly using specialized hardware such as structured-light or time-of-flight sensors, or computed algorithmically through techniques like stereo vision matching, focus cues, coded apertures, and neural network depth estimation. Depth maps are fundamental to various applications, including three-dimensional reconstruction, view interpolation, image-based rendering, scene segmentation, and robotic spatial perception.
9 items

GLIGEN: Open-Set Grounded Text-to-Image Generation
Yuheng Li, Haotian Liu, Qingyang Wu, Fangzhou Mu, Jianwei Yang, Jianfeng Gao, Chunyuan Li, Yong Jae Lee
Why you should read this
Proposes GLIGEN, a framework that injects spatial grounding inputs into frozen pre-trained diffusion models via gated layers, enabling precise layout-controlled image generation across open-world concepts without retraining the base model.
Large-scale text-to-image diffusion models have made amazing advances. However, the status quo is to use text input alone, which can impede controllability. In this work, we propose GLIGEN, Grounded-Language-to-Image Generation, a novel approach that builds upon and extends the functionality of existing pre-trained text-to-image diffusion models by enabling them to also be conditioned on grounding inputs. To preserve the vast concept knowledge of the pre-trained model, we freeze all of its weights and inject the grounding information into new trainable layers via a gated mechanism. Our model achieves open-world grounded text2img generation with caption and bounding box condition inputs, and the grounding ability generalizes well to novel spatial configurations and concepts. GLIGEN's zero-shot performance on COCO and LVIS outperforms that of existing supervised layout-to-image baselines by a large margin.
Added
2026-10-05

Shape from Focus
S. Nayar, Y. Nakagawa
Why you should read this
Proposes a high-precision shape recovery framework that pairs a Sum-Modified-Laplacian focus operator with Gaussian interpolation across varying focus levels to accurately reconstruct dense 3D depth maps of microscopic and rough-textured industrial surfaces.
The shape from focus method presented here uses different focus levels to obtain a sequence of object images. The sum-modified-Laplacian (SML) operator is developed to provide local measures of the quality of image focus. The operator is applied to the image sequence to determine a set of focus measures at each image point. A depth estimation algorithm interpolates a small number of focus measure values to obtain accurate depth estimates. A fully automated shape from focus system has been implemented using an optical microscope and tested on a variety of industrial samples. Experimental results are presented that demonstrate the accuracy and robustness of the proposed method. These results suggest shape from focus to be an effective approach for a variety of challenging visual inspection tasks.
Added
2026-09-25

Layered depth images
Jonathan Shade, Steven Gortler, Li-wei He, Richard Szeliski
Why you should read this
Introduces the Layered Depth Image representation to render complex novel views at interactive frame rates by storing multiple depth pixels along each line of sight and using a back-to-front warp ordering that avoids explicit depth sorting.
In this paper we present a set of efficient image based rendering methods capable of rendering multiple frames per second on a PC. The first method warps Sprites with Depth representing smooth surfaces without the gaps found in other techniques. A second method for more general scenes performs warping from an intermediate representation called a Layered Depth Image (LDI). An LDI is a view of the scene from a single input camera view, but with multiple pixels along each line of sight. The size of the representation grows only linearly with the observed depth complexity in the scene. Moreover, because the LDI data are represented in a single image coordinate system, McMillan’s warp ordering algorithm can be successfully adapted. As a result, pixels are drawn in the output image in back-to-front order. No z-buffer is required, so alpha-compositing can be done efficiently without depth sorting. This makes splatting an efficient solution to the resampling problem.
Added
2026-09-25

Image and depth from a conventional camera with a coded aperture
Anat Levin, R. Fergus, F. Durand, W. Freeman
Why you should read this
Proposes inserting an optimized coded aperture mask into a standard camera lens to simultaneously extract depth maps and recover all-focus images from a single exposure without sacrificing spatial resolution.
A conventional camera captures blurred versions of scene information away from the plane of focus. Camera systems have been proposed that allow for recording all-focus images, or for extracting depth, but to record both simultaneously has required more extensive hardware and reduced spatial resolution. We propose a simple modification to a conventional camera that allows for the simultaneous recovery of both (a) high resolution image information and (b) depth information adequate for semi-automatic extraction of a layered depth representation of the image. Our modification is to insert a patterned occluder within the aperture of the camera lens, creating a coded aperture. We introduce a criterion for depth discriminability which we use to design the preferred aperture pattern. Using a statistical model of images, we can recover both depth information and an all-focus image from single photographs taken with the modified camera. A layered depth map is then extracted, requiring user-drawn strokes to clarify layer assignments in some cases. The resulting sharp image and layered depth map can be combined for various photographic applications, including automatic scene segmentation, post-exposure refocusing, or re-rendering of the scene from an alternate viewpoint.
Added
2026-09-24

Stereo Matching Using Belief Propagation
Jian Sun, Harry Shum, N. Zheng
Why you should read this
Proposes a Bayesian belief propagation framework on coupled Markov random fields that simultaneously estimates disparity, detects depth discontinuities, and handles occlusions for accurate dense stereo matching.
In this paper, we formulate the stereo matching problem as a Markov network and solve it using Bayesian belief propagation. The stereo Markov network consists of three coupled Markov random fields that model the following: a smooth field for depth/disparity, a line process for depth discontinuity, and a binary process for occlusion. After eliminating the line process and the binary process by introducing two robust functions, we apply the belief propagation algorithm to obtain the maximum a posteriori (MAP) estimation in the Markov network. Other low-level visual cues (e.g., image segmentation) can also be easily incorporated in our stereo model to obtain better stereo results. Experiments demonstrate that our methods are comparable to the state-of-the-art stereo algorithms for many test cases.
Added
2026-09-24

Unsupervised CNN for Single View Depth Estimation: Geometry to the Rescue
Ravi Garg, Vijay Kumar BG, Gustavo Carneiro, Ian Reid
Why you should read this
Proposes an unsupervised framework for single-view depth estimation that trains a convolutional neural network using image reconstruction loss across stereo pairs, matching the accuracy of fully supervised models on KITTI without requiring ground-truth depth data.
A significant weakness of most current deep Convolutional Neural Networks is the need to train them using vast amounts of manu- ally labelled data. In this work we propose a unsupervised framework to learn a deep convolutional neural network for single view depth predic- tion, without requiring a pre-training stage or annotated ground truth depths. We achieve this by training the network in a manner analogous to an autoencoder. At training time we consider a pair of images, source and target, with small, known camera motion between the two such as a stereo pair. We train the convolutional encoder for the task of predicting the depth map for the source image. To do so, we explicitly generate an inverse warp of the target image using the predicted depth and known inter-view displacement, to reconstruct the source image; the photomet- ric error in the reconstruction is the reconstruction loss for the encoder. The acquisition of this training data is considerably simpler than for equivalent systems, requiring no manual annotation, nor calibration of depth sensor to camera. We show that our network trained on less than half of the KITTI dataset (without any further augmentation) gives com- parable performance to that of the state of art supervised methods for single view depth estimation.
Added
2026-09-24

High-quality video view interpolation using a layered representation
C. Lawrence Zitnick, S. Kang, M. Uyttendaele, Simon A. J. Winder, R. Szeliski
Why you should read this
Presents a novel two-layer depth and matting representation with a segmentation-based stereo algorithm that enables real-time, interactive free-viewpoint video synthesis from a sparse set of synchronized cameras.
The ability to interactively control viewpoint while watching a video is an exciting application of image-based rendering. The goal of our work is to render dynamic scenes with interactive viewpoint control using a relatively small number of video cameras. In this paper, we show how high-quality video-based rendering of dynamic scenes can be accomplished using multiple synchronized video streams combined with novel image-based modeling and rendering algorithms. Once these video streams have been processed, we can synthesize any intermediate view between cameras at any time, with the potential for space-time manipulation. In our approach, we first use a novel color segmentation-based stereo algorithm to generate high-quality photoconsistent correspondences across all camera views. Mattes for areas near depth discontinuities are then automatically extracted to reduce artifacts during view synthesis. Finally, a novel temporal two-layer compressed representation that handles matting is developed for rendering at interactive rates.
Added
2026-09-24

Single Image Haze Removal Using Dark Channel Prior
Kaiming He, Jian Sun, X. Tang
Why you should read this
Introduces the dark channel prior, a simple statistical discovery that enables accurate estimation of haze thickness to recover clear outdoor scenes and generate depth maps from a single degraded photograph.
In this paper, we propose a simple but effective image prior - dark channel prior to remove haze from a single input image. The dark channel prior is a kind of statistics of the haze-free outdoor images. It is based on a key observation - most local patches in haze-free outdoor images contain some pixels which have very low intensities in at least one color channel. Using this prior with the haze imaging model, we can directly estimate the thickness of the haze and recover a high quality haze-free image. Results on a variety of outdoor haze images demonstrate the power of the proposed prior. Moreover, a high quality depth map can also be obtained as a by-product of haze removal.
Added
2026-09-24

An Iterative Image Registration Technique with an Application to Stereo Vision
Bruce D. Lucas, Takeo Kanade
Why you should read this
Proposes a gradient-based iterative optimization framework for fast image registration across affine transformations, establishing the foundational Lucas-Kanade method used across computer vision for optical flow and stereo matching.
Image registration finds a variety of applications in computer vision. Unfortunately, traditional image registration techniques tend to be costly. We present a new image registration technique that makes use of the spatial intensity gradient of the images to find a good match using a type of Newton-Raphson iteration. Our technique is faster because it examines far fewer potential matches between the images than existing techniques. Furthermore, this registration technique can be generalized to handle rotation, scaling and shearing. We show show our technique can be adapted for use in a stereo vision system.
Added
2026-09-06
