keyword
depth maps
A depth map is an image or data channel where each pixel value represents the distance between a viewpoint and the surfaces of objects in a scene. Unlike conventional color images that record light intensity and chromaticity, depth maps record geometric spatial information from a specific camera perspective. These maps can be captured directly using active range sensors, such as structured-light and time-of-flight devices, or estimated computationally from standard 2D photographs using techniques like stereo vision, structure from motion, and deep neural networks. Depth maps are widely used in computer vision, robotics, and 3D computer graphics to enable tasks such as 3D scene reconstruction, spatial reasoning, object localization, and visual rendering.
12 items

Depth-supervised NeRF: Fewer Views and Faster Training for Free
Kangle Deng, Andrew Liu, Jun-Yan Zhu, Deva Ramanan
Why you should read this
Presents DS-NeRF, a method that uses sparse 3D keypoints freely generated by structure-from-motion pipelines to supervise ray termination depth, achieving 2-3x faster training and higher-quality novel view synthesis with fewer input images.
A commonly observed failure mode of Neural Radiance Field (NeRF) is fitting incorrect geometries when given an insufficient number of input views. One potential reason is that standard volumetric rendering does not enforce the constraint that most of a scene’s geometry consist of empty space and opaque surfaces. We formalize the above assumption through DS-NeRF (Depth-supervised Neural Radiance Fields), a loss for learning radiance fields that takes advantage of readily-available depth supervision. We leverage the fact that current NeRF pipelines require images with known camera poses that are typically estimated by running structure-from-motion (SFM). Crucially, SFM also produces sparse 3D points that can be used as “free” depth supervision during training: we add a loss to encourage the distribution of a ray’s terminating depth matches a given 3D keypoint, incorporating depth uncertainty. DS-NeRF can render better images given fewer training views while training 2-3x faster. Further, we show that our loss is compatible with other recently proposed NeRF methods, demonstrating that depth is a cheap and easily digestible supervisory signal. And finally, we find that DS-NeRF can support other types of depth supervision such as scanned depth sensors and RGB-D reconstruction outputs.
Added
2026-10-04

Loc3R-VLM: Language-based Localization and 3D Reasoning with Vision-Language Models
Kevin Qu, Haozhe Qi, Mihai Dusmanu, Mahdi Rad, Rui Wang, Marc Pollefeys
Why you should read this
Proposes Loc3R-VLM, a framework that equips 2D vision-language models with 3D spatial reasoning and language-based localization capabilities from monocular video by coupling global layout reconstruction with viewpoint-aware situation modeling.
Multimodal Large Language Models (MLLMs) have made impressive progress in connecting vision and language, but they still struggle with spatial understanding and viewpoint-aware reasoning. Recent efforts aim to augment the input representations with geometric cues rather than explicitly teaching models to reason in 3D space. We introduce Loc3R-VLM, a framework that equips 2D Vision-Language Models with advanced 3D understanding capabilities from monocular video input. Inspired by human spatial cognition, Loc3R-VLM relies on two joint objectives: global layout reconstruction to build a holistic representation of the scene structure, and explicit situation modeling to anchor egocentric perspective. These objectives provide direct spatial supervision that grounds both perception and language in a 3D context. To ensure geometric consistency and metric-scale alignment, we leverage lightweight camera pose priors extracted from a pre-trained 3D foundation model. Loc3R-VLM achieves state-of-the-art performance in language-based localization and outperforms existing 2D- and video-based approaches on situated and general 3D question-answering benchmarks, demonstrating that our spatial supervision framework enables strong 3D understanding. Project page: this https URL
Added
2026-09-29

Learning 3D Representations from 2D Pre-Trained Models via Image-to-Point Masked Autoencoders
Renrui Zhang, Liuhui Wang, Yu Qiao, Peng Gao, Hongsheng Li
Why you should read this
Proposes I2P-MAE, a self-supervised framework that transfers rich visual knowledge from off-the-shelf 2D pre-trained models to 3D point cloud learning through semantically guided masking and multi-view feature reconstruction.
Pre-training by numerous image data has become de-facto for robust 2D representations. In contrast, due to the expensive data processing, a paucity of 3D datasets severely hinders the learning for high-quality 3D features. In this paper, we propose an alternative to obtain superior 3D representations from 2D pre-trained models via Image-to-Point Masked Autoencoders, named as I2P-MAE. By self-supervised pre-training, we leverage the well learned 2D knowledge to guide 3D masked autoencoding, which reconstructs the masked point tokens with an encoder-decoder architecture. Specifically, we first utilize off-the-shelf 2D models to extract the multi-view visual features of the input point cloud, and then conduct two types of image-to-point learning schemes. For one, we introduce a 2D-guided masking strategy that maintains semantically important point tokens to be visible. Compared to random masking, the network can better concentrate on significant 3D structures with key spatial cues. For another, we enforce these visible tokens to reconstruct multi-view 2D features after the decoder. This enables the network to effectively inherit high-level 2D semantics for discriminative 3D modeling. Aided by our image-to-point pre-training, the frozen I2P-MAE, without any fine-tuning, achieves 93.4% accuracy for linear SVM on ModelNet40, competitive to existing fully trained methods. By further fine-tuning on on ScanObjectNN's hardest split, I2P-MAE attains the state-of-the-art 90.11% accuracy, +3.68% to the second-best, demonstrating superior transferable capacity. Code is available at https://github.com/ZrrSkywalker/I2P-MAE.
Added
2026-09-26

Joint bilateral upsampling
Johannes Kopf, Michael F. Cohen, Dani Lischinski, Matt Uyttendaele
Why you should read this
Proposes a fast edge-preserving method that uses high-resolution guide images to accurately upsample low-resolution computational solutions in tone mapping, stereo depth, colorization, and segmentation tasks.
Image analysis and enhancement tasks such as tone mapping, colorization, stereo depth, and photomontage, often require computing a solution (e.g., for exposure, chromaticity, disparity, labels) over the pixel grid. Computational and memory costs often require that a smaller solution be run over a downsampled image. Although general purpose upsampling methods can be used to interpolate the low resolution solution to the full resolution, these methods generally assume a smoothness prior for the interpolation. We demonstrate that in cases, such as those above, the available high resolution input image may be leveraged as a prior in the context of a joint bilateral upsampling procedure to produce a better high resolution solution. We show results for each of the applications above and compare them to traditional upsampling methods.
Added
2026-09-25

Single Image Dehazing via Multi-scale Convolutional Neural Networks
Wenqi Ren, Sibo Liu, Hua Zhang, Jin-shan Pan, Xiaochun Cao, Ming-Hsuan Yang
Why you should read this
Proposes a data-driven, multi-scale convolutional neural network that directly learns transmission maps through coarse global estimation and fine local refinement, replacing hand-crafted priors to achieve faster and higher-quality single-image dehazing.
The performance of existing image dehazing methods is limited by hand-designed features, such as the dark channel, color disparity and maximum contrast, with complex fusion schemes. In this paper, we propose a multi-scale deep neural network for single-image dehazing by learning the mapping between hazy images and their corresponding transmission maps. The proposed algorithm consists of a coarse-scale net which predicts a holistic transmission map based on the entire image, and a fine-scale net which refines results locally. To train the multi-scale deep network, we synthesize a dataset comprised of hazy images and corresponding transmission maps based on the NYU Depth dataset. Extensive experiments demonstrate that the proposed algorithm performs favorably against the state-of-the-art methods on both synthetic and real-world images in terms of quality and speed.
Added
2026-09-24

Accurate, Dense, and Robust Multiview Stereopsis
Yasutaka Furukawa, Jean Ponce
Why you should read this
Presents a patch-based multiview stereo algorithm (PMVS) that reconstructs highly accurate, dense 3D surface meshes from calibrated images without requiring bounding volume initializations or artificial smoothing.
This paper proposes a novel algorithm for calibrated multi-view stereopsis that outputs a (quasi) dense set of rectangular patches covering the surfaces visible in the input images. This algorithm does not require any initialization in the form of a bounding volume, and it detects and discards automatically outliers and obstacles. It does not perform any smoothing across nearby features, yet is currently the top performer in terms of both coverage and accuracy for four of the six benchmark datasets presented in [20]. The keys to its performance are effective techniques for enforcing local photometric consistency and global visibility constraints. Stereopsis is implemented as a match, expand, and filter procedure, starting from a sparse set of matched keypoints, and repeatedly expanding these to nearby pixel correspondences before using visibility constraints to filter away false matches. A simple but effective method for turning the resulting patch model into a mesh appropriate for image-based modeling is also presented. The proposed approach is demonstrated on various datasets including objects with fine surface details, deep concavities, and thin structures, outdoor scenes observed from a restricted set of viewpoints, and “crowded” scenes where moving obstacles appear in different places in multiple images of a static structure of interest.
Source
https://vigir.missouri.edu/~gdesouza/Research/Conference_CDs/IEEE_CVPR_2007/data/papers/0278.pdfAdded
2026-09-24

Shape from Shading: A Survey
Ruo Zhang, Ping-Sing Tsai, J. Cryer, M. Shah
Why you should read this
Presents a comprehensive empirical benchmark of six leading shape-from-shading algorithms, providing standardized implementations and quantitative performance comparisons across depth accuracy, gradient error, and computational efficiency on synthetic and real imagery.
Since the first shape-from-shading (SFS) technique was developed by Horn in the early 1970s, many different approaches have emerged. In this paper, six well-known SFS algorithms are implemented and compared. The performance of the algorithms was analyzed on synthetic images using mean and standard deviation of depth (Z) error, mean of surface gradient (p, q) error and CPU timing. Each algorithm works well for certain images, but performs poorly for others. In general, minimization approaches are more robust, while the other approaches are faster. The implementation of these algorithms in C, and images used in this paper, are available by anonymous ftp under the pub/tech-paper/survey directory at eustis.cs.ucf.edu (132.170.108.42). These are also part of the electronic version of paper.
Added
2026-09-18

T2I-Adapter: Learning Adapters to Dig Out More Controllable Ability for Text-to-Image Diffusion Models
Chong Mou, Xintao Wang, Liangbin Xie, Yanze Wu, Jian Zhang, Zhongang Qi, Ying Shan
Why you should read this
Introduces lightweight, composable adapter modules that enable precise structural and color control in pretrained text-to-image diffusion models while keeping the base network frozen.
The incredible generative ability of large-scale text-to-image (T2I) models has demonstrated strong power of learning complex structures and meaningful semantics. However, relying solely on text prompts cannot fully take advantage of the knowledge learned by the model, especially when flexible and accurate controlling (e.g., color and structure) is needed. In this paper, we aim to ``dig out" the capabilities that T2I models have implicitly learned, and then explicitly use them to control the generation more granularly. Specifically, we propose to learn simple and lightweight T2I-Adapters to align internal knowledge in T2I models with external control signals, while freezing the original large T2I models. In this way, we can train various adapters according to different conditions, achieving rich control and editing effects in the color and structure of the generation results. Further, the proposed T2I-Adapters have attractive properties of practical value, such as composability and generalization ability. Extensive experiments demonstrate that our T2I-Adapter has promising generation quality and a wide range of applications.
Added
2026-09-18

Deeper Depth Prediction with Fully Convolutional Residual Networks
Iro Laina, Christian Rupprecht, Vasileios Belagiannis, Federico Tombari, Nassir Navab
Why you should read this
Proposes a fully convolutional residual network with novel up-sampling layers and a reverse Huber loss to achieve accurate, real-time monocular depth estimation without post-processing.
This paper addresses the problem of estimating the depth map of a scene given a single RGB image. We propose a fully convolutional architecture, encompassing residual learning, to model the ambiguous mapping between monocular images and depth maps. In order to improve the output resolution, we present a novel way to efficiently learn feature map up-sampling within the network. For optimization, we introduce the reverse Huber loss that is particularly suited for the task at hand and driven by the value distributions commonly present in depth maps. Our model is composed of a single architecture that is trained end-to-end and does not rely on post-processing techniques, such as CRFs or other additional refinement steps. As a result, it runs in real-time on images or videos. In the evaluation, we show that the proposed model contains fewer parameters and requires fewer training data than the current state of the art, while outperforming all approaches on depth estimation. Code and models are publicly available.
Added
2026-09-17

SUN RGB-D: A RGB-D scene understanding benchmark suite
Shuran Song, Samuel P. Lichtenberg, Jianxiong Xiao
Why you should read this
Introduces a large-scale RGB-D dataset of over 10,000 images paired with dense 2D and 3D bounding annotations to benchmark core scene understanding tasks, including 3D object detection, semantic segmentation, and room layout estimation.
Scene understanding is one of the most important and challenging tasks in computer vision. Although remarkable progress has been made in the past few decades, the performance of general-purpose scene understanding is still far from satisfactory. The recent arrival of affordable depth sensors in consumer markets enables us to obtain reliable depth maps at a very low cost, greatly simplifying some common challenges in computer vision and enabling breakthroughs for several tasks, such as body pose estimation [5], intrinsic image [1], and 3D modeling [2]. RGB-D sensors have also enabled rapid progress for scene understanding. However, although we can download color images from the Internet easily, it is hard to obtain RGB-D data online. Therefore the existing RGB-D recognition benchmarks, such as NYU Depth v2 [4], are an order-of-magnitude smaller than modern recognition datasets (e.g. PASCAL VOC) for color images. Although these small datasets successfully bootstrapped initial research and enabled early progress in the past few years, the size limit is now the common bottleneck in advancing research to the next level. It causes easy overfitting of the algorithm during evaluation, and it cannot support learning for data-hungry algorithms that achieve state-of-the-art performance in color-based recognition tasks. Furthermore, although the RGB-D images these datasets provide contain depth maps, the annotation and evaluation metrics are mostly in the 2D image domain but not directly in 3D, as shown in Figure 1. However, scene understanding will be much more useful in the real 3D space. Hence we advocate that the community should use the depth map to reasoning about scenes and evaluating algorithms in 3D. We introduce SUN RGB-D, a dataset containing 10,335 RGB-D images with dense annotations in both 2D and 3D, for both objects and rooms. The goal of our dataset construction is to obtain an image dataset captured by various RGB-D sensors ( Intel RealSense, Asus Xtion LIVE PRO, Microsoft Kinect versions 1 and 2 ) at a similar scale as the PASCAL VOC object detection benchmark. We capture 3,784 images using Kinect v2 and 1,159 images using Intel RealSense. We included the 1,449 images from the NYU Depth V2 captured by Kinect v1 and We also choosed 554 realistic scene images from the Berkeley B3DO Dataset captured by Kinect v1, manually selected 3,389 distinguished frames without significant motion blur from the SUN3D videos captured by Asus Xtion. In total, we obtain 10,335 RGB-D images. To improve the depth map quality, we take short videos and use our SIFT+ICP algorithm to do the refinement. For each image, we annotate the objects with both 2D polygons and 3D bounding boxes and the room layout with 3D polygons. For the 10,335 RGB-D images, we have 146,617 2D polygons and 64,595 3D bounding boxes (with accurate orientations for objects) annotated. Therefore, there are 14.2 objects in each image on average. In total, there are 47 scene categories and 800 object categories. To evaluate all major tasks that work towards total scene understanding, we focus on six important recognition tasks towards total scene understanding that integrates objects, room layout and scene class, includes scene categorization, semantic segmentation, object detection, object orientation, room layout estimation, as well as a final total scene understanding task that integrates everything. The final task for our scene understanding benchmark is to estimate the whole scene including objects and room layout in 3D [3]. This task is also referred to “Basic Level Scene Understanding” [6]. We propose this benchmark task as the final goal to integrate both object detection and room layout estimation to obtain a total scene understanding, recognizing and localizing all the objects and the room structure. We choose state-of-the-art algorithms to evaluate each task. For the tasks without existing algorithm or implementation, we adapt popular algorithms from other tasks. For each task, whenever possible, we try to evaluate algorithms using color, depth, as well as RGB-D images to study the relative importance of color and depth, and gauge to what extent the information from both is complementary. Various evaluation results show that we can apply standard techniques that invented design for color (e.g. hand craft features, deep learning features, detector, sift flow label transfer) to depth domain and it can achieve comparable performance for various tasks.In most of cases, when we combining these two source of information the performance get improved. By constructing a PASCAL-scale dataset using various sensors and defining a benchmark for all major scene understanding tasks with 3D evaluation metrics, we hope to lay the foundation for advancing RGB-D scene understanding in the coming years.
Added
2026-09-16

Modeling and rendering architecture from photographs: a hybrid geometry- and image-based approach
Paul E. Debevec, Camillo J. Taylor, Jitendra Malik
Why you should read this
Presents a hybrid modeling framework that reconstructs photorealistic 3D architectural scenes from a sparse set of photographs by combining interactive block-based photogrammetry, model-based stereo, and view-dependent texture mapping.
We present a new approach for modeling and rendering existing architectural scenes from a sparse set of still photographs. Our modeling approach, which combines both geometry-based and image-based techniques, has two components. The first component is a photogrammetric modeling method which facilitates the recovery of the basic geometry of the photographed scene. Our photogrammetric modeling approach is effective, convenient, and robust because it exploits the constraints that are characteristic of architectural scenes. The second component is a model-based stereo algorithm, which recovers how the real scene deviates from the basic model. By making use of the model, our stereo technique robustly recovers accurate depth from widely-spaced image pairs. Consequently, our approach can model large architectural environments with far fewer photographs than current image-based modeling approaches. For producing renderings, we present view-dependent texture mapping, a method of compositing multiple views of a scene that better simulates geometric detail on basic models. Our approach can be used to recover models for use in either geometry-based or image-based rendering systems. We present results that demonstrate our approach’s ability to create realistic renderings of architectural scenes from viewpoints far from the original photographs.
Added
2026-09-16

Playing for Data: Ground Truth from Computer Games
Stephan R. Richter, Vibhav Vineet, Stefan Roth, Vladlen Koltun
Why you should read this
Demonstrates how intercepting communication between commercial video games and graphics hardware enables automatic generation of pixel-accurate semantic annotations, cutting the real-world manual training data needed for semantic segmentation by two-thirds.
Recent progress in computer vision has been driven by high-capacity models trained on large datasets. Unfortunately, creating large datasets with pixel-level labels has been extremely costly due to the amount of human effort required. In this paper, we present an approach to rapidly creating pixel-accurate semantic label maps for images extracted from modern computer games. Although the source code and the internal operation of commercial games are inaccessible, we show that associations between image patches can be reconstructed from the communication between the game and the graphics hardware. This enables rapid propagation of semantic labels within and across images synthesized by the game, with no access to the source code or the content. We validate the presented approach by producing dense pixel-level semantic annotations for 25 thousand images synthesized by a photorealistic open-world computer game. Experiments on semantic segmentation datasets show that using the acquired data to supplement real-world images significantly increases accuracy and that the acquired data enables reducing the amount of hand-labeled real-world data: models trained with game data and just 1/3 of the CamVid training set outperform models trained on the complete CamVid training set.
Added
2026-09-16
