Built independently by an author, for readers. Read the story and support ChapterPal

keyword

NYU Depth

NYU Depth is a standard computer vision benchmark and RGB-D dataset designed for 3D indoor scene understanding, monocular depth estimation, and semantic segmentation. Collected using Microsoft Kinect sensors, it contains synchronized color images paired with aligned, dense depth maps across diverse indoor environments, including residential and commercial spaces. In addition to raw video recordings, the dataset provides curated subsets featuring in-painted depth maps to fill sensor holes, accelerometer data, and detailed per-pixel object class annotations. Widely adopted across artificial intelligence research, NYU Depth serves as a primary standard for evaluating how effectively machine learning models can infer three-dimensional geometry, surface normals, and scene semantics from two-dimensional images.

4 items

Deeper Depth Prediction with Fully Convolutional Residual Networks

Deeper Depth Prediction with Fully Convolutional Residual Networks

Iro Laina, Christian Rupprecht, Vasileios Belagiannis, Federico Tombari, Nassir Navab

OrganizationsJohns Hopkins UniversityTechnical University of MunichUniversity of BolognaUniversity of Oxford

Why you should read this

Proposes a fully convolutional residual network with novel up-sampling layers and a reverse Huber loss to achieve accurate, real-time monocular depth estimation without post-processing.

This paper addresses the problem of estimating the depth map of a scene given a single RGB image. We propose a fully convolutional architecture, encompassing residual learning, to model the ambiguous mapping between monocular images and depth maps. In order to improve the output resolution, we present a novel way to efficiently learn feature map up-sampling within the network. For optimization, we introduce the reverse Huber loss that is particularly suited for the task at hand and driven by the value distributions commonly present in depth maps. Our model is composed of a single architecture that is trained end-to-end and does not rely on post-processing techniques, such as CRFs or other additional refinement steps. As a result, it runs in real-time on images or videos. In the evaluation, we show that the proposed model contains fewer parameters and requires fewer training data than the current state of the art, while outperforming all approaches on depth estimation. Code and models are publicly available.

Added

2026-09-17

Deep Ordinal Regression Network for Monocular Depth Estimation

Deep Ordinal Regression Network for Monocular Depth Estimation

Huan Fu, Mingming Gong, Chaohui Wang, Kayhan Batmanghelich, Dacheng Tao

OrganizationsCarnegie Mellon UniversityUniversité Paris-EstUniversity of PittsburghUniversity of Sydney

Why you should read this

Presents an ordinal regression framework for monocular depth estimation that combines spacing-increasing discretization with a multi-scale architecture to achieve faster training convergence and superior accuracy across major benchmarks.

Monocular depth estimation, which plays a crucial role in understanding 3D scene geometry, is an ill-posed problem. Recent methods have gained significant improvement by exploring image-level information and hierarchical features from deep convolutional neural networks (DCNNs). These methods model depth estimation as a regression problem and train the regression networks by minimizing mean squared error, which suffers from slow convergence and unsatisfactory local solutions. Besides, existing depth estimation networks employ repeated spatial pooling operations, resulting in undesirable low-resolution feature maps. To obtain high-resolution depth maps, skip-connections or multi-layer deconvolution networks are required, which complicates network training and consumes much more computations. To eliminate or at least largely reduce these problems, we introduce a spacing-increasing discretization (SID) strategy to discretize depth and recast depth network learning as an ordinal regression problem. By training the network using an ordinary regression loss, our method achieves much higher accuracy and \dd{faster convergence in synch}. Furthermore, we adopt a multi-scale network structure which avoids unnecessary spatial pooling and captures multi-scale information in parallel. The method described in this paper achieves state-of-the-art results on four challenging benchmarks, i.e., KITTI [17], ScanNet [9], Make3D [50], and NYU Depth v2 [42], and win the 1st prize in Robust Vision Challenge 2018. Code has been made available at: this https URL.

Added

2026-09-17