Deeper Depth Prediction with Fully Convolutional Residual Networks

Iro LainaChristian RupprechtVasileios BelagiannisFederico TombariNassir Navab

article20163DV2,046 citations

Proposes a fully convolutional residual network with novel up-sampling layers and a reverse Huber loss to achieve accurate, real-time monocular depth estimation without post-processing.

Listen

Estimating 3D depth from a single standard 2D image is an essential capability for computer vision applications such as autonomous navigation, 3D reconstruction, and scene understanding, especially when dedicated depth sensors are impractical. However, existing deep learning methods for monocular depth estimation rely on complex multi-stage pipelines, heavy parameter counts, or post-processing refinement steps like Conditional Random Fields, which significantly inflate computational and data requirements.

The article evaluates a unified, end-to-end fully convolutional neural network architecture designed to predict high-resolution depth maps directly from single color images. The objective is to demonstrate that a streamlined residual learning architecture can achieve state-of-the-art accuracy with fewer parameters and significantly less training data, operating fast enough for real-time applications.

The approach introduces a fully convolutional architecture based on a 50-layer residual network (ResNet-50) paired with novel residual up-sampling modules called up-projection blocks. The authors also reformulated the up-convolution operation to avoid computing zero values, reducing training time by roughly 15%. To train the network effectively on depth data distributions, the authors adopted a reverse Huber (berHu) loss function. The framework was evaluated on standard indoor (NYU Depth v2) and outdoor (Make3D) benchmark datasets and tested within an experimental 3D Simultaneous Localization and Mapping (SLAM) system.

The findings confirm that the proposed ResNet-UpProjection model sets a new state-of-the-art across standard benchmarks. On indoor scenes, it reduced relative error to 0.127 and root mean squared error to 0.573, outperforming prior multi-scale CNN models. On outdoor scenes, it achieved a relative error of 0.176 compared to previous results above 0.278. The network uses roughly 63.6 million parameters—about 3.5 times fewer than leading multi-scale architectures—and achieves top performance using approximately 12,000 unique training images, an order of magnitude fewer than competing deep models. Furthermore, single-image inference takes only 55 milliseconds (and 14 milliseconds per image in batches), demonstrating real-time capability and generating coherent 3D reconstructions in SLAM testing.

These results demonstrate that single-view depth estimation does not require cumbersome multi-network pipelines or costly post-processing to produce sharp, geometrically consistent depth boundaries. By lowering memory footprints, training overhead, and inference latency, this approach makes high-quality monocular depth estimation feasible for resource-constrained embedded systems and real-time robotic platforms.

Organizations developing vision-based navigation or augmented reality systems should adopt residual up-projection architectures and specialized loss functions like berHu over traditional fully-connected or multi-scale alternatives. For deployment in spatial mapping, further work is recommended to integrate temporal consistency mechanisms across consecutive video frames to approach the metric fidelity of multi-view geometry systems.

Confidence in these findings is high for standard indoor and outdoor operational domains covered by the benchmark datasets. However, stakeholders should note that the model relies on pre-trained classification weights and was evaluated within specific distance bounds (up to 70 meters outdoors), meaning extreme environmental conditions or long-range outdoor sensing may require additional domain-specific validation.

  • Paper: Deep Ordinal Regression Network for Monocular Depth Estimation, Huan Fu et al. (2018). This later method moves beyond continuous depth regression by recasting monocular prediction as ordinal regression, offering a useful next step from the source’s direct depth-map prediction.
  • Paper: Digging Into Self-Supervised Monocular Depth Estimation, Clément Godard et al. (2018). This work advances monocular depth estimation with self-supervised training and targeted fixes for occlusions and moving scenes, extending the source’s focus on practical prediction.
  • Paper: Depth Anything: Unleashing the Power of Large-Scale Unlabeled Data, Lihe Yang et al. (2024). This foundation-model approach scales monocular depth estimation with massive unlabeled data, extending the source’s benchmark-focused model toward broader generalization.
  • Paper: Depth Anything V2, Lihe Yang et al. (2024). Depth Anything V2 continues the push for efficient monocular depth prediction, pairing broad real-world coverage with sharper detail than the source’s benchmark-era system.
Cover for Deeper Depth Prediction with Fully Convolutional Residual Networks

Abstract

This paper addresses the problem of estimating the depth map of a scene given a single RGB image. We propose a fully convolutional architecture, encompassing residual learning, to model the ambiguous mapping between monocular images and depth maps. In order to improve the output resolution, we present a novel way to efficiently learn feature map up-sampling within the network. For optimization, we introduce the reverse Huber loss that is particularly suited for the task at hand and driven by the value distributions commonly present in depth maps. Our model is composed of a single architecture that is trained end-to-end and does not rely on post-processing techniques, such as CRFs or other additional refinement steps. As a result, it runs in real-time on images or videos. In the evaluation, we show that the proposed model contains fewer parameters and requires fewer training data than the current state of the art, while outperforming all approaches on depth estimation. Code and models are publicly available.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Methodology
  • 3.1 CNN Architecture
  • 3.2 Loss Function
  • 4 Experimental Results
  • 4.1 Experimental Setup
  • 4.2 NYU Depth Dataset
  • 4.3 Make3D Dataset
  • 4.4 Application to SLAM
  • 5 Conclusion
  • References

Knowls

  1. Knowl 1 — Up-Projection Block for Feature Map Upsampling

    model/method

    The up-projection block is a residual up-sampling module designed to increase the spatial resolution of feature representations within deep convolutional networks while preserving gradient propagation across scales.

    Given an input feature map X∈RH×W×Cin\mathbf{X} \in \mathbb{R}^{H \times W \times C_{\text{in}}}, the spatial resolution is first doubled via an unpooling operation Unpool(⋅)\text{Unpool}(\cdot), which places each input element into the top-left position of a 2×22 \times 2 zero-filled cell, yielding an intermediate map of size 2H×2W×Cin2H \times 2W \times C_{\text{in}}. The block then processes this unpooled map through two parallel paths:

    1. Main Branch: A 5×55 \times 5 convolution reducing or altering channels, followed by a Rectified Linear Unit (ReLU) activation, and a subsequent 3×33 \times 3 convolution: Fmain(X)=Wmain(2)∗ReLU(Wmain(1)∗Unpool(X))\mathbf{F}_{\text{main}}(\mathbf{X}) = \mathcal{W}_{\text{main}}^{(2)} * \text{ReLU}\left(\mathcal{W}_{\text{main}}^{(1)} * \text{Unpool}(\mathbf{X})\right)

    2. Projection (Residual) Branch: A separate 5×55 \times 5 convolution applied directly to the shared unpooled feature map to match the channel dimensions of the main branch output: Fproj(X)=Wproj∗Unpool(X)\mathbf{F}_{\text{proj}}(\mathbf{X}) = \mathcal{W}_{\text{proj}} * \text{Unpool}(\mathbf{X})

    The outputs of both branches are combined via elementwise addition and passed through a final ReLU activation: Y=ReLU(Fmain(X)+Fproj(X))\mathbf{Y} = \text{ReLU}\left(\mathbf{F}_{\text{main}}(\mathbf{X}) + \mathbf{F}_{\text{proj}}(\mathbf{X})\right) where Y∈R2H×2W×Cout\mathbf{Y} \in \mathbb{R}^{2H \times 2W \times C_{\text{out}}}. Stacking multiple up-projection blocks allows the network to progressively expand low-resolution deep feature representations to high-resolution spatial predictions without requiring fully connected layers.

  2. Knowl 2 — Fast Up-Convolution via 2D Filter Decomposition and Interleaving

    algorithm

    Standard up-convolution unpools an input feature map by a factor of 2 (leaving 75% of spatial locations as zeros) and applies a 5×55 \times 5 convolution filter W\mathbf{W}. To eliminate redundant multiplications with zero entries and avoid explicit unpooling in memory, the 5×55 \times 5 convolution is reformulated into four smaller, dense convolutions followed by a coordinate-based interleaving step.

    The original 5×55 \times 5 filter kernel is decomposed into four non-overlapping sub-filters based on the coordinate parity of non-zero entries in the unpooled grid:

    • Sub-filter WA\mathbf{W}_A of spatial size 3×33 \times 3, capturing even-row, even-column offsets.
    • Sub-filter WB\mathbf{W}_B of spatial size 3×23 \times 2, capturing even-row, odd-column offsets.
    • Sub-filter WC\mathbf{W}_C of spatial size 2×32 \times 3, capturing odd-row, even-column offsets.
    • Sub-filter WD\mathbf{W}_D of spatial size 2×22 \times 2, capturing odd-row, odd-column offsets.

    Each sub-filter is convolved directly with the dense input feature map X∈RH×W×C\mathbf{X} \in \mathbb{R}^{H \times W \times C}, producing four compact feature maps YA,YB,YC,YD∈RH×W×Cout\mathbf{Y}_A, \mathbf{Y}_B, \mathbf{Y}_C, \mathbf{Y}_D \in \mathbb{R}^{H \times W \times C_{\text{out}}}. These four maps are interleaved into an output tensor Y∈R2H×2W×Cout\mathbf{Y} \in \mathbb{R}^{2H \times 2W \times C_{\text{out}}}.

    Input: Feature map X∈RH×W×Cin\mathbf{X} \in \mathbb{R}^{H \times W \times C_{\text{in}}}, sub-filters WA∈R3×3\mathbf{W}_A \in \mathbb{R}^{3 \times 3}, WB∈R3×2\mathbf{W}_B \in \mathbb{R}^{3 \times 2}, WC∈R2×3\mathbf{W}_C \in \mathbb{R}^{2 \times 3}, WD∈R2×2\mathbf{W}_D \in \mathbb{R}^{2 \times 2}
    Output: Up-sampled feature map Y∈R2H×2W×Cout\mathbf{Y} \in \mathbb{R}^{2H \times 2W \times C_{\text{out}}}
    YA←X∗WA\mathbf{Y}_A \leftarrow \mathbf{X} * \mathbf{W}_A
    YB←X∗WB\mathbf{Y}_B \leftarrow \mathbf{X} * \mathbf{W}_B
    YC←X∗WC\mathbf{Y}_C \leftarrow \mathbf{X} * \mathbf{W}_C
    YD←X∗WD\mathbf{Y}_D \leftarrow \mathbf{X} * \mathbf{W}_D
    for r=0r = 0 to H−1H - 1 do
        for c=0c = 0 to W−1W - 1 do
            Y(2r,2c)←YA(r,c)\mathbf{Y}(2r, 2c) \leftarrow \mathbf{Y}_A(r, c)
            Y(2r,2c+1)←YB(r,c)\mathbf{Y}(2r, 2c + 1) \leftarrow \mathbf{Y}_B(r, c)
            Y(2r+1,2c)←YC(r,c)\mathbf{Y}(2r + 1, 2c) \leftarrow \mathbf{Y}_C(r, c)
            Y(2r+1,2c+1)←YD(r,c)\mathbf{Y}(2r + 1, 2c + 1) \leftarrow \mathbf{Y}_D(r, c)
        end for
    end for
    return Y\mathbf{Y}

    This exact reformulation achieves an identical mathematical result to unpooling followed by a 5×55 \times 5 convolution while reducing arithmetic operations by a theoretical factor of 4 and yielding training time speedups of approximately 15% across the full network.

  3. Knowl 3 — Reverse Huber Loss with Batch-Adaptive Threshold for Depth Estimation

    equation

    The reverse Huber (berHu) loss function B(x)\mathcal{B}(x) is a regression loss that applies an L1L_1 penalty for small residuals and an L2L_2 penalty for large residuals:

    B(x)={∣x∣if ∣x∣≤c,x2+c22cif ∣x∣>c,\mathcal{B}(x) = \begin{cases} |x| & \text{if } |x| \le c, \\[6pt] \dfrac{x^2 + c^2}{2c} & \text{if } |x| > c, \end{cases}

    where x=y~i−yi∈Rx = \tilde{y}_i - y_i \in \mathbb{R} is the residual between predicted depth y~i\tilde{y}_i and ground-truth depth yiy_i at pixel index ii, and c>0c > 0 is a scalar cutoff parameter. The function is continuous and first-order differentiable at ∣x∣=c|x| = c.

    In every gradient descent iteration, the cutoff cc is dynamically set to 20%20\% of the maximal absolute residual across all pixels in the current mini-batch:

    c=15max⁡i∣y~i−yi∣,c = \frac{1}{5} \max_{i} |\tilde{y}_i - y_i|,

    where ii ranges over all valid pixels across all images in the mini-batch.

    For residuals ∣x∣≤c|x| \le c, the derivative magnitude is constant (∣B′(x)∣=1|\mathcal{B}'(x)| = 1), preventing gradients from vanishing on small errors and encouraging fine structural boundaries. For residuals ∣x∣>c|x| > c, the quadratic term heavily penalizes large errors, matching the heavy-tailed error distributions typical of monocular depth estimation without ignoring high-residual samples.

  4. Knowl 4 — Fully Convolutional Residual Network for Monocular Depth Estimation

    model/method

    The Fully Convolutional Residual Network (FCRN) predicts a dense depth map from a single monocular RGB image in a single-scale, end-to-end framework without post-processing (such as Conditional Random Fields).

    1. Contractive Backbone: A ResNet-50 network initialized with ImageNet classification weights, modified by removing the final global pooling and fully connected layers. Given an RGB input image cropped to 304×228×3304 \times 228 \times 3, the backbone outputs 20482048 feature maps at a spatial resolution of 10×810 \times 8.
    2. Expansion Head: A sequence of four cascaded up-projection blocks. Each block performs 2×2\times spatial upsampling and reduces the feature channel depth, producing successive spatial dimensions of 20×16×51220 \times 16 \times 512, 40×32×25640 \times 32 \times 256, 80×64×12880 \times 64 \times 128, and 160×128×64160 \times 128 \times 64, totaling a 16×16\times spatial upsampling from the bottleneck.
    3. Prediction Layer: A dropout layer followed by a final convolutional layer that maps the 160×128×64160 \times 128 \times 64 representation to a single-channel depth prediction map of size 160×128160 \times 128.

    The total architecture contains 63.6×10663.6 \times 10^6 parameters. In contrast, replacing the up-projection stages with a direct fully connected layer to produce a 160×128160 \times 128 output would introduce approximately 3.3×1093.3 \times 10^9 parameters (12.6 GB memory).

  5. Knowl 5 — Component Ablation on NYU Depth v2

    data/table

    The performance of different down-sampling backbones (AlexNet, VGG-16, ResNet-50), up-sampling strategies (fully connected layers FC, deconvolution DeConv, simple up-convolution UpConv, and residual up-projection UpProj), and loss functions (L2L_2 vs. berHu) was evaluated on the 654 test images of the NYU Depth v2 dataset.

    Evaluation metrics include:

    • Relative error: rel=1∣T∣∑y∈T∣y~−y∣y\text{rel} = \frac{1}{|T|} \sum_{y \in T} \frac{|\tilde{y} - y|}{y}
    • Root mean squared error: rms=1∣T∣∑y∈T(y~−y)2\text{rms} = \sqrt{\frac{1}{|T|} \sum_{y \in T} (\tilde{y} - y)^2}
    • log⁡10\log_{10} error: log⁡10=1∣T∣∑y∈T∣log⁡10y~−log⁡10y∣\log_{10} = \frac{1}{|T|} \sum_{y \in T} |\log_{10} \tilde{y} - \log_{10} y|
    • Threshold accuracy: δi=% of pixels where max⁡(y~y,yy~)<1.25i\delta_i = \% \text{ of pixels where } \max\left(\frac{\tilde{y}}{y}, \frac{y}{\tilde{y}}\right) < 1.25^i for i∈{1,2,3}i \in \{1, 2, 3\}
    Architecture Type Loss #params rel rms log⁡10\log_{10} δ1\delta_1 δ2\delta_2 δ3\delta_3
    AlexNet FC L2L_2 104.4×106104.4 \times 10^6 0.209 0.845 0.090 0.586 0.869 0.967
    AlexNet FC berHu 104.4×106104.4 \times 10^6 0.207 0.842 0.091 0.581 0.872 0.969
    AlexNet UpConv L2L_2 6.3×1066.3 \times 10^6 0.218 0.853 0.094 0.576 0.855 0.957
    AlexNet UpConv berHu 6.3×1066.3 \times 10^6 0.215 0.855 0.094 0.574 0.855 0.958
    VGG UpConv L2L_2 18.5×10618.5 \times 10^6 0.194 0.746 0.083 0.626 0.894 0.974
    VGG UpConv berHu 18.5×10618.5 \times 10^6 0.194 0.790 0.083 0.629 0.889 0.971
    ResNet FC-160x128 berHu 359.1×106359.1 \times 10^6 0.181 0.784 0.080 0.649 0.894 0.971
    ResNet FC-64x48 berHu 73.9×10673.9 \times 10^6 0.154 0.679 0.066 0.754 0.938 0.984
    ResNet DeConv L2L_2 28.5×10628.5 \times 10^6 0.152 0.621 0.065 0.749 0.934 0.985
    ResNet UpConv L2L_2 43.1×10643.1 \times 10^6 0.139 0.606 0.061 0.778 0.944 0.985
    ResNet UpConv berHu 43.1×10643.1 \times 10^6 0.132 0.604 0.058 0.789 0.946 0.986
    ResNet UpProj L2L_2 63.6×10663.6 \times 10^6 0.138 0.592 0.060 0.785 0.952 0.987
    ResNet UpProj berHu 63.6×10663.6 \times 10^6 0.127 0.573 0.055 0.811 0.953 0.988

    ResNet-50 combined with UpProj and the berHu loss achieves the best overall performance, with a relative error of 0.127 and δ1\delta_1 accuracy of 0.811. In all configurations, the berHu loss outperforms the L2L_2 loss, showing the largest gains on relative error and δ1\delta_1.

  6. Knowl 6 — State-of-the-Art Monocular Depth Prediction on NYU Depth v2

    data/table

    The proposed ResNet-UpProj model was compared against prior single-image depth estimation methods on the standard 654-image test split of the NYU Depth v2 dataset. Predictions are bilinearly upsampled to 640×480640 \times 480 for evaluation against in-painted ground truth.

    Method rel rms rms(log) log⁡10\log_{10} δ1\delta_1 δ2\delta_2 δ3\delta_3
    Karsch et al. (2012) 0.374 1.12 – 0.134 – – –
    Ladicky et al. (2014) – – – – 0.542 0.829 0.941
    Liu et al. (2014) 0.335 1.06 – 0.127 – – –
    Li et al. (2015) 0.232 0.821 – 0.094 0.621 0.886 0.968
    Liu et al. (2015) 0.230 0.824 – 0.095 0.614 0.883 0.971
    Wang et al. (2015) 0.220 0.745 0.262 0.094 0.605 0.890 0.970
    Eigen et al. (2014) 0.215 0.907 0.285 – 0.611 0.887 0.971
    Roy and Todorovic (2016) 0.187 0.744 – 0.078 – – –
    Eigen and Fergus (2015) 0.158 0.641 0.214 – 0.769 0.950 0.988
    ResNet-UpProj (Ours) 0.127 0.573 0.195 0.055 0.811 0.953 0.988

    The ResNet-UpProj model outperforms all previous approaches across all error and accuracy metrics without requiring multi-scale refinement cascades or CRF post-processing. Furthermore, the model uses approximately 12k unique raw training images (expanded to ~95k via augmentations), an order of magnitude fewer than the 120k unique images used by Eigen and Fergus (2015) and the 800k patches used by Li et al. (2015).

  7. Knowl 7 — Monocular Depth Benchmark on Make3D Dataset

    data/table

    The ResNet-UpProj network was evaluated on the Make3D outdoor benchmark (consisting of 400 training images and 134 test images). Predictions are upscaled to 345×460345 \times 460 via bilinear interpolation and evaluated on pixels with ground truth distance less than 70 meters (C1 criterion).

    Method rel rms log⁡10\log_{10}
    Karsch et al. (2012) 0.355 9.20 0.127
    Liu et al. (2014) 0.335 9.49 0.137
    Liu et al. (2015) 0.314 8.60 0.119
    Li et al. (2015) 0.278 7.19 0.092
    ResNet-UpProj (L2L_2) 0.223 4.89 0.089
    ResNet-UpProj (berHu) 0.176 4.46 0.072

    When optimized with the berHu loss, the model reduces relative error to 0.176 (compared to 0.278 for the prior best method by Li et al.) and root mean squared error to 4.46 meters (compared to 7.19 meters). Training with berHu outperforms L2L_2 optimization substantially (rel error 0.176 vs. 0.223).

  8. Knowl 8 — Direct 3D Dense SLAM from Single-Image Depth Predictions

    model/method

    Single-view depth predictions from the fully convolutional residual network (ResNet-UpProj) can be integrated directly into a dense RGB-D Simultaneous Localization and Mapping (SLAM) pipeline without requiring stereo pairs or temporal multi-view depth initialization.

    The SLAM system operates as follows:

    1. Frame-to-Frame Tracking: Camera motion between consecutive RGB frames is estimated via Gauss-Newton optimization on pixelwise intensity differences.
    2. Model Fusion: Depth maps predicted per frame by the ResNet-UpProj model are projected into 3D points and fused into a global surfel-based 3D map representation using point-based fusion.

    Because the depth estimator relies on learned monocular appearance cues rather than visual feature matching or temporal baseline parallax, it produces dense 3D reconstructions even across low-textured, planar surfaces (e.g., blank walls and floors) where traditional monocular feature-based SLAM and Structure-from-Motion algorithms fail.

  9. Knowl 9 — Inference Latency and Layer Speedups of Fast Up-Projection

    empirical result

    On an NVIDIA GeForce GTX TITAN GPU (12GB memory), execution times were measured for individual up-sampling blocks and complete inference pipelines:

    • Single Layer Latency: A standard unpooling plus 5×55 \times 5 convolution block requires 1.5 ms per image, whereas the fast decomposed up-projection block executes in 0.14 ms per image (a speedup of >10×>10\times, surpassing the theoretical 4×4\times arithmetic reduction due to better GPU memory linearization in cuDNN with smaller filter sizes).
    • Single Image Inference: Full forward-pass depth estimation for a single 304×228304 \times 228 RGB image takes 55 ms using up-projection blocks versus 78 ms using standard up-convolutions.
    • Batched Inference: At a batch size of 16, per-image latency drops to 14 ms (~71 frames per second) with up-projections versus 28 ms with up-convolutions, enabling real-time operation on video streams.
  10. Knowl 10 — Training Protocol and Data Augmentation for FCRN Depth Estimation

    experimental setup

    The FCRN depth estimation model was implemented in MatConvNet and trained on an NVIDIA GeForce GTX TITAN (12GB GPU) using mini-batch stochastic gradient descent with momentum 0.9 and a batch size of 16 images.

    • Layer Initialization: Weights in the ResNet-50 contractive backbone are initialized from ImageNet (ILSVRC) classification pre-training. Newly added up-sampling layers are initialized randomly from a zero-mean Gaussian distribution N(0,0.012)\mathcal{N}(0, 0.01^2).
    • Optimization Schedule: Training runs for ~20 epochs on NYU Depth v2 and ~30 epochs on Make3D. The initial learning rate is set to 10−210^{-2} for berHu optimization (and 0.0050.005 for L2L_2 on Make3D), reduced gradually every 6–8 epochs upon reaching loss plateaus.
    • Data Augmentation: Input RGB frames and ground-truth depth maps undergo online random transformations including small rotations, uniform scaling, color jitter, horizontal flips with probability 0.5, and random cropping down to the target network input resolution (304×228304 \times 228 for NYU Depth v2, half-resolution 172×230172 \times 230 for Make3D).
    • Loss Masking: On Make3D, pixels with ground-truth distance >70m> 70\text{m} (including sky regions) are masked out during loss calculation.

Coverage note — None was omitted; all key contributions—including the up-projection block, fast 2D filter decomposition, reverse Huber loss formulation, full FCRN architecture, ablation analysis, benchmark comparisons on NYU Depth v2 and Make3D, dense SLAM integration, and computational latency measurements—have been extracted.

References

  1. 1.R. Achanta, A. Shaji, K. Smith, A. Lucchi, P. Fua, and S. Susstrunk. SLIC superpixels compared to state-of-the-art superpixel methods. Pattern Analysis and Machine Intelligence, IEEE Transactions on, 34(11):2274–2282, 2012.
  2. 2.V. Belagiannis, C. Rupprecht, G. Carneiro, and N. Navab. Robust optimization for deep regression. Proceedings of the International Conference on Computer Vision, 2015.
  3. 3.N. Dalal and B. Triggs. Histograms of oriented gradients for human detection. In Proc. Conf. Computer Vision and Pattern Recognition (CVPR), pages 886–893, 2005.
  4. 4.A. Dosovitskiy, J. Tobias Springenberg, and T. Brox. Learning to generate chairs with convolutional neural networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 1538–1546, 2015.
  5. 5.D. Eigen and R. Fergus. Predicting depth, surface normals and semantic labels with a common multi-scale convolutional architecture. In Proc. Int. Conf. Computer Vision (ICCV), 2015.
  6. 6.D. Eigen, C. Puhrsch, and R. Fergus. Prediction from a single image using a multi-scale deep network. In Proc. Conf. Neural Information Processing Systems (NIPS), 2014.
  7. 7.K. He, X. Zhang, S. Ren, and J. Sun. Deep residual learning for image recognition. arXiv preprint arXiv:1512.03385, 2015.
  8. 8.D. Hoiem, A. Efros, M. Hebert, et al. Geometric context from a single image. In Computer Vision, 2005. ICCV 2005. Tenth IEEE International Conference on, volume 1, pages 654–661. IEEE, 2005.
  9. 9.S. Ioffe and C. Szegedy. Batch normalization: Accelerating deep network training by reducing internal covariate shift. In Proceedings of The 32nd International Conference on Machine Learning, pages 448–456, 2015.
  10. 10.K. Karsch, C. Liu, and S. B. Kang. Depth extraction from video using non-parametric sampling. In Proc. Europ. Conf. Computer Vision (ECCV), pages 775–788, 2012.
  11. 11.M. Keller, D. Lefloch, M. Lambers, S. Izadi, T. Weyrich, and A. Kolb. Real-time 3d reconstruction in dynamic scenes using point-based fusion. In Proc. Int. Conf. 3D Vision(3DV), pages 1–8, 2013.
  12. 12.C. Kerl, J. Sturm, and D. Cremers. Robust odometry estimation for RGB-D cameras. In Proc. Int. Conf. on Robotics and Automation (ICRA), 2013.
  13. 13.J. Konrad, M. Wang, and P. Ishwar. 2d-to-3d image conversion by learning depth from examples. In Proc. Conf. Computer Vision and Pattern Recognition Workshops, pages 16–22, 2012.
  14. 14.A. Krizhevsky, I. Sutskever, and G. E. Hinton. Imagenet classification with deep convolutional neural networks. In Advances in neural information processing systems, pages 1097–1105, 2012.
  15. 15.L. Ladicky, J. Shi, and M. Pollefeys. Pulling things out of perspective. In Proc. Conf. Computer Vision and Pattern Recognition (CVPR), pages 89–96, 2014.
  16. 16.B. Li, C. Shen, Y. Dai, A. V. den Hengel, and M. He. Depth and surface normal estimation from monocular images using regression on deep features and hierarchical CRFs. In Proc. Conf. Computer Vision and Pattern Recognition (CVPR), pages 1119–1127, 2015.
  17. 17.B. Liu, S. Gould, and D. Koller. Single image depth estimation from predicted semantic labels. In Computer Vision and Pattern Recognition (CVPR), 2010 IEEE Conference on, pages 1253–1260. IEEE, 2010.
  18. 18.C. Liu, J. Yuen, and A. Torralba. Sift flow: Dense correspondence across scenes and its applications. Pattern Analysis and Machine Intelligence, IEEE Transactions on, 33(5):978–994, 2011.
  19. 19.F. Liu, C. Shen, and G. Lin. Deep convolutional neural fields for depth estimation from a single image. In Proc. Conf. Computer Vision and Pattern Recognition (CVPR), pages 5162–5170, 2015.
  20. 20.M. Liu, M. Salzmann, and X. He. Discrete-continuous depth estimation from a single image. In Proc. Conf. Computer Vision and Pattern Recognition (CVPR), pages 716–723, 2014.
  21. 21.J. Long, E. Shelhamer, and T. Darrell. Fully convolutional networks for semantic segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 3431–3440, 2015.
  22. 22.R. Memisevic and C. Conrad. Stereopsis via deep learning. In NIPS Workshop on Deep Learning, volume 1, 2011.
  23. 23.P. K. Nathan Silberman, Derek Hoiem and R. Fergus. Indoor segmentation and support inference from RGBD images. In ECCV, 2012.
  24. 24.A. Oliva and A. Torralba. Modeling the shape of the scene: A holistic representation of the spatial envelope. Int. Journal of Computer Vision (IJCV), pages 145–175, 2014.
  25. 25.A. B. Owen. A robust hybrid of lasso and ridge regression. Contemporary Mathematics, 443:59–72, 2007.
  26. 26.X. Ren, L. Bo, and D. Fox. Rgb-(d) scene labeling: Features and algorithms. In Computer Vision and Pattern Recognition (CVPR), 2012 IEEE Conference on, pages 2759–2766. IEEE, 2012.
  27. 27.A. Roy and S. Todorovic. Monocular depth estimation using neural regression forest. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition CVPR (CVPR), 2016.
  28. 28.O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, A. C. Berg, and L. Fei-Fei. ImageNet Large Scale Visual Recognition Challenge. International Journal of Computer Vision (IJCV), 115(3):211–252, 2015.
  29. 29.A. Saxena, S. H. Chung, and A. Y. Ng. Learning depth from single monocular images. In Advances in Neural Information Processing Systems, pages 1161–1168, 2005.
  30. 30.A. Saxena, M. Sun, and A. Ng. Make3d: Learning 3d scene structure from a single still image. IEEE Trans. Pattern Analysis and Machine Intelligence (PAMI), 12(5):824–840, 2009.
  31. 31.K. Simonyan and A. Zisserman. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556, 2014.
  32. 32.F. H. Sinz, J. Q. Candela, G. H. Bakır, C. E. Rasmussen, and M. O. Franz. Learning depth from stereo. In Pattern Recognition, pages 245–252. Springer, 2004.
  33. 33.S. Suwajanakorn and C. Hernandez. Depth from focus with your mobile phone. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2015.
  34. 34.R. Szeliski. Structure from motion. In Computer Vision, Texts in Computer Science, pages 303–334. Springer London, 2011.
  35. 35.J. Taylor, J. Shotton, T. Sharp, and A. Fitzgibbon. The vitruvian manifold: Inferring dense correspondences for one-shot human pose estimation. In Computer Vision and Pattern Recognition (CVPR), 2012 IEEE Conference on, pages 103–110. IEEE, 2012.
  36. 36.A. Vedaldi and K. Lenc. Matconvnet – convolutional neural networks for matlab. 2015.
  37. 37.P. Wang, X. Shen, Z. Lin, S. Cohen, B. Price, and A. L. Yuille. Towards unified depth and semantic prediction from a single image. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 2800–2809, 2015.
  38. 38.M. D. Zeiler, G. W. Taylor, and R. Fergus. Adaptive deconvolutional networks for mid and high level feature learning. In Computer Vision (ICCV), 2011 IEEE International Conference on, pages 2018–2025. IEEE, 2011.
  39. 39.R. Zhang, P.-S. Tsai, J. E. Cryer, and M. Shah. Shape-from-shading: a survey. Pattern Analysis and Machine Intelligence, IEEE Transactions on, 21(8):690–706, 1999.
  40. 40.L. Zwald and S. Lambert-Lacroix. The berhu penalty and the grouped effect. arXiv preprint arXiv:1207.6868, 2012.

Citation

MLA
Laina, I., et al. “Deeper Depth Prediction with Fully Convolutional Residual Networks”. arXiv, 2016, http://arxiv.org/abs/1606.00373v2.
APA
Laina, I., Rupprecht, C., Belagiannis, V., Tombari, F., & Navab, N. (2016). Deeper Depth Prediction with Fully Convolutional Residual Networks. arXiv. http://arxiv.org/abs/1606.00373v2
Chicago
Laina, I., C. Rupprecht, V. Belagiannis, F. Tombari, and N. Navab. 2016. “Deeper Depth Prediction with Fully Convolutional Residual Networks”. arXiv. http://arxiv.org/abs/1606.00373v2.
Harvard
Laina, I. et al. (2016) “Deeper Depth Prediction with Fully Convolutional Residual Networks”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1606.00373v2.
Vancouver
1. Laina I, Rupprecht C, Belagiannis V, Tombari F, Navab N (2016) Deeper Depth Prediction with Fully Convolutional Residual Networks. arXiv

BibTeX

@article{laina2016deeper,
  title = {Deeper Depth Prediction with Fully Convolutional Residual Networks},
  author = {Laina, Iro and Rupprecht, Christian and Belagiannis, Vasileios and Tombari, Federico and Navab, Nassir},
  year = {2016},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1606.00373v2},
  eprint = {1606.00373}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF