Unsupervised Learning of Depth and Ego-Motion from Video

Tinghui ZhouMatthew BrownNoah SnavelyDavid G. Lowe

article2017CVPR2,960 citations

Proposes an unsupervised framework that trains monocular depth and camera pose networks jointly using view synthesis supervision, achieving accuracy comparable to supervised baselines without relying on ground-truth 3D labels.

Listen

This paper introduces an unsupervised framework for jointly learning single-view depth estimation and multi-view camera pose from unlabeled monocular video. The approach addresses the long-standing difficulty of recovering 3D scene structure and ego-motion in real-world environments where labeled depth or pose data are scarce or unavailable, and where traditional geometric pipelines often fail due to textureless regions, occlusions, or non-rigid motion. By training directly on raw video, the method aims to scale to large unlabeled datasets while still producing models usable at test time without any auxiliary inputs.

The authors train two convolutional networksone that predicts per-pixel depth from a single target image and another that predicts 6-DoF relative camera poses between the target and nearby framesby minimizing a photometric reconstruction loss that warps source frames to the target using the predicted depth and pose. An auxiliary explainability mask is learned jointly to down-weight regions that violate the rigid-scene assumption, such as moving objects or disocclusions. Training uses short image sequences from the KITTI and Cityscapes driving datasets (approximately 40,000 sequences after filtering), with no ground-truth depth or pose provided. Evaluation is performed on held-out KITTI test splits and the unseen Make3D dataset.

On KITTI, the unsupervised depth model achieves absolute relative error of 0.1980.208, comparable to several supervised baselines that use either ground-truth depth or calibrated stereo pairs, though it trails the best concurrent stereo-supervised result. Pose estimation on the KITTI odometry split yields average trajectory error of 0.0200.021 on 5-frame snippets, outperforming both the dataset mean motion and short-term ORB-SLAM while remaining below full-sequence ORB-SLAM that exploits loop closure. Pre-training on Cityscapes followed by fine-tuning on KITTI provides modest further gains, and direct transfer to Make3D produces plausible global layouts despite a remaining performance gap to fully supervised methods.

These results demonstrate that view-synthesis error alone can serve as an effective training signal for geometric reasoning, removing the need for expensive labeled supervision and enabling models that generalize across datasets. For applications such as robotics, augmented reality, or autonomous driving, the approach offers a practical route to training on the vast quantities of raw video already collected by vehicles and cameras. The learned depth and pose networks can also be run independently at inference time, supporting flexible deployment.

Further progress would benefit from explicit modeling of scene dynamics, incorporation of left-right consistency constraints, relaxation of the known-intrinsics assumption to handle internet video, and extension beyond depth maps to full volumetric representations. The current framework still assumes largely rigid scenes and photometric consistency; performance degrades on thin structures, large open areas, and fast-moving objects near the camera. Results on Make3D and certain KITTI failure cases indicate that domain shift and limited scene coverage remain practical constraints. Overall, the evidence supports the core claim that unsupervised view-synthesis training can reach parity with supervised methods on standard benchmarks, but users should validate on target domains and consider hybrid supervision when highest accuracy is required.

No sufficiently relevant recommendations were found.

Cover for Unsupervised Learning of Depth and Ego-Motion from Video

Abstract

We present an unsupervised learning framework for the task of monocular depth and camera motion estimation from unstructured video sequences. We achieve this by simultaneously training depth and camera pose estimation networks using the task of view synthesis as the supervisory signal. The networks are thus coupled via the view synthesis objective during training, but can be applied independently at test time. Empirical evaluation on the KITTI dataset demonstrates the effectiveness of our approach: 1) monocular depth performing comparably with supervised methods that use either ground-truth pose or depth for training, and 2) pose estimation performing favorably with established SLAM systems under comparable input settings.

Table of Contents

  • 1 Introduction
  • 2 Related work
  • 3 Approach
  • 3.1 View synthesis as supervision
  • 3.2 Differentiable depth image-based rendering
  • 3.3 Modeling the model limitation
  • 3.4 Overcoming the gradient locality
  • 3.5 Network architecture
  • 4 Experiments
  • 4.1 Single-view depth estimation
  • 4.2 Pose estimation
  • 4.3 Visualizing the explainability prediction
  • 5 Discussion
  • References

Knowls

  1. Knowl 1 — Unsupervised Monocular Depth and Ego-Motion Learning Framework

    model/method

    The framework jointly trains two convolutional neural networks on unlabeled monocular video sequences: a single-view depth network and a multi-view camera pose network, coupled only through a photometric view-synthesis loss function during training.

    Given a sequence of video frames I1,,IN\langle I_1, \dots, I_N \rangle with a chosen target frame ItI_t and adjacent source frames IsI_s (1sN,st1 \le s \le N, s \ne t):

    1. The single-view depth network takes solely the target image ItI_t as input and estimates a per-pixel depth map D^t\hat{D}_t.
    2. The pose network takes the channel-wise concatenation of the target image ItI_t and source images IsI_s as input and predicts relative 6-DoF rigid camera transformations T^ts\hat{T}_{t \to s} for each source frame.
    3. A differentiable inverse depth-image-based rendering module warps pixels from each source image IsI_s to synthesize the target view I^s\hat{I}_s using D^t\hat{D}_t, T^ts\hat{T}_{t \to s}, and the known camera intrinsics KK.
    4. Photometric reconstruction error between ItI_t and I^s\hat{I}_s serves as the unsupervised supervision signal.

    At test time, the depth and pose networks decouple and operate independently on single images and image pairs/sequences, respectively.

  2. Knowl 2 — Differentiable Depth-Image-Based Warping Formulation

    equation

    Let ptp_t denote the homogeneous coordinate vector of a pixel in the target frame ItI_t, KK denote the 3×33 \times 3 intrinsic camera calibration matrix, D^t(pt)R+\hat{D}_t(p_t) \in \mathbb{R}^+ denote the predicted depth at pixel ptp_t, and T^tsSE(3)\hat{T}_{t \to s} \in \mathrm{SE}(3) denote the predicted 4×44 \times 4 rigid coordinate transformation matrix from target to source camera coordinates (parameterized by 3 Euler angles and a 3D translation vector). The corresponding projected coordinate psp_s in source frame IsI_s is given by:

    psKT^tsD^t(pt)K1ptp_s \sim K \hat{T}_{t \to s} \hat{D}_t(p_t) K^{-1} p_t

    Because the projected coordinates ps=(xs,ys)p_s = (x_s, y_s) are continuous, pixel values of the synthesized target image I^s(pt)\hat{I}_s(p_t) are obtained via differentiable bilinear interpolation over the four discrete neighboring pixel locations {psiji{t,b},j{l,r}}\{p_s^{ij} \mid i \in \{t, b\}, j \in \{l, r\}\} in IsI_s:

    I^s(pt)=Is(ps)=i{t,b},j{l,r}wijIs(psij)\hat{I}_s(p_t) = I_s(p_s) = \sum_{i \in \{t, b\}, j \in \{l, r\}} w^{ij} I_s(p_s^{ij})

    where wij=xsxsiysysjw^{ij} = |x_s - x_s^{i'}| \cdot |y_s - y_s^{j'}| represents the spatial proximity weight between psp_s and psijp_s^{ij} such that i,jwij=1\sum_{i,j} w^{ij} = 1.

  3. Knowl 3 — Explainability Mask Prediction for Non-Rigidity and Occlusions

    model/method

    Photometric view synthesis assumes a static scene, Lambertian surface reflectance, and absence of occlusion or disocclusion between viewpoints. Violations of these conditions produce corrupted gradients during backpropagation.

    To improve training robustness, an explainability convolutional neural network is jointly trained to output a soft confidence mask E^s(p)[0,1]\hat{E}_s(p) \in [0, 1] across multi-scale spatial resolutions for each target-source pair (It,Is)(I_t, I_s). The mask indicates the probability that direct view synthesis holds at pixel pp.

    The view-synthesis loss is weighted pixel-wise by the mask:

    Lvs=spE^s(p)It(p)I^s(p)\mathcal{L}_{vs} = \sum_{s} \sum_{p} \hat{E}_s(p) \,|I_t(p) - \hat{I}_s(p)|

    To prevent the degenerate solution where E^s(p)=0\hat{E}_s(p) = 0 everywhere, an explainability regularization loss is applied via cross-entropy with a constant ground-truth label of 11 at each pixel:

    Lreg(E^s)=plogE^s(p)\mathcal{L}_{reg}(\hat{E}_s) = - \sum_{p} \log \hat{E}_s(p)

  4. Knowl 4 — Multi-Scale Photometric and Depth Smoothness Loss Function

    equation

    To overcome gradient locality in low-texture regions and support large pixel displacements, the learning pipeline minimizes a loss function computed across multiple pyramid scales ll:

    Lfinal=l(Lvsl+λsLsmoothl+λesLreg(E^sl))\mathcal{L}_{final} = \sum_{l} \left( \mathcal{L}_{vs}^l + \lambda_s \mathcal{L}_{smooth}^l + \lambda_e \sum_{s} \mathcal{L}_{reg}(\hat{E}_s^l) \right)

    where ll indexes the 4 output resolutions, ss indexes the source views, λe=0.2\lambda_e = 0.2, and λs=0.52l\lambda_s = \frac{0.5}{2^l} where 2l2^l is the downscaling factor at scale ll.

    The depth smoothness term Lsmoothl\mathcal{L}_{smooth}^l penalizes the L1L_1 norm of the second-order spatial gradients of the predicted depth map D^l\hat{D}^l:

    Lsmoothl=p(2D^lx2(p)+2D^ly2(p)+22D^lxy(p))\mathcal{L}_{smooth}^l = \sum_{p} \left( \left| \frac{\partial^2 \hat{D}^l}{\partial x^2}(p) \right| + \left| \frac{\partial^2 \hat{D}^l}{\partial y^2}(p) \right| + 2 \left| \frac{\partial^2 \hat{D}^l}{\partial x \partial y}(p) \right| \right)

  5. Knowl 5 — Depth, Pose, and Explainability Network Architectures

    model/method

    The system utilizes two neural network pipelines:

    1. Single-View Depth Architecture: Built upon the DispNet encoder-decoder structure with skip connections and multi-scale side predictions across 4 resolutions. The first four convolutional layers use kernel sizes of 7,7,5,57, 7, 5, 5 with 32 initial output channels, and all subsequent layers use kernel size 3. All convolutions use ReLU activations except for prediction layers, which apply a scaled inverse sigmoid non-linearity to bound the depth strictly to positive values:

    D^=1αsigmoid(x)+β\hat{D} = \frac{1}{\alpha \cdot \text{sigmoid}(x) + \beta}

    with α=10\alpha = 10 and β=0.01\beta = 0.01.

    1. Pose and Explainability Architecture: The pose and explainability networks share five initial feature-encoding convolutional layers (16 initial channels, kernel sizes 7,5,5,3,37, 5, 5, 3, 3 with stride 2). The pose head continues with two stride-2 convolutions and a 1×11 \times 1 convolution producing 6(N1)6(N-1) channels for an NN-frame sequence (3 Euler angles and 3 translation components per source frame relative to the target), followed by global average pooling without activation. The explainability head branches into 5 deconvolutional layers with multi-scale side predictions, outputting 2(N1)2(N-1) channels per scale normalized via softmax over each 2-channel pair to obtain soft masks E^s\hat{E}_s.
  6. Knowl 6 — Training Specifications and Median Scale Alignment Protocol

    experimental setup

    Training is conducted with the Adam optimizer (learning rate 0.00020.0002, β1=0.9\beta_1 = 0.9, β2=0.999\beta_2 = 0.999, mini-batch size 4) using batch normalization on all non-output layers, converging around 150k iterations. Training sequences use 3-frame clips (central frame as target, ±1\pm 1 frames as source) resized to 128×416128 \times 416 pixels with static sequences (mean optical flow <1< 1 pixel) removed. Pose benchmarking uses 5-frame snippets on KITTI Odometry sequences 00–08 for training and 09–10 for testing.

    Because monocular unsupervised training exhibits an inherent global scale ambiguity, predicted depth maps D^pred\hat{D}_{pred} are aligned to sparse ground-truth depth DgtD_{gt} during evaluation by scaling with the ratio of medians:

    s^=median(Dgt)median(D^pred)\hat{s} = \frac{\text{median}(D_{gt})}{\text{median}(\hat{D}_{pred})}

    Similarly, trajectory predictions for camera motion are scale-aligned with ground-truth odometry prior to computing the Absolute Trajectory Error (ATE).

  7. Knowl 7 — Single-View Depth Estimation Performance on the KITTI Dataset

    data/table

    The single-view depth estimation performance evaluated on the 697 test images of the KITTI Eigen split is reported below. Lower is better for error metrics (Abs Rel, Sq Rel, RMSE, RMSE log); higher is better for accuracy thresholds (δ<1.25k\delta < 1.25^k for k{1,2,3}k \in \{1, 2, 3\} where δ=max(DgtD^,D^Dgt)\delta = \max(\frac{D_{gt}}{\hat{D}}, \frac{\hat{D}}{D_{gt}})). Training datasets are denoted K (KITTI) and CS (Cityscapes).

    Method Dataset Supervision Error metric Accuracy metric
    Depth Pose Abs Rel Sq Rel RMSE RMSE log δ<1.25\delta < 1.25 δ<1.252\delta < 1.25^2
    Train set mean K 0.403 5.530 8.709 0.403 0.593 0.776
    Eigen et al. Coarse K 0.214 1.605 6.563 0.292 0.673 0.884
    Eigen et al. Fine K 0.203 1.548 6.307 0.282 0.702 0.890
    Liu et al. K 0.202 1.614 6.523 0.275 0.678 0.895
    Godard et al. K 0.148 1.344 5.927 0.247 0.803 0.922
    Godard et al. CS + K 0.124 1.076 5.311 0.219 0.847 0.942
    Ours (w/o explain.) K 0.221 2.226 7.527 0.294 0.676 0.885
    Ours K 0.208 1.768 6.856 0.283 0.678 0.885
    Ours CS 0.267 2.686 7.580 0.334 0.577 0.840
    Ours CS + K 0.198 1.836 6.565 0.275 0.718 0.901
    Garg et al. cap 50m K 0.169 1.080 5.104 0.273 0.740 0.904
    Ours (w/o explain.) cap 50m K 0.208 1.551 5.452 0.273 0.695 0.900
    Ours cap 50m K 0.201 1.391 5.181 0.264 0.696 0.900
    Ours cap 50m CS 0.260 2.232 6.148 0.321 0.590 0.852
    Ours cap 50m CS + K 0.190 1.436 4.975 0.258 0.735 0.915

    The fully unsupervised model trained with Cityscapes pretraining and KITTI fine-tuning achieves an Abs Rel of 0.198, which is comparable to supervised baselines such as Eigen et al. Fine (0.203) and Liu et al. (0.202).

  8. Knowl 8 — Camera Ego-Motion Estimation on KITTI Odometry Split

    data/table

    Ego-motion performance evaluated by the Absolute Trajectory Error (ATE in meters, lower is better) averaged over all 5-frame snippets on test sequences 09 and 10 of the KITTI Odometry dataset:

    Method Seq. 09 Seq. 10
    ORB-SLAM (full) 0.014±0.0080.014 \pm 0.008 0.012±0.0110.012 \pm 0.011
    ORB-SLAM (short) 0.064±0.1410.064 \pm 0.141 0.064±0.1300.064 \pm 0.130
    Mean Odom. 0.032±0.0260.032 \pm 0.026 0.028±0.0230.028 \pm 0.023
    Ours 0.021±0.0170.021 \pm 0.017 0.020±0.0150.020 \pm 0.015

    When evaluated on 5-frame local snippets, the unsupervised pose network achieves ATE of 0.021 m0.021\text{ m} (Seq 09) and 0.020 m0.020\text{ m} (Seq 10), significantly outperforming the local 5-frame monocular baseline ORB-SLAM (short) (0.064 m0.064\text{ m}) and dataset mean odometry (0.032 m0.032\text{ m} and 0.028 m0.028\text{ m}). It falls slightly short of full ORB-SLAM (0.014 m0.014\text{ m} and 0.012 m0.012\text{ m}), which performs global loop closure and bundle adjustment over the entire sequence of >1200>1200 frames.

  9. Knowl 9 — Zero-Shot Cross-Dataset Depth Evaluation on Make3D

    data/table

    To evaluate cross-dataset generalization, the single-view depth model trained on Cityscapes and KITTI without fine-tuning on Make3D is evaluated on the Make3D test dataset (for pixels with depth <70 m< 70\text{ m} in a central image crop):

    Method Supervision Error metric
    Depth Pose Abs Rel Sq Rel RMSE RMSE log
    Train set mean 0.876 13.98 12.27 0.307
    Karsch et al. 0.428 5.079 8.389 0.149
    Liu et al. 0.475 6.562 10.05 0.165
    Laina et al. 0.204 1.840 5.683 0.084
    Godard et al. 0.544 10.94 11.76 0.193
    Ours 0.383 5.321 10.47 0.478

    Without seeing any Make3D training images, the unsupervised model achieves an Abs Rel error of 0.383 and Sq Rel error of 5.321, outperforming non-parametric and stereo-supervised transfer baselines such as Godard et al. (0.544 Abs Rel).

  10. Knowl 10 — Assumptions and Structural Limitations of Unsupervised Monocular SfM

    limitation

    The framework exhibits several structural limitations and domain assumptions:

    1. Camera Intrinsics Dependency: The differentiable rendering pipeline assumes the camera intrinsic calibration matrix KK is known and fixed during training, preventing direct application to uncalibrated web video collections.
    2. Scale Indeterminacy: Monocular view synthesis cannot resolve metric physical scale; depth and 3D translation are resolved only up to an arbitrary global scalar factor.
    3. Simplified 2.5D Scene Representation: Scene geometry is modeled as 2.5D depth maps from reference viewpoints rather than complete 3D volumetric or mesh structures, limiting representation of occluded or complex geometries.
    4. Implicit Dynamic Object Discarding: Moving objects and occlusions are identified and masked out by the explainability mask rather than explicitly modeled through motion segmentation or 3D scene flow.
    5. Over-Masking of Thin Structures: The depth network often displays low confidence on thin structures (e.g., street poles, distant trees) or vast open scenes, leading the explainability network to discount these regions during optimization.

Coverage note — No substantial contributed material was omitted. All core architectural components, equations, training losses, experimental benchmarks (KITTI depth, KITTI odometry, Make3D generalization), and model limitations have been captured.

References

  1. 1.M. Abadi, A. Agarwal, P. Barham, E. Brevdo, Z. Chen, C. Citro, G. S. Corrado, A. Davis, J. Dean, M. Devin, et al. TensorFlow: Large-scale machine learning on heterogeneous distributed systems. arXiv preprint arXiv:1603.04467, 2016.
  2. 2.P. Agrawal, J. Carreira, and J. Malik. Learning to see by moving. In Int. Conf. Computer Vision, 2015.
  3. 3.J. Bergen, P. Anandan, K. Hanna, and R. Hingorani. Hierarchical model-based motion estimation. In Computer VisionECCV’92, pages 237–252. Springer, 1992.
  4. 4.S. E. Chen and L. Williams. View interpolation for image synthesis. In Proceedings of the 20th annual conference on Computer graphics and interactive techniques, pages 279–288. ACM, 1993.
  5. 5.M. Cordts, M. Omran, S. Ramos, T. Rehfeld, M. Enzweiler, R. Benenson, U. Franke, S. Roth, and B. Schiele. The Cityscapes dataset for semantic urban scene understanding. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 3213–3223, 2016.
  6. 6.P. E. Debevec, C. J. Taylor, and J. Malik. Modeling and rendering architecture from photographs: A hybrid geometryand image-based approach. In Proceedings of the 23rd annual conference on Computer graphics and interactive techniques, pages 11–20. ACM, 1996.
  7. 7.D. Eigen, C. Puhrsch, and R. Fergus. Depth map prediction from a single image using a multi-scale deep network. In Advances in Neural Information Processing Systems, 2014.
  8. 8.C. Fehn. Depth-image-based rendering (dibr), compression, and transmission for a new approach on 3d-tv. In Electronic Imaging 2004, pages 93–104. International Society for Optics and Photonics, 2004.
  9. 9.A. Fitzgibbon, Y. Wexler, and A. Zisserman. Image-based rendering using image-based priors. Int. Journal of Computer Vision, 63(2):141–151, 2005.
  10. 10.J. Flynn, I. Neulander, J. Philbin, and N. Snavely. DeepStereo: Learning to predict new views from the world’s imagery. In Computer Vision and Pattern Recognition, 2016.
  11. 11.D. F. Fouhey, W. Hussain, A. Gupta, and M. Hebert. Single image 3D without a single 3D image. In Proceedings of the IEEE International Conference on Computer Vision, pages 1053–1061, 2015.
  12. 12.Y. Furukawa, B. Curless, S. M. Seitz, and R. Szeliski. Towards internet-scale multi-view stereo. In Computer Vision and Pattern Recognition, pages 1434–1441. IEEE, 2010.
  13. 13.M. Gadelha, S. Maji, and R. Wang. 3d shape induction from 2d views of multiple objects. arXiv preprint arXiv:1612.05872, 2016.
  14. 14.R. Garg, V. K. BG, G. Carneiro, and I. Reid. Unsupervised CNN for single view depth estimation: Geometry to the rescue. In European Conf. Computer Vision, 2016.
  15. 15.A. Geiger, P. Lenz, and R. Urtasun. Are we ready for autonomous driving? The KITTI vision benchmark suite. In Computer Vision and Pattern Recognition (CVPR), 2012 IEEE Conference on, pages 3354–3361. IEEE, 2012.
  16. 16.C. Godard, O. Mac Aodha, and G. J. Brostow. Unsupervised monocular depth estimation with left-right consistency. In Computer Vision and Pattern Recognition, 2017.
  17. 17.R. Goroshin, J. Bruna, J. Tompson, D. Eigen, and Y. LeCun. Unsupervised learning of spatiotemporally coherent metrics. In Proceedings of the IEEE International Conference on Computer Vision, pages 4086–4093, 2015.
  18. 18.X. Han, T. Leung, Y. Jia, R. Sukthankar, and A. C. Berg. MatchNet: Unifying feature and metric learning for patchbased matching. In Computer Vision and Pattern Recognition, pages 3279–3286, 2015.
  19. 19.A. Handa, M. Bloesch, V. Patraucean, S. Stent, J. McCormac, and A. Davison. gvnn: Neural network library for geometric computer vision. arXiv preprint arXiv:1607.07405, 2016.
  20. 20.D. Hoiem, A. A. Efros, and M. Hebert . Automatic photo pop-up. In Proc. SIGGRAPH, 2005.
  21. 21.S. Ioffe and C. Szegedy. Batch normalization: Accelerating deep network training by reducing internal covariate shift. arXiv preprint arXiv:1502.03167, 2015.
  22. 22.M. Irani and P. Anandan. About direct methods. In International Workshop on Vision Algorithms, pages 267–277. Springer, 1999.
  23. 23.M. Jaderberg, K. Simonyan, A. Zisserman, et al. Spatial transformer networks. In Advances in Neural Information Processing Systems, pages 2017–2025, 2015.
  24. 24.D. Jayaraman and K. Grauman. Learning image representations tied to egomotion. In Int. Conf. Computer Vision, 2015.
  25. 25.K. Karsch, C. Liu, and S. B. Kang. Depth transfer: Depth extraction from video using non-parametric sampling. IEEE transactions on pattern analysis and machine intelligence, 36(11):2144–2158, 2014.
  26. 26.A. Kendall, M. Grimes, and R. Cipolla. PoseNet: A convolutional network for real-time 6-DOF camera relocalization. In Int. Conf. Computer Vision, pages 2938–2946, 2015.
  27. 27.A. Kendall, H. Martirosyan, S. Dasgupta, P. Henry, R. Kennedy, A. Bachrach, and A. Bry. End-to-end learning of geometry and context for deep stereo regression. arXiv preprint arXiv:1703.04309, 2017.
  28. 28.D. Kingma and J. Ba. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014.
  29. 29.T. D. Kulkarni, W. F. Whitney, P. Kohli, and J. Tenenbaum. Deep convolutional inverse graphics network. In C. Cortes, N. D. Lawrence, D. D. Lee, M. Sugiyama, and R. Garnett, editors, Advances in Neural Information Processing Systems, pages 2539–2547. Curran Associates, Inc., 2015.
  30. 30.Y. Kuznietsov, J. Stuckler, and B. Leibe. Semi-supervised deep learning for monocular depth map prediction. arXiv preprint arXiv:1702.02706, 2017.
  31. 31.I. Laina, C. Rupprecht, V. Belagiannis, F. Tombari, and N. Navab. Deeper depth prediction with fully convolutional residual networks. In 3D Vision (3DV), 2016 Fourth International Conference on, pages 239–248. IEEE, 2016.
  32. 32.F. Liu, C. Shen, G. Lin, and I. Reid. Learning depth from single monocular images using deep convolutional neural fields. IEEE transactions on pattern analysis and machine intelligence, 38(10):2024–2039, 2016.
  33. 33.M. Liu, M. Salzmann, and X. He. Discrete-continuous depth estimation from a single image. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 716–723, 2014.
  34. 34.M. M. Loper and M. J. Black. OpenDR: An approximate differentiable renderer. In European Conf. Computer Vision, pages 154–169. Springer, 2014.
  35. 35.N. Mayer, E. Ilg, P. Hausser, P. Fischer, D. Cremers, A. Dosovitskiy, and T. Brox. A large dataset to train convolutional networks for disparity, optical flow, and scene flow estimation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 4040–4048, 2016.
  36. 36.I. Misra, C. L. Zitnick, and M. Hebert. Shuffle and learn: unsupervised learning using temporal order verification. In European Conference on Computer Vision, pages 527–544. Springer, 2016.
  37. 37.R. Mur-Artal, J. M. M. Montiel, and J. D. Tardos. ORBSLAM: a versatile and accurate monocular SLAM system. IEEE Transactions on Robotics, 31(5), 2015.
  38. 38.R. A. Newcombe, S. J. Lovegrove, and A. J. Davison. DTAM: Dense tracking and mapping in real-time. In Int. Conf. Computer Vision, pages 2320–2327. IEEE, 2011.
  39. 39.D. Pathak, R. Girshick, P. Dollar, T. Darrell, and B. Hariharan. Learning features by watching objects move. In CVPR, 2017.
  40. 40.R. Ranftl, V. Vineet, Q. Chen, and V. Koltun. Dense monocular depth estimation in complex dynamic scenes. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 4058–4066, 2016.
  41. 41.D. J. Rezende, S. A. Eslami, S. Mohamed, P. Battaglia, M. Jaderberg, and N. Heess. Unsupervised learning of 3d structure from images. In Advances In Neural Information Processing Systems, pages 4997–5005, 2016.
  42. 42.A. Saxena, M. Sun, and A. Y. Ng. Make3D: Learning 3D scene structure from a single still image. Pattern Analysis and Machine Intelligence, 31(5):824–840, May 2009.
  43. 43.S. M. Seitz and C. R. Dyer. View morphing. In Proceedings of the 23rd annual conference on Computer graphics and interactive techniques, pages 21–30. ACM, 1996.
  44. 44.R. Szeliski. Prediction error as a quality metric for motion and stereo. In Int. Conf. Computer Vision, volume 2, pages 781–788. IEEE, 1999.
  45. 45.M. Tatarchenko, A. Dosovitskiy, and T. Brox. Multi-view 3d models from single images with a convolutional network. In European Conference on Computer Vision, pages 322–337. Springer, 2016.
  46. 46.S. Tulsiani, T. Zhou, A. A. Efros, and J. Malik. Multi-view supervision for single-view reconstruction via differentiable ray consistency. In Computer Vision and Pattern Recognition, 2017.
  47. 47.B. Ummenhofer, H. Zhou, J. Uhrig, N. Mayer, E. Ilg, A. Dosovitskiy, and T. Brox. DeMoN: Depth and motion network for learning monocular stereo. arXiv preprint arXiv:1612.02401, 2016.
  48. 48.S. Vijayanarasimhan, S. Ricco, C. Schmid, R. Sukthankar, and K. Fragkiadaki. SfM-Net: Learning of structure and motion from video. arXiv preprint, 2017.
  49. 49.X. Wang and A. Gupta. Unsupervised learning of visual representations using videos. In Proceedings of the IEEE International Conference on Computer Vision, pages 2794–2802, 2015.
  50. 50.C. Wu. VisualSFM: A visual structure from motion system. http://ccwu.me/vsfm, 2011.
  51. 51.J. Xie, R. B. Girshick, and A. Farhadi. Deep3D: Fully automatic 2D-to-3D video conversion with deep convolutional neural networks. In European Conf. Computer Vision, 2016.
  52. 52.X. Yan, J. Yang, E. Yumer, Y. Guo, and H. Lee. Perspective transformer nets: Learning single-view 3d object reconstruction without 3d supervision. In Advances in Neural Information Processing Systems, pages 1696–1704, 2016.
  53. 53.J. Zbontar and Y. LeCun. Stereo matching by training a convolutional neural network to compare image patches. Journal of Machine Learning Research, 17(1-32):2, 2016.
  54. 54.T. Zhou, S. Tulsiani, W. Sun, J. Malik, and A. A. Efros. View synthesis by appearance flow. In European Conference on Computer Vision, pages 286–301. Springer, 2016.
  55. 55.C. L. Zitnick, S. B. Kang, M. Uyttendaele, S. Winder, and R. Szeliski. High-quality video view interpolation using a layered representation. In ACM Transactions on Graphics (TOG), volume 23, pages 600–608. ACM, 2004.

Citation

MLA
Zhou, T., et al. “Unsupervised Learning of Depth and Ego-Motion from Video”. 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017, pp. 6612–19, https://doi.org/10.1109/CVPR.2017.700.
APA
Zhou, T., Brown, M., Snavely, N., & Lowe, D. G. (2017). Unsupervised Learning of Depth and Ego-Motion from Video. 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 6612–6619. https://doi.org/10.1109/CVPR.2017.700
Chicago
Zhou, T., M. Brown, N. Snavely, and D. G. Lowe. 2017. “Unsupervised Learning of Depth and Ego-Motion from Video”. 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 6612–19. https://doi.org/10.1109/CVPR.2017.700.
Harvard
Zhou, T. et al. (2017) “Unsupervised Learning of Depth and Ego-Motion from Video”, 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp. 6612–6619. Available at: https://doi.org/10.1109/CVPR.2017.700.
Vancouver
1. Zhou T, Brown M, Snavely N, Lowe DG (2017) Unsupervised Learning of Depth and Ego-Motion from Video. In: 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp 6612–6619

BibTeX

@inproceedings{Zhou_2017, title={Unsupervised Learning of Depth and Ego-Motion from Video}, url={http://dx.doi.org/10.1109/CVPR.2017.700}, DOI={10.1109/cvpr.2017.700}, booktitle={2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)}, publisher={IEEE}, author={Zhou, Tinghui and Brown, Matthew and Snavely, Noah and Lowe, David G.}, year={2017}, month=July, pages={6612–6619} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF

License: IEEE