MVSNet: Depth Inference for Unstructured Multi-view Stereo

Yao YaoZixin LuoShiwei LiTian FangLong Quan

article2018ECCV1,714 citations

Proposes an end-to-end deep learning network that reconstructs dense 3D geometry from arbitrary multi-view images using differentiable homography warping and variance-based cost volumes, dramatically improving reconstruction speed and accuracy over traditional multi-view stereo methods.

Listen

Multi-view stereo technology reconstructs dense 3D representations from sets of overlapping two-dimensional images. While traditional methods perform accurately in ideal conditions, they struggle with low-textured, specular, or reflective surfaces, leading to incomplete reconstructions. Recent deep learning approaches show promise but have suffered from severe scalability and memory limitations because they attempt to process entire 3D scenes at once using rigid volumetric grids. Overcoming these reconstruction bottlenecks is crucial for modern applications requiring fast, high-quality 3D models from ordinary imagery.

The article demonstrates an end-to-end deep learning framework, named MVSNet, designed to infer depth maps view-by-view from unstructured multi-view images. The core objective is to evaluate whether decoupling 3D reconstruction into individual per-view depth map estimations using learned geometric features can simultaneously boost reconstruction completeness, overall accuracy, and runtime efficiency.

The authors approach the problem by extracting multi-scale visual features from input images and warping them into the reference camera frustum using differentiable geometric transformations. A variance-based metric aggregates features across an arbitrary number of views into a compact 3D cost volume, which a 3D neural network regularizes to predict continuous depth values and measure estimation confidence. The model was trained on the indoor DTU dataset containing over 27,000 samples and evaluated across standard benchmarks, with additional out-of-domain testing performed on the complex outdoor Tanks and Temples dataset.

Evaluation results show substantial improvements across key operational metrics. On the DTU dataset, MVSNet outperformed all competing algorithms in reconstruction completeness (0.527 mm) and achieved the best overall score (0.462 mm), surpassing previous leading methods like Gipuma (0.578 mm) and SurfaceNet (0.745 mm). In outdoor testing on the Tanks and Temples benchmark, the model achieved the top overall ranking among all submissions—including leading commercial and open-source software—without requiring any dataset-specific fine-tuning. Furthermore, MVSNet demonstrated exceptional speed, processing one DTU scan in approximately 230 seconds (4.7 seconds per view), which is roughly 5 times faster than Gipuma, 100 times faster than COLMAP, and 160 times faster than SurfaceNet.

These findings indicate that deep learning can be deployed for large-scale multi-view 3D reconstruction without prohibitive computational bottlenecks. Decoupling the problem into per-view depth maps enables organizations to achieve higher reconstruction completeness on difficult surfaces—such as reflections and textureless zones—while drastically reducing runtime and hardware costs. The strong out-of-the-box generalization to outdoor scenes confirms that the model learns fundamental geometric matching principles rather than simply memorizing training environments.

Organizations implementing dense 3D reconstruction pipelines should consider adopting frustum-based learning frameworks to accelerate throughput and improve model completeness. For production integration, standard depth filtering and fusion techniques should be maintained to remove background noise and occlusions. Where hardware permits, teams can balance speed and point-cloud quality by adjusting the number of input views and depth hypothesis resolution to match available GPU capacity.

The results should be interpreted in the context of specific training constraints. Ground-truth depth maps rendered from incomplete training meshes occasionally introduce background artifacts or fail to account for pixels that are occluded across all views. While the authors demonstrate high confidence in the method across indoor and outdoor settings, performance in deployment will still depend on the accuracy of upstream camera parameter estimation and sufficient graphics memory to handle high-resolution inputs.

  • Paper: Pixelwise View Selection for Unstructured Multi-View Stereo, Johannes L. Schönberger et al. (2016). This work establishes modern patch-based depth estimation and pixelwise view-selection strategies for unstructured multi-view stereo that classical and learning-based MVS pipelines directly build upon.
  • Paper: Structure-from-Motion Revisited, Johannes L. Schönberger et al. (2016). This paper presents the foundational COLMAP structure-from-motion pipeline and geometric principles used to provide calibrated camera poses for unstructured multi-view stereo.
  • Paper: Accurate, Dense, and Robust Multiview Stereopsis, Yasutaka Furukawa et al. (2010). This classic paper introduces patch-based multi-view stereopsis (PMVS), defining standard photometric consistency and geometric filtering baselines in 3D multi-view reconstruction.
  • Paper: A Comparison and Evaluation of Multi-View Stereo Reconstruction Algorithms, Steven M. Seitz et al. (2006). This benchmark paper establishes the standardized evaluation taxonomy, metrics, and ground-truth comparison methodologies essential for evaluating multi-view stereo algorithms.
  • Paper: Modeling the World from Internet Photo Collections, Noah Snavely et al. (2008). This foundational work introduces the paradigm of recovering calibrated 3D scene geometry from unstructured, uncontrolled Internet photo collections.
Cover for MVSNet: Depth Inference for Unstructured Multi-view Stereo

Abstract

We present an end-to-end deep learning architecture for depth map inference from multi-view images. In the network, we first extract deep visual image features, and then build the 3D cost volume upon the reference camera frustum via the differentiable homography warping. Next, we apply 3D convolutions to regularize and regress the initial depth map, which is then refined with the reference image to generate the final output. Our framework flexibly adapts arbitrary N-view inputs using a variance-based cost metric that maps multiple features into one cost feature. The proposed MVSNet is demonstrated on the large-scale indoor DTU dataset. With simple post-processing, our method not only significantly outperforms previous state-of-the-arts, but also is several times faster in runtime. We also evaluate MVSNet on the complex outdoor Tanks and Temples dataset, where our method ranks first before April 18, 2018 without any fine-tuning, showing the strong generalization ability of MVSNet.

Table of Contents

  • 1 Introduction
  • 2 Related work
  • 3 MVSNet
  • 3.1 Image Features
  • 3.2 Cost Volume
  • Differentiable Homography
  • Cost Metric
  • Cost Volume Regularization
  • 3.3 Depth Map
  • Initial Estimation
  • Probability Map
  • Depth Map Refinement
  • 3.4 Loss
  • 4 Implementations
  • 4.1 Training
  • Data Preparation
  • View Selection
  • 4.2 Post-processing
  • Depth Map Filter
  • Depth Map Fusion
  • 5 Experiments
  • 5.1 Benchmarking on DTU dataset
  • 5.2 Generalization on Tanks and Temples dataset
  • 5.3 Ablations
  • View Number
  • Image Features
  • Cost Metric
  • Depth Refinement
  • 5.4 Discussions
  • Running Time
  • GPU Memory
  • Training Data
  • 6 Conclusion
  • References

Knowls

  1. Knowl 1 — MVSNet Architecture for Multi-View Stereo Depth Map Inference

    model/method

    MVSNet is an end-to-end deep neural network architecture designed for depth map estimation from unstructured multi-view images. Given a reference image I1∈RH×W×3I_1 \in \mathbb{R}^{H \times W \times 3} and N−1N-1 source images {Ii}i=2N\{I_i\}_{i=2}^N along with known camera intrinsics and extrinsics {Ki,Ri,ti}i=1N\{K_i, R_i, t_i\}_{i=1}^N, MVSNet infers the depth map corresponding to the reference camera view.

    The pipeline consists of the following sequential stages:

    1. 2D Feature Extraction: A weight-shared 8-layer 2D convolutional neural network extracts deep feature representations {Fi}i=1N\{F_i\}_{i=1}^N from all input images. The network divides into three scales with strided convolutions (stride 2 at layers 3 and 6), producing 32-channel feature maps downscaled by a factor of 4 in each spatial dimension (Fi∈RH4×W4×32F_i \in \mathbb{R}^{\frac{H}{4} \times \frac{W}{4} \times 32}).
    2. Differentiable Homography Warping: Using plane sweep homography, source feature maps {Fi}i=2N\{F_i\}_{i=2}^N and the reference feature map F1F_1 are warped into DD fronto-parallel planes within the reference camera frustum [dmin⁡,dmax⁡][d_{\min}, d_{\max}] via differentiable bilinear sampling, constructing NN feature volumes {Vi}i=1N\{V_i\}_{i=1}^N of shape W4×H4×D×32\frac{W}{4} \times \frac{H}{4} \times D \times 32.
    3. Cost Volume Aggregation: The NN feature volumes are aggregated into a single matching cost volume C∈RW4×H4×D×32C \in \mathbb{R}^{\frac{W}{4} \times \frac{H}{4} \times D \times 32} using a view-symmetric, variance-based operation.
    4. 3D Cost Volume Regularization: A multi-scale 3D CNN structured as a 4-scale 3D U-Net regularizes the raw cost volume to filter matching noise and output a 1-channel volume. A depth-wise softmax operation is applied along the depth dimension to generate a probability volume P∈RW4×H4×DP \in \mathbb{R}^{\frac{W}{4} \times \frac{H}{4} \times D}.
    5. Initial Depth Regression: A differentiable soft argmin operation computes the expectation value along the depth hypotheses to retrieve a continuous initial depth map D^i\hat{D}_i.
    6. Depth Map Refinement: A 2D residual CNN conditioned on the reference image I1I_1 refines D^i\hat{D}_i to recover sharp object boundaries.
  2. Knowl 2 — Frustum-Based Differentiable Homography for Multi-View Cost Volume Warping

    equation

    To build matching cost volumes directly within the reference camera's frustum rather than in Euclidean regular grids, 2D feature maps from source views are warped into fronto-parallel planes of the reference camera at depth hypothesis dd. The planar homography Hi(d)∈R3×3H_i(d) \in \mathbb{R}^{3 \times 3} mapping pixel coordinates xx from the reference feature plane to coordinates x′∼Hi(d)xx' \sim H_i(d) x in the ii-th source feature map is defined as:

    Hi(d)=Ki⋅Ri⋅(I−(t1−ti)⋅n1Td)⋅R1T⋅K1−1H_i(d) = K_i \cdot R_i \cdot \left( I - \frac{(t_1 - t_i) \cdot n_1^T}{d} \right) \cdot R_1^T \cdot K_1^{-1}

    where:

    • K1,Ki∈R3×3K_1, K_i \in \mathbb{R}^{3 \times 3} denote the camera intrinsic calibration matrices of the reference view and the ii-th source view, respectively.
    • R1,Ri∈SO(3)R_1, R_i \in SO(3) and t1,ti∈R3t_1, t_i \in \mathbb{R}^3 are the camera rotation matrices and translation vectors.
    • n1∈R3n_1 \in \mathbb{R}^3 is the principal optical axis of the reference camera (the normal vector of the fronto-parallel sweeping planes).
    • d∈[dmin⁡,dmax⁡]d \in [d_{\min}, d_{\max}] denotes the discrete depth hypothesis sampled along the reference camera optical axis.
    • I∈R3×3I \in \mathbb{R}^{3 \times 3} is the identity matrix.

    For the reference image feature map F1F_1, H1(d)=IH_1(d) = I. Warped feature values are sampled via differentiable bilinear interpolation from FiF_i, enabling end-to-end backpropagation.

  3. Knowl 3 — Variance-Based Multi-View Cost Aggregation Metric

    equation

    To aggregate feature volumes from an arbitrary number of input views N≥2N \ge 2 into a single matching cost volume without inducing reference-view bias, the matching cost volume C∈RVC \in \mathbb{R}^{V} (where V=W4×H4×D×FV = \frac{W}{4} \times \frac{H}{4} \times D \times F) is computed via the element-wise variance across all warped feature volumes:

    C=M(V1,…,VN)=1N∑i=1N(Vi−Vˉ)2C = \mathcal{M}(V_1, \dots, V_N) = \frac{1}{N} \sum_{i=1}^N (V_i - \bar{V})^2

    where Vˉ=1N∑i=1NVi\bar{V} = \frac{1}{N} \sum_{i=1}^N V_i is the element-wise average feature volume among all NN warped feature volumes {Vi}i=1N\{V_i\}_{i=1}^N.

    Unlike pairwise correlation or mean-based aggregation, the variance metric treats all input views symmetrically and explicitly measures multi-view feature discrepancies across depth planes.

  4. Knowl 4 — Soft Argmin Depth Regression and Probability Confidence Estimation

    model/method

    The regularized probability volume P∈RW4×H4×DP \in \mathbb{R}^{\frac{W}{4} \times \frac{H}{4} \times D} provides a normalized probability distribution P(p,d)P(p, d) over DD uniformly sampled discrete depth hypotheses d∈{dmin⁡,dmin⁡+Δd,…,dmax⁡}d \in \{d_{\min}, d_{\min} + \Delta d, \dots, d_{\max}\} for each pixel pp.

    1. Continuous Depth Regression: Instead of the non-differentiable winner-take-all argmax, the initial continuous depth map D^i\hat{D}_i is computed as the expected depth along the depth axis (soft argmin):

    D^i(p)=∑d=dmin⁡dmax⁡d⋅P(p,d)\hat{D}_i(p) = \sum_{d=d_{\min}}^{d_{\max}} d \cdot P(p, d)

    1. Estimation Quality / Probability Map: The confidence of the depth estimation at pixel pp is quantified by summing the probability values over the four nearest discrete depth hypotheses N4(D^i(p))\mathcal{N}_4(\hat{D}_i(p)) surrounding the estimated depth:

    P(p)=∑d∈N4(D^i(p))P(p,d)\mathcal{P}(p) = \sum_{d \in \mathcal{N}_4(\hat{D}_i(p))} P(p, d)

    Pixels with unimodal, sharp distributions yield high confidence scores (close to 1), whereas mismatched or occluded pixels produce dispersed, multi-modal probability distributions with low confidence scores.

  5. Knowl 5 — Reference-Guided Depth Map Residual Refinement

    model/method

    Due to the large spatial receptive field of the multi-scale 3D regularizing convolutions, the initial depth map D^i\hat{D}_i often suffers from oversmoothed depth discontinuities at object boundaries. A depth residual learning sub-network refines the initial depth map using the reference RGB image as guidance.

    The refinement procedure operates as follows:

    1. The initial depth map D^i\hat{D}_i is pre-scaled to the range [0,1][0, 1].
    2. The pre-scaled initial depth map and the downsampled reference image I1I_1 (matching the spatial resolution H4×W4\frac{H}{4} \times \frac{W}{4}) are concatenated into a 4-channel tensor.
    3. The concatenated tensor is passed through three 2D convolutional layers with 32 channels (each with batch normalization and ReLU activations), followed by a final 1-channel convolutional layer without batch normalization or ReLU to output a signed depth residual ΔD\Delta D.
    4. The residual is added to the initial depth map and transformed back to the original depth scale to produce the refined depth map D^r=D^i+ΔD\hat{D}_r = \hat{D}_i + \Delta D.
  6. Knowl 6 — MVSNet Multi-Stage Depth Training Loss

    equation

    MVSNet is trained end-to-end using the mean absolute error (L1L_1 loss) between the predicted depth values and the ground-truth depth labels across all valid labeled pixels, penalizing both the initial and refined depth predictions:

    Loss=∑p∈pvalid∥d(p)−d^i(p)∥1+λ⋅∥d(p)−d^r(p)∥1\text{Loss} = \sum_{p \in p_{\text{valid}}} \|d(p) - \hat{d}_i(p)\|_1 + \lambda \cdot \|d(p) - \hat{d}_r(p)\|_1

    where:

    • pvalidp_{\text{valid}} is the set of image pixels with valid ground-truth depth labels.
    • d(p)∈R+d(p) \in \mathbb{R}^+ denotes the ground-truth depth value at pixel pp.
    • d^i(p)∈R+\hat{d}_i(p) \in \mathbb{R}^+ is the initial depth estimate regressed via soft argmin from the regularized probability volume.
    • d^r(p)∈R+\hat{d}_r(p) \in \mathbb{R}^+ is the refined depth estimate produced by the guided residual network.
    • λ\lambda is a balancing weighting hyperparameter, set to λ=1.0\lambda = 1.0.
  7. Knowl 7 — Two-Step Depth Map Filtering and Fusion Pipeline

    algorithm

    Before generating dense point clouds, raw estimated depth maps are filtered for outliers and merged using photometric and geometric consistency criteria:

    Input: Depth maps DkD_k, probability confidence maps Pk\mathcal{P}_k, and camera projection parameters (Kk,Rk,tk)(K_k, R_k, t_k) for each view k∈{1,…,M}k \in \{1, \dots, M\}.
    Output: Reconstructed 3D point cloud X\mathcal{X}.
    for each view k∈{1,…,M}k \in \{1, \dots, M\} do
        for each pixel p1p_1 in view kk do
            // Step 1: Photometric filtering
            if Pk(p1)<0.8\mathcal{P}_k(p_1) < 0.8 then
                Mark p1p_1 as invalid (outlier)
                continue
            end if
            // Step 2: Multi-view geometric consistency check
            Initialize consistent view set Vconsistent←{k}V_{\text{consistent}} \leftarrow \{k\}
            Initialize reprojected depth list Dreproj←{Dk(p1)}\mathcal{D}_{\text{reproj}} \leftarrow \{D_k(p_1)\}
            for each source view j≠kj \ne k do
                Project pixel p1p_1 with depth d1=Dk(p1)d_1 = D_k(p_1) into view jj to obtain coordinate pjp_j
                Retrieve depth dj=Dj(pj)d_j = D_j(p_j) from view jj
                Reproject pjp_j with depth djd_j back to view kk obtaining coordinate preprojp_{\text{reproj}} and depth dreprojd_{\text{reproj}}
                if ∥preproj−p1∥2<1.0 pixel\|p_{\text{reproj}} - p_1\|_2 < 1.0\text{ pixel} and ∣dreproj−d1∣d1<0.01\frac{|d_{\text{reproj}} - d_1|}{d_1} < 0.01 then
                    Vconsistent←Vconsistent∪{j}V_{\text{consistent}} \leftarrow V_{\text{consistent}} \cup \{j\}
                    Dreproj←Dreproj∪{dreproj}\mathcal{D}_{\text{reproj}} \leftarrow \mathcal{D}_{\text{reproj}} \cup \{d_{\text{reproj}}\}
                end if
            end for
            // Check minimum view support (at least 3-view consistent)
            if ∣Vconsistent∣≥3|V_{\text{consistent}}| \ge 3 then
                Dk∗(p1)←mean(Dreproj)D_k^*(p_1) \leftarrow \text{mean}(\mathcal{D}_{\text{reproj}})
                Mark p1p_1 as valid
            else
                Mark p1p_1 as invalid
            end if
        end for
    end for
    Reproject all valid pixels p1p_1 with fused depth Dk∗(p1)D_k^*(p_1) into 3D Euclidean space to form point cloud X\mathcal{X}.
    return X\mathcal{X}
  8. Knowl 8 — Quantitative Reconstruction Performance on the DTU Dataset

    data/table

    The reconstruction quality of MVSNet evaluated on the 22 evaluation scans of the DTU dataset demonstrates state-of-the-art completeness and overall accuracy-completeness metrics across distance thresholds compared to traditional and learning-based MVS baselines:

    Method Mean Distance (mm) Percentage (<<1 mm) Percentage (<<2 mm)
    Acc. Comp. overall Acc. Comp. f-score Acc. Comp. f-score
    Camp 0.835 0.554 0.695 71.75 64.94 66.31 84.83 67.82 73.02
    Furu 0.613 0.941 0.777 69.55 61.52 63.26 78.99 67.88 70.93
    Tola 0.342 1.190 0.766 90.49 57.83 68.07 93.94 63.88 73.61
    Gipuma 0.283 0.873 0.578 94.65 59.93 70.64 96.42 63.81 74.16
    SurfaceNet 0.450 1.040 0.745 83.80 63.38 69.95 87.15 67.99 74.40
    MVSNet (Ours) 0.396 0.527 0.462 86.46 71.13 75.69 91.06 75.31 80.25

    For the mean distance metric (lower is better; overall score defined as Acc.+Comp.2\frac{\text{Acc.} + \text{Comp.}}{2}), MVSNet achieves the lowest mean completeness error (0.527 mm0.527\text{ mm}) and best overall score (0.462 mm0.462\text{ mm}). Under the percentage metrics (higher is better), MVSNet achieves the highest ff-scores (75.69%75.69\% at 1 mm1\text{ mm} and 80.25%80.25\% at 2 mm2\text{ mm}).

  9. Knowl 9 — Benchmark Performance on the Tanks and Temples Dataset

    data/table

    Generalization performance evaluated on the outdoor Tanks and Temples benchmark using the MVSNet model trained exclusively on the indoor DTU dataset without any fine-tuning:

    Method Rank Mean Family Francis Horse Lighthouse M60 Panther Playground Train
    MVSNet (Ours) 3.00 43.48 55.99 28.55 25.07 50.79 53.96 50.86 47.90 34.69
    Pix4D 3.12 43.24 64.45 31.91 26.43 54.41 50.58 35.37 47.78 34.96
    COLMAP 3.50 42.14 50.41 22.25 25.63 56.43 44.83 46.97 48.53 42.04
    OpenMVG + OpenMVS 3.62 41.71 58.86 32.59 26.25 43.12 44.73 46.85 45.97 35.27
    OpenMVG + MVE 6.00 38.00 49.91 28.19 20.75 43.35 44.51 44.76 36.58 35.95
    OpenMVG + SMVS 10.38 30.67 31.93 19.92 15.02 39.38 36.51 41.61 35.89 25.12
    OpenMVG-G + OpenMVS 10.88 22.86 56.50 29.63 21.69 6.55 39.54 28.48 0.00 0.53
    MVE 11.25 25.37 48.59 23.84 12.70 5.07 39.62 38.16 5.81 29.19
    OpenMVG + PMVS 11.88 29.66 41.03 17.70 12.83 36.68 35.93 33.20 31.78 28.10

    MVSNet achieved the highest mean ff-score (43.4843.48) and best average rank (3.003.00) among all submissions (prior to April 18, 2018), outperforming established traditional pipelines including COLMAP, Pix4D, and OpenMVS without domain-specific retraining.

  10. Knowl 10 — Ablation Studies on View Count, Image Features, Cost Metric, and Depth Refinement

    empirical result

    Ablation experiments on MVSNet demonstrate the following properties:

    1. Arbitrary View Scalability: Evaluating an MVSNet model trained on N=3N=3 views with N=2,3,5N=2, 3, 5 input views reveals that validation loss strictly decreases as the number of input views increases. Testing with N=5N=5 yields superior results to N=3N=3, confirming generalization to variable camera counts.
    2. Deep vs. Shallow Patch Features: Replacing the 8-layer 2D feature extractor with a single 32-channel 7×77 \times 7 convolutional layer (stride 4, simulating traditional patch matching) results in substantially higher validation loss, indicating that multi-scale deep descriptors are critical for dense matching.
    3. Variance vs. Mean Cost Metric: Substituting the variance-based feature metric with a mean aggregation metric results in slower optimization convergence and higher validation loss, confirming that explicit multi-view variance provides a superior matching signal.
    4. Depth Refinement Impact: Incorporating reference-guided residual refinement improves the DTU evaluation ff-score from 75.58%75.58\% to 75.69%75.69\% at the <1 mm<1\text{ mm} threshold, and from 79.98%79.98\% to 80.25%80.25\% at the <2 mm<2\text{ mm} threshold.
    5. Runtime Efficiency: Reconstructing an entire DTU scan requires approximately 230 s230\text{ s} (4.7 s4.7\text{ s} per view), which is ≈5×\approx 5\times faster than Gipuma, ≈100×\approx 100\times faster than COLMAP, and ≈160×\approx 160\times faster than SurfaceNet.
  11. Knowl 11 — Limitations in Ground-Truth Depth Rendering and Multi-View Occlusion

    limitation

    The supervision of MVSNet depends on depth maps rendered from 3D surface meshes generated via Screened Poisson Surface Reconstruction (SPSR) on raw point clouds. This setup exhibits two primary limitations:

    1. Incomplete Mesh Artifacts: Incomplete ground-truth point clouds yield non-watertight mesh surfaces with holes. Rendering depth maps from such meshes can cause background triangles behind foreground geometry to be falsely rendered as valid foreground depths, injecting false labels into the L1L_1 loss.
    2. Multi-View Occlusions: Pixels occluded across all source views should ideally be excluded from the matching cost loss. However, without complete ground-truth scene meshes, occluded pixels cannot be rigorously identified and filtered out during training.
    3. Dataset Dependency: Datasets that lack 3D point normal vectors or dense surface meshes (such as Tanks and Temples) cannot be directly rendered into ground-truth depth maps for fine-tuning.

Coverage note — The baseline view selection scoring formula based on piecewise Gaussian angular weighting used during training data preprocessing was omitted as an auxiliary implementation detail.

References

  1. 1.Aanæs, H., Jensen, R.R., Vogiatzis, G., Tola, E., Dahl, A.B.: Large-scale data for multiple-view stereopsis. Int. J. Comput. Vis. (IJCV) 120, 153–168 (2016)
  2. 2.Abadi, M., et al.: TensorFlow: Large-scale machine learning on heterogeneous systems (2015). https://www.tensorflow.org/
  3. 3.Campbell, N.D.F., Vogiatzis, G., Hernández, C., Cipolla, R.: Using multiple hypotheses to improve depth-maps for multi-view stereo. In: Forsyth, D., Torr, P., Zisserman, A. (eds.) ECCV 2008. LNCS, vol. 5302, pp. 766–779. Springer, Heidelberg (2008). https://doi.org/10.1007/978-3-540-88682-2_58
  4. 4.Chen, L.C., Papandreou, G., Kokkinos, I., Murphy, K., Yuille, A.L.: DeepLab: semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected CRFs. IEEE Trans. Pattern Anal. Mach. Intell. (TPAMI) 40, 834–848 (2017)
  5. 5.Collins, R.T.: A space-sweep approach to true multi-image matching. In: Computer Vision and Pattern Recognition (CVPR) (1996)
  6. 6.Fuhrmann, S., Langguth, F., Goesele, M.: MVE-a multi-view reconstruction environment. In: Eurographics Workshop on Graphics and Cultural Heritage (GCH) (2014)
  7. 7.Furukawa, Y., Ponce, J.: Accurate, dense, and robust multiview stereopsis. IEEE Trans. Pattern Anal. Mach. Intell. (TPAMI) 32, 1362–1376 (2010)
  8. 8.Galliani, S., Lasinger, K., Schindler, K.: Massively parallel multiview stereopsis by surface normal diffusion. In: International Conference on Computer Vision (ICCV) (2015)
  9. 9.Geiger, A., Lenz, P., Urtasun, R.: Are we ready for autonomous driving? the KITTI vision benchmark suite. In: Computer Vision and Pattern Recognition (CVPR) (2012)
  10. 10.Han, X., Leung, T., Jia, Y., Sukthankar, R., Berg, A.C.: MatchNet: unifying feature and metric learning for patch-based matching. In: Computer Vision and Pattern Recognition (CVPR) (2015)
  11. 11.Hartmann, W., Galliani, S., Havlena, M., Van Gool, L., Schindler, K.: Learned multi-patch similarity. In: International Conference on Computer Vision (ICCV) (2017)
  12. 12.Hirschmuller, H.: Stereo processing by semiglobal matching and mutual information. IEEE Trans. Pattern Anal. Mach. Intell. (TPAMI) 30, 328–341 (2008)
  13. 13.Hirschmuller, H., Scharstein, D.: Evaluation of cost functions for stereo matching. In: Computer Vision and Pattern Recognition (CVPR) (2007)
  14. 14.Ji, M., Gall, J., Zheng, H., Liu, Y., Fang, L.: SurfaceNet: an end-to-end 3D neural network for multiview stereopsis. In: International Conference on Computer Vision (ICCV) (2017)
  15. 15.Kar, A., Häne, C., Malik, J.: Learning a multi-view stereo machine. In: Advances in Neural Information Processing Systems (NIPS) (2017)
  16. 16.Kazhdan, M., Hoppe, H.: Screened poisson surface reconstruction. ACM Trans. Graph. (TOG) 32, 29 (2013)
  17. 17.Kendall, A., Martirosyan, H., Dasgupta, S., Henry, P.: End-to-end learning of geometry and context for deep stereo regression. In: Computer Vision and Pattern Recognition (CVPR) (2017)
  18. 18.Knapitsch, A., Park, J., Zhou, Q.Y., Koltun, V.: Tanks and temples: benchmarking large-scale scene reconstruction. ACM Trans. Graph. (TOG) 36, 78 (2017)
  19. 19.Knöbelreiter, P., Reinbacher, C., Shekhovtsov, A., Pock, T.: End-to-end training of hybrid CNN-CRF models for stereo. In: Computer Vision and Pattern Recognition (CVPR) (2017)
  20. 20.Kutulakos, K.N., Seitz, S.M.: A theory of shape by space carving. Int. J. Comput. Vis. (IJCV) 38, 199–218 (2000)
  21. 21.Langguth, F., Sunkavalli, K., Hadap, S., Goesele, M.: Shading-aware multi-view stereo. In: Leibe, B., Matas, J., Sebe, N., Welling, M. (eds.) ECCV 2016. LNCS, vol. 9907, pp. 469–485. Springer, Cham (2016). https://doi.org/10.1007/978-3-319-46487-9_29
  22. 22.Lhuillier, M., Quan, L.: A quasi-dense approach to surface reconstruction from uncalibrated images. IEEE Trans. Pattern Anal. Mach. Intell. (TPAMI) 27, 418–433 (2005)
  23. 23.Luo, W., Schwing, A.G., Urtasun, R.: Efficient deep learning for stereo matching. In: Computer Vision and Pattern Recognition (CVPR) (2016)
  24. 24.Mayer, N., et al.: A large dataset to train convolutional networks for disparity, optical flow, and scene flow estimation. In: Computer Vision and Pattern Recognition (CVPR) (2016)
  25. 25.Menze, M., Geiger, A.: Object scene flow for autonomous vehicles. In: Computer Vision and Pattern Recognition (CVPR) (2015)
  26. 26.Merrell, P., et al.: Real-time visibility-based fusion of depth maps. In: International Conference on Computer Vision (ICCV) (2007)
  27. 27.Moulon, P., Monasse, P., Marlet, R., et al.: OpenMVG: an open multiple view geometry library. https://github.com/openMVG/openMVG
  28. 28.Newcombe, R.A., et al.: KinectFusion: real-time dense surface mapping and tracking. In: IEEE International Symposium on Mixed and Augmented Reality (ISMAR) (2011)
  29. 29.OpenMVS: open multi-view stereo reconstruction library. https://github.com/cdcseacave/openMVS
  30. 30.Pix4D: https://pix4d.com/
  31. 31.Ronneberger, O., Fischer, P., Brox, T.: U-Net: convolutional networks for biomedical image segmentation. In: Navab, N., Hornegger, J., Wells, W.M., Frangi, A.F. (eds.) MICCAI 2015. LNCS, vol. 9351, pp. 234–241. Springer, Cham (2015). https://doi.org/10.1007/978-3-319-24574-4_28
  32. 32.Schönberger, J.L., Zheng, E., Frahm, J.-M., Pollefeys, M.: Pixelwise view selection for unstructured multi-view stereo. In: Leibe, B., Matas, J., Sebe, N., Welling, M. (eds.) ECCV 2016. LNCS, vol. 9907, pp. 501–518. Springer, Cham (2016). https://doi.org/10.1007/978-3-319-46487-9_31
  33. 33.Seitz, S.M., Dyer, C.R.: Photorealistic scene reconstruction by voxel coloring. Int. J. Comput. Vis. (IJCV) 35, 151–173 (1999)
  34. 34.Seki, A., Pollefeys, M.: SGM-Nets: semi-global matching with neural networks. In: Computer Vision and Pattern Recognition Workshops (CVPRW) (2017)
  35. 35.Tola, E., Strecha, C., Fua, P.: Efficient large-scale multi-view stereo for ultra high-resolution image sets. In: Machine Vision and Applications (MVA) (2012)
  36. 36.Vu, H.H., Labatut, P., Pons, J.P., Keriven, R.: High accuracy and visibility-consistent dense multiview stereo. IEEE Trans. Pattern Anal. Mach. Intell. (TPAMI) 34, 889–901 (2012)
  37. 37.Xu, N., Price, B., Cohen, S., Huang, T.: Deep image matting. In: Computer Vision and Pattern Recognition (CVPR) (2017)
  38. 38.Yao, Y., Li, S., Zhu, S., Deng, H., Fang, T., Quan, L.: Relative camera refinement for accurate dense reconstruction. In: 3D Vision (3DV) (2017)
  39. 39.Zbontar, J., LeCun, Y.: Stereo matching by training a convolutional neural network to compare image patches. J. Mach. Learn. Res. (JMLR) 17, 2 (2016)
  40. 40.Zhang, R., Li, S., Fang, T., Zhu, S., Quan, L.: Joint camera clustering and surface segmentation for large-scale multi-view stereo. In: International Conference on Computer Vision (ICCV) (2015)

Citation

MLA
Yao, Y., et al. “MVSNet: Depth Inference for Unstructured Multi-view Stereo”. arXiv, 2018, http://arxiv.org/abs/1804.02505v2.
APA
Yao, Y., Luo, Z., Li, S., Fang, T., & Quan, L. (2018). MVSNet: Depth Inference for Unstructured Multi-view Stereo. arXiv. http://arxiv.org/abs/1804.02505v2
Chicago
Yao, Y., Z. Luo, S. Li, T. Fang, and L. Quan. 2018. “MVSNet: Depth Inference for Unstructured Multi-view Stereo”. arXiv. http://arxiv.org/abs/1804.02505v2.
Harvard
Yao, Y. et al. (2018) “MVSNet: Depth Inference for Unstructured Multi-view Stereo”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1804.02505v2.
Vancouver
1. Yao Y, Luo Z, Li S, Fang T, Quan L (2018) MVSNet: Depth Inference for Unstructured Multi-view Stereo. arXiv

BibTeX

@article{yao2018mvsnet,
  title = {MVSNet: Depth Inference for Unstructured Multi-view Stereo},
  author = {Yao, Yao and Luo, Zixin and Li, Shiwei and Fang, Tian and Quan, Long},
  year = {2018},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1804.02505v2},
  eprint = {1804.02505}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF