A Comparison and Evaluation of Multi-View Stereo Reconstruction Algorithms

S. SeitzB. CurlessJ. DiebelD. ScharsteinR. Szeliski

article2006CVPR2,913 citations

Establishes a standardized benchmark and evaluation methodology for multi-view stereo reconstruction by providing calibrated image datasets with high-precision laser-scanned ground truth alongside a comprehensive algorithmic taxonomy.

Listen

This paper addresses the absence of standardized, calibrated multi-view image datasets with accurate ground-truth 3D models, which had blocked direct quantitative comparisons among reconstruction algorithms and slowed progress in the field. Without such benchmarks, researchers could not reliably identify strengths, weaknesses, or priority areas for improvement, unlike the situation in binocular stereo where shared test data had already accelerated gains.

The work set out to create the first public collection of high-quality multi-view datasets registered to laser-scanned ground truth, together with a taxonomy of algorithm properties and an evaluation framework that measures both geometric accuracy and surface completeness.

The authors first surveyed existing techniques and organized them by scene representation, photo-consistency measure, visibility handling, shape priors, reconstruction strategy, and initialization needs. They then captured roughly 300360 calibrated images per object for two subjects using a precision robotic gantry, produced reference meshes via dense laser scanning, and refined the alignment of each mesh to the images by minimizing photo-consistency error. Six leading algorithms were run on full-hemisphere, ring, and sparse-ring subsets of the data; accuracy was reported as the distance within which 90 percent of reconstructed points lie from the ground truth, and completeness as the fraction of ground-truth points recovered within a 1.25 mm tolerance.

The evaluation shows that several methods reach sub-millimeter accuracy from standard video-resolution images, with the best result placing 90 percent of points within 0.36 mm on the full temple set. Accuracy remains high across view counts for textured objects but varies more for low-texture surfaces, where regularization influences outcomes. Completeness is generally strong when silhouettes are available, yet drops for methods that leave holes in uncertain regions. Offsets of several tenths of a millimeter between independently produced models required explicit alignment before comparison.

These results demonstrate that current multi-view stereo techniques can already deliver models accurate enough for many practical uses, while also revealing that performance still depends on scene texture, view density, and the use of silhouette constraints. The availability of common test data and repeatable metrics should therefore focus future effort on the remaining gaps, such as handling specular surfaces and operating without silhouettes.

The authors plan to release additional datasets with specularities and no silhouettes, to acquire higher-resolution imagery paired with industrial-grade ground truth, and to keep the evaluation open for new submissions. Researchers should treat the current numbers as a baseline rather than a final ranking, because the study covers only algorithms that supplied results by the original deadline and assumes largely Lambertian reflectance.

Cover for A Comparison and Evaluation of Multi-View Stereo Reconstruction Algorithms

Abstract

This paper presents a quantitative comparison of several multi-view stereo reconstruction algorithms. Until now, the lack of suitable calibrated multi-view image datasets with known ground truth (3D shape models) has prevented such direct comparisons. In this paper, we first survey multi-view stereo algorithms and compare them qualitatively using a taxonomy that differentiates their key properties. We then describe our process for acquiring and calibrating multi-view image datasets with high-accuracy ground truth and introduce our evaluation methodology. Finally, we present the results of our quantitative comparison of state-of-the-art multi-view stereo reconstruction algorithms on six benchmark datasets. The datasets, evaluation details, and instructions for submitting new models are available online at http://vision.middlebury.edu/mview.

Table of Contents

  • 1. Introduction
  • 2. A multi-view stereo taxonomy
  • 2.1. Scene representation
  • 2.2. Photo›consistency measure
  • 2.3. Visibility model
  • 2.4. Shape prior
  • 2.5. Reconstruction algorithm
  • 2.6. Initialization requirements
  • 3. Multi-view data sets
  • 4. Evaluation methodology
  • 5. Results
  • 6. Conclusions
  • References

Knowls

  1. Knowl 1 — Geometric Evaluation Metrics for Multi-View Stereo: Accuracy and Completeness

    model/method

    Multi-view stereo reconstructions are quantitatively evaluated against a high-accuracy ground-truth 3D model through two complementary metrics: accuracy (measuring surface proximity) and completeness (measuring surface coverage).

    Let GG denote the ground-truth surface mesh and RR denote the reconstructed triangle surface mesh to be evaluated. To avoid false error penalties in regions where GG is incomplete due to scanner occlusions, GG is augmented with hole-filled surface patches generated via volumetric space carving, producing an augmented mesh GG'. In addition, scanner sampling confidence weights c(v)c(v) are assigned to each vertex vGv \in G.

    Accuracy Metric: For each vertex rRr \in R, the closest point on GG' is determined. If that closest point falls on a hole-filled region or on a low-confidence vertex of GG, the vertex rr is excluded from the evaluation set, yielding a subset of valid vertices RvalidRR_{\text{valid}} \subseteq R. For each remaining vertex rRvalidr \in R_{\text{valid}}, the Euclidean distance mingGrg2\min_{g \in G} \|r - g\|_2 to the ground-truth surface GG is computed. Accuracy is reported as the distance threshold dd (in millimeters) such that a fraction XX (standardized at X=90%X = 90\%) of valid vertices on RR lie within distance dd of GG:

    Accuracy(R,G;X)=inf{dR0  |  {rRvalidmingGrg2d}RvalidX100}\text{Accuracy}(R, G; X) = \inf \left\{ d \in \mathbb{R}_{\ge 0} \;\middle|\; \frac{\left|\{r \in R_{\text{valid}} \mid \min_{g \in G} \|r - g\|_2 \le d\}\right|}{|R_{\text{valid}}|} \ge \frac{X}{100} \right\}

    Signed distance for a query vertex rRr \in R relative to its nearest surface point gGg \in G with outward unit normal ng\mathbf{n}_g is given by sgn(ng(rg))rg2\operatorname{sgn}(\mathbf{n}_g \cdot (r - g)) \|r - g\|_2.

    Completeness Metric: For each vertex gGg \in G, the Euclidean distance minrRgr2\min_{r \in R} \|g - r\|_2 to the nearest point on the reconstruction RR is evaluated. For a fixed inlier distance threshold dinlierd_{\text{inlier}} (standardized to 1.25 mm1.25\text{ mm}), completeness is defined as the percentage of ground-truth vertices gGg \in G that lie within dinlierd_{\text{inlier}} of RR:

    Completeness(G,R;dinlier)={gGminrRgr2dinlier}G×100%\text{Completeness}(G, R; d_{\text{inlier}}) = \frac{\left|\{g \in G \mid \min_{r \in R} \|g - r\|_2 \le d_{\text{inlier}}\}\right|}{|G|} \times 100\%

    Ground-truth vertices whose projection onto RR falls on an open mesh boundary or beyond dinlierd_{\text{inlier}} are marked as uncovered.

  2. Knowl 2 — Photo-Consistency-Based Alignment of Ground-Truth 3D Meshes to Multi-View Images

    algorithm

    To register laser-scanned ground-truth 3D meshes to calibrated multi-view image sets, a non-linear optimization estimates rigid transformation parameters and uniform scale T=(Rrot,t,s)\mathbf{T} = (\mathbf{R}_{\text{rot}}, \mathbf{t}, s) by minimizing photo-consistency variance across all views where mesh vertices are visible.

    Input: Ground-truth mesh G=(VG,FG)G = (V_G, F_G), calibrated multi-view images {Ik}k=1K\{I_k\}_{k=1}^K with projection matrices {πk}k=1K\{\pi_k\}_{k=1}^K, initial reconstruction RinitR_{\text{init}}, vertex confidences {c(v)}vVG\{c(v)\}_{v \in V_G}
    Output: Transformation parameters T=(Rrot,t,s)\mathbf{T}^* = (\mathbf{R}_{\text{rot}}^*, \mathbf{t}^*, s^*) registering GG to the camera coordinate frame
    Initialize TICP(G,Rinit)\mathbf{T} \leftarrow \text{ICP}(G, R_{\text{init}})
    step_size \leftarrow \text{initial_step_size}
    Define objective function E(T)\mathcal{E}(\mathbf{T}):
      E(T)=vVGc(v)V(v,T)Var({Ik(πk(T(v)))kV(v,T)})\mathcal{E}(\mathbf{T}) = \sum_{v \in V_G} c(v) \cdot |\mathcal{V}(v, \mathbf{T})| \cdot \operatorname{Var}(\{ I_k(\pi_k(\mathbf{T}(v))) \mid k \in \mathcal{V}(v, \mathbf{T}) \})
      where V(v,T)\mathcal{V}(v, \mathbf{T}) is the set of camera indices where transformed vertex T(v)\mathbf{T}(v) is visible
    while step_size > \text{convergence_threshold} do
        for each transformation parameter θT\theta \in \mathbf{T} do
            Compute finite difference Newton step along coordinate θ\theta
            TcandT\mathbf{T}_{\text{cand}} \leftarrow \mathbf{T} with bounded step along θ\theta
            if E(Tcand)<E(T)\mathcal{E}(\mathbf{T}_{\text{cand}}) < \mathcal{E}(\mathbf{T}) then
                TTcand\mathbf{T} \leftarrow \mathbf{T}_{\text{cand}}
            else
                step_size \leftarrow \text{step_size} / 2
    return T\mathbf{T}

    Initialization via Iterative Closest Point (ICP) against an arbitrary candidate multi-view reconstruction yields consistent convergence regardless of which initial reconstruction is selected. Final alignment quality is verified through visual reprojection inspection, yielding maximum reprojection errors 1 pixel\le 1\text{ pixel}.

  3. Knowl 3 — Middlebury Multi-View Stereo Benchmark Dataset Acquisition Setup

    experimental setup

    The Multi-View Stereo benchmark datasets provide calibrated multi-view imagery paired with registered sub-millimeter ground-truth 3D meshes.

    • Image Acquisition Platform: A robotic Stanford spherical gantry positioning a 640×480640 \times 480 pixel CCD camera along a 1-meter radius sphere with approximately 0.010.01^\circ angular positioning precision. A camera pixel projects to approximately 0.25 mm0.25\text{ mm} on the target object surface.
    • Camera Calibration: 68 views of a planar calibration grid are captured over the hemisphere to estimate intrinsic and extrinsic parameters and determine the rigid 6-DOF offset relative to the gantry arm tip.
    • Target Objects:
      1. Temple: Physical dimensions 10 cm×16 cm×8 cm10\text{ cm} \times 16\text{ cm} \times 8\text{ cm}, characterized by sharp polygonal edges, complex topology, strong concavities, and high surface texture.
      2. Dino: Physical dimensions 7 cm×9 cm×7 cm7\text{ cm} \times 9\text{ cm} \times 7\text{ cm}, characterized by smooth organic curvature and low surface texture with subtle shading variations.
    • Image Subsets: Objects are illuminated by three stationary spotlights. Gantry shadow artifacts are eliminated by double-covering the sphere from two arm configurations, yielding ~80% hemispherical coverage divided into three benchmark splits per object:
      • Full hemisphere: 317 views for Temple; 363 views for Dino.
      • Ring: 47 views for Temple; 48 views for Dino (single equatorial orbit).
      • Sparse Ring: 16 views for Temple; 16 views for Dino (regularly subsampled equatorial orbit).
    • Ground-Truth 3D Mesh Generation: Scanned using a Cyberware Model 15 laser stripe scanner (0.25 mm0.25\text{ mm} scan resolution, 0.050.2 mm0.05\text{--}0.2\text{ mm} measurement accuracy). For each object, approximately 200 individual range scans are aligned and merged onto a 0.25 mm0.25\text{ mm} grid using volumetric range image processing (VRIP) with sub-voxel isosurface extraction.
  4. Knowl 4 — Six-Axis Categorization Taxonomy for Multi-View Stereo Algorithms

    model/method

    Dense multi-view stereo reconstruction algorithms are categorized along six fundamental technical axes:

    1. Scene Representation:
      • 3D Volumes / Grids: Regular discrete voxel occupancy grids or signed distance functions encoded via level-set formulations.
      • Polygon Meshes: Explicit connected planar triangular facets, enabling direct surface deformation and visibility queries.
      • Depth Maps: Multi-view sets of 2D depth maps or relief surfaces defined relative to base geometric proxies.
    2. Photo-Consistency Measure:
      • Scene-Space Measures: Evaluated by projecting 3D points or patches into multiple images to calculate pixel color variance, sum of squared differences (SSD), or normalized cross-correlation (NCC). Integration is performed over 3D surface area (inherently favoring smaller surface areas).
      • Image-Space Measures: Warping reference images to target viewpoints to evaluate photometric prediction error. Integration occurs over 2D image planes (weighting frequently observed or large-projection regions higher).
      • Radiometric Models: Lambertian diffuse reflectance versus non-Lambertian BRDF modeling, silhouette contours, or shadow constraints.
    3. Visibility Model:
      • Geometric: Explicit occlusion determination using the current 3D surface estimate (e.g., conservative carving where visible camera sets monotonically expand).
      • Quasi-Geometric: View clustering heuristics (restricting matching to nearby camera baselines) or visual hull approximations.
      • Outlier-Based: Treating occluded views as statistical outliers without computing explicit 3D ray intersections.
    4. Shape Prior:
      • Minimal Surface Priors: Favoring minimal surface area (common in level-set PDEs, volumetric graph cuts, and mesh Laplacian smoothing); tends to smooth high-curvature features.
      • Maximal Surface Priors: Preserving the maximal photo-consistent volume (the photo hull via space carving); preserves thin structures but bulges in low-texture regions.
      • Smoothness Priors: Enforcing 2D MRF pairwise depth smoothness (with fronto-parallel bias) or 3D surface curvature continuity.
    5. Reconstruction Algorithm:
      • Volumetric cost optimization: Constructing a 3D cost volume followed by optimal surface extraction (voxel coloring, volumetric MRF max-flow / graph cuts).
      • Iterative surface evolution: Deforming an initial surface to minimize an energy functional (space carving, level sets, deformable meshes).
      • Multi-depth-map fusion: Computing independent 2D depth maps with cross-view consistency constraints followed by 3D volumetric merging (VRIP).
      • Feature matching and fitting: Reconstructing sparse feature points followed by surface mesh fitting.
    6. Initialization Requirements:
      • Bounding boxes, foreground/background silhouettes (visual hulls), or disparity/depth search intervals [znear,zfar][z_{\text{near}}, z_{\text{far}}].
  5. Knowl 5 — Benchmark Quantitative Results across Multi-View Stereo Algorithms

    data/table

    Quantitative evaluation of six state-of-the-art multi-view stereo algorithms on the Temple and Dino benchmark datasets across Full, Ring, and Sparse Ring camera configurations.

    Accuracy is defined as the distance threshold dd (in millimeters) within which 90%90\% of the reconstructed mesh vertices lie relative to the ground-truth mesh (lower is better). Completeness is defined as the percentage of ground-truth mesh vertices lying within 1.25 mm1.25\text{ mm} of the reconstructed mesh (higher is better). Ground-truth meshes were rigidly aligned to each submitted reconstruction via ICP prior to metric evaluation.

    Full Ring Sparse Ring
    Algorithm Acc. (mm) Comp. (%) Acc. (mm) Comp. (%) Acc. (mm) Comp. (%)
    Temple Dataset
    Furukawa 0.65 98.7% 0.58 98.5% 0.82 94.3%
    Goesele 0.42 98.0% 0.61 86.2% 0.87 56.6%
    Hernandez 0.36 99.7% 0.52 99.5% 0.75 95.3%
    Kolmogorov 1.86 90.4%
    Pons 0.60 99.5% 0.90 95.4%
    Vogiatzis 1.07 90.7% 0.76 96.2% 2.77 79.4%
    Dino Dataset
    Furukawa 0.52 99.2% 0.42 98.8% 0.58 96.9%
    Goesele 0.56 80.0% 0.46 57.8% 0.56 26.0%
    Hernandez 0.49 99.6% 0.45 97.9% 0.60 98.5%
    Kolmogorov 2.80 85.7%
    Pons 0.55 99.0% 0.71 97.7%
    Vogiatzis 0.42 99.0% 0.49 96.7% 1.18 90.8%

    The benchmark demonstrates that state-of-the-art algorithms consistently achieve sub-millimeter geometric accuracy (<0.5 mm< 0.5\text{ mm}) from standard video-resolution (640×480640 \times 480) imagery where pixel ground resolution is 0.25 mm\approx 0.25\text{ mm}. Hernandez achieves the highest accuracy on the textured Temple datasets (0.36 mm0.36\text{ mm} on Full, 0.52 mm0.52\text{ mm} on Ring). Goesele maintains high local accuracy but produces lower completeness on sparse views (56.6%56.6\% on Temple Sparse Ring) and weakly textured surfaces (26.0%26.0\% on Dino Sparse Ring) because low-confidence patches are left unconstructed.

  6. Knowl 6 — Empirical Performance and Methodological Trade-offs in Multi-View Stereo

    empirical result

    Quantitative benchmarking of dense multi-view stereo algorithms on controlled datasets reveals several behavioral properties across varying scene conditions:

    • Texture vs. View Density Dynamics: On strongly textured objects (Temple), reconstruction accuracy improves monotonically with increasing view count (Full > Ring > Sparse Ring). On weakly textured objects with subtle shading (Dino), algorithms counterintuitively achieve better accuracy on Ring (48 views) than on Full (363 views), because excessive dense view comparisons in textureless areas introduce photometric matching noise, whereas fewer views allow shape regularization priors to effectively smooth the surface.
    • Completeness vs. Accuracy Conservatism: Methods using confidence-gated depth map fusion (such as Goesele) prioritize high precision in reliable regions and reject ambiguous matches, leading to sharp completeness drops on low-texture or sparsely sampled datasets. In contrast, global energy-minimization methods (Hernandez, Furukawa, Pons, Vogiatzis) interpolate over uncertain regions, preserving completeness >90%> 90\% across most configurations.
    • Silhouette Reliance: Algorithms requiring silhouette constraints (Hernandez, Vogiatzis, Furukawa) perform well when background segmentation is clean. However, silhouette-free formulations (such as Pons) achieve comparable sub-millimeter accuracy (0.60 mm0.60\text{ mm} on Temple Ring, 0.55 mm0.55\text{ mm} on Dino Ring) while remaining applicable to unsegmented scenes.
    • Computational Runtime: Computational cost varies by more than an order of magnitude; the level-set prediction-error method of Pons was the fastest (31 minutes on Temple Ring), whereas the multi-depth-map fusion method of Goesele was the slowest (exceeding 24 hours on Temple Ring).
  7. Knowl 7 — Sub-Millimeter Translational Offsets and Pre-Evaluation ICP Compensation

    model/method

    Independent multi-view stereo reconstructions exhibit small, systematic translational offsets (up to 0.6 mm0.6\text{ mm}) relative to one another and to the laser-scanned ground-truth mesh. On high-accuracy datasets where algorithms attain errors on the order of 0.360.50 mm0.36\text{--}0.50\text{ mm}, uncompensated translational shifts distort geometric benchmarking comparisons.

    These global shifts arise from minute physical calibration tolerances (such as gantry arm flexure variations across orbital latitudes) and algorithm-specific formulation biases (e.g., visual hull shrinkage versus smoothness regularization offsets). To evaluate intrinsic shape reconstruction fidelity rather than global coordinate frame calibration discrepancies, each reconstruction RR is rigidly registered to the ground-truth mesh GG using Iterative Closest Point (ICP) before computing accuracy and completeness metrics.

Coverage note — None was omitted; all contributed taxonomy axes, data acquisition procedures, alignment algorithms, evaluation metrics, and benchmark results are fully represented.

References

  1. 1.D. Scharstein and R. Szeliski. A taxonomy and evaluation of dense two-frame stereo correspondence algorithms. IJCV, 47(1):7–42, 2002.
  2. 2.S. Seitz et al. Multi-view stereo evaluation web page. http://vision.middlebury.edu/mview/.
  3. 3.C. Dyer. Volumetric scene reconstruction from multiple views. In L. S. Davis, editor, Foundations of Image Understanding, pp. 469–489. Kluwer, 2001.
  4. 4.G. Slabaugh, B. Culbertson, T. Malzbender, and R. Shafer. A survey of methods for volumetric scene reconstruction from photographs. In Intl. WS on Volume Graphics, 2001.
  5. 5.T. Fromherz and M. Bichsel. Shape from multiple cues: Integrating local brightness information. In ICYCS, 1995.
  6. 6.S. Roy and I. Cox. A maximum-flow formulation of the N-camera stereo correspondence problem. In ICCV, pp. 492–499, 1998.
  7. 7.R. Szeliski and P. Golland. Stereo matching with transparency and matting. IJCV, 32(1):45–61, 1999.
  8. 8.S. Seitz and C. Dyer. Photorealistic scene reconstruction by voxel coloring. IJCV, 35(2):151–173, 1999.
  9. 9.P. Eisert, E. Steinbach, and B. Girod. Multi-hypothesis, volumetric reconstruction of 3-D objects from multiple calibrated camera views. In ICASSP 99, pp. 3509–3512, 1999.
  10. 10.J. De Bonet and P. Viola. Poxels: Probabilistic voxelized volume reconstruction. In ICCV, pp. 418–425, 1999.
  11. 11.K. Kutulakos and S. Seitz. A theory of shape by space carving. IJCV, 38(3):199–218, 2000.
  12. 12.K. Kutulakos. Approximate N-view stereo. In ECCV, vol. I, pp. 67–83, 2000.
  13. 13.A. Broadhurst, T. Drummond, and R. Cipolla. A probabilistic framework for the space carving algorithm. In ICCV, pp. 388–393, 2001.
  14. 14.R. Bhotika, D. Fleet, and K. Kutulakos. A probabilistic theory of occupancy and emptiness. In ECCV, vol. 3, pp. 112–132, 2002.
  15. 15.R. Yang, M. Pollefeys, and G. Welch. Dealing with textureless regions and specular highlights – a progressive space carving scheme using a novel photo-consistency measure. In ICCV, pp. 576–584, 2003.
  16. 16.T. Bonfort and P. Sturm. Voxel carving for specular surfaces. In ICCV, pp. 591–596, 2003.
  17. 17.A. Treuille, A. Hertzmann, and S. Seitz. Example-based stereo with general BRDFs. In ECCV, vol. II, pp. 457–469, 2004.
  18. 18.G. Slabaugh, B. Culbertson, T. Malzbender, and M. Stevens. Methods for volumetric reconstruction of visual scenes. IJCV, 57(3):179–199, 2004.
  19. 19.G. Vogiatzis, P. Torr, and R. Cipolla. Multi-view stereo via volumetric graph-cuts. In CVPR, pp. 391–398, 2005.
  20. 20.O. Faugeras and R. Keriven. Variational principles, surface evolution, PDE’s, level set methods and the stereo problem. IEEE Trans. on Image Processing, 7(3):336–344, 1998.
  21. 21.J.-P. Pons, R. Keriven, O. Faugeras, and G. Hermosillo. Variational stereovision and 3D scene flow estimation with statistical similarity measures. In ICCV, pp. 597–602, 2003.
  22. 22.S. Soatto, A. Yezzi, and H. Jin. Tales of shape and radiance in multiview stereo. In ICCV, pp. 974–981, 2003.
  23. 23.H. Jin, S. Soatto, and A. Yezzi. Multi-view stereo beyond lambert. In CVPR, vol. 1, pp. 171–178, 2003.
  24. 24.Y. Duan, L. Yang, H. Qin, and D. Samaras. Shape reconstruction from 3D and 2D data using PDE-based deformable surfaces. In ECCV, vol. 3, pp. 238–251, 2004.
  25. 25.H. Jin, S. Soatto, and A. Yezzi. Multi-view stereo reconstruction of dense shape and complex appearance. IJCV, 63(3):175–189, 2005.
  26. 26.J.-P. Pons, R. Keriven, and O. Faugeras. Modelling dynamic scenes by registering multi-view image sequences. In CVPR, vol. II, pp. 822–827, 2005.
  27. 27.P. Fua and Y. Leclerc. Object-centered surface reconstruction: Combining multi-image stereo and shading. IJCV, 16:35–56, 1995.
  28. 28.A. Rockwood and J. Winget. Three-dimensional object reconstruction from two-dimensional images. Computer-Aided Design, 29(4):279–285, 1997.
  29. 29.L. Zhang and S. Seitz. Image-based multiresolution shape recovery by surface deformation. In SPIE: Videometrics and Optical Methods for 3D Shape Measurement, pp. 51–61, 2001.
  30. 30.J. Isidoro and S. Sclaroff. Stochastic refinement of the visual hull to satisfy photometric and silhouette consistency constraints. In ICCV, pp. 1335–1342, 2003.
  31. 31.C. Hernandez and F. Schmitt. Silhouette and stereo fusion for 3D object modeling. CVIU, 96(3):367–392, 2004.
  32. 32.T. Yu, N. Xu, and N. Ahuja. Shape and view independent reflectance map from multiple views. In ECCV, pp. 602–616, 2004.
  33. 33.R. Szeliski. A multi-view approach to motion and stereo. In CVPR, vol. 1, pp. 157–163, 1999.
  34. 34.S.-B. Kang, R. Szeliski, and J. Chai. Handling occlusions in dense multi-view stereo. In CVPR, vol. I, pp. 103–110, 2001.
  35. 35.V. Kolmogorov and R. Zabih. Multi-camera scene reconstruction via graph cuts. In ECCV, vol. III, pp. 82–96, 2002.
  36. 36.C. Zitnick, S.-B. Kang, M. Uyttendaele, S. Winder, and R. Szeliski. High-quality video view interpolation using a layered representation. ACM Trans. on Graphics, 23(3):600–608, 2004.
  37. 37.P. Gargallo and P. Sturm. Bayesian 3D modeling from images using multiple depth maps. In CVPR, vol. II, pp. 885–891, 2005.
  38. 38.M.-A. Drouin, M. Trudeau, and S. Roy. Geo-consistency for wide multi-camera stereo. In CVPR, vol. I, pp. 351–358, 2005.
  39. 39.G. Vogiatzis, P. Torr, S. M. Seitz, and R. Cipolla. Reconstructing relief surfaces. In BMVC, pp. 117–126, 2004.
  40. 40.G. Zeng, S. Paris, L. Quan, and F. Sillion. Progressive surface reconstruction from images using a local prior. In ICCV, pp. 1230–1237, 2005.
  41. 41.R. Szeliski. Prediction error as a quality metric for motion and stereo. In ICCV, pp. 781–788, 1999.
  42. 42.S. Savarese, H. Rushmeier, F. Bernardini, and P. Perona. Shadow carving. In ICCV, pp. 190–197, 2001.
  43. 43.M. Okutomi and T. Kanade. A multiple-baseline stereo. TPAMI, 15(4):353–363, 1993.
  44. 44.A. Prock and C. Dyer. Towards real-time voxel coloring. In Image Understanding WS, pp. 315–321, 1998.
  45. 45.P. Narayanan, P. Rander, and T. Kanade. Constructing virtual worlds using dense stereo. In ICCV, pp. 3–10, 1998.
  46. 46.A. Laurentini. The visual hull concept for silhouette-based image understanding. TPAMI, 16(2):150–162, 1994.
  47. 47.S. Sinha and M. Pollefeys. Multi-view reconstruction using photo-consistency and exact silhouette constraints: A maximum-flow formulation. In ICCV, pp. 349–356, 2005.
  48. 48.Y. Furukawa and J. Ponce. High-fidelity image-based modeling. Technical Report 2006-02, UIUC, 2006.
  49. 49.C. Stewart. Robust parameter estimation in computer vision. SIAM Reviews, 41(3):513–537, 1999.
  50. 50.S. Baker, T. Sim, and T. Kanade. When is the shape of a scene unique given its light-field: A fundamental theorem of 3D vision? TPAMI, 25(1):100–109, 2003.
  51. 51.T. Tasdizen and R. Whitaker. Higher-order nonlinear priors for surface reconstruction. TPAMI, 26(7):878–891, 2004.
  52. 52.J. Diebel, S. Thrun, and M. Bruenig. A Bayesian method for probable surface reconstruction and decimation. ACM Trans. on Graphics, 25(1), 2006.
  53. 53.H. Saito and T. Kanade. Shape reconstruction in projective grid space from large number of images. In CVPR, vol. 2, pp. 49–54, 1999.
  54. 54.G. Slabaugh, T. Malzbender, B. Culbertson, and R. Schafer. Improved voxel coloring via volumetric optimization. TR 3, Center for Signal and Image Processing, 2000.
  55. 55.O. Faugeras, E. Bras-Mehlman, and J.-D. Boissonnat. Representing stereo data with the Delaunay triangulation. Artificial Intelligence, 44(1–2):41–87, 1990.
  56. 56.A. Manessis, A. Hilton, P. Palmer, P. McLauchlan, and X. Shen. Reconstruction of scene models from sparse 3D structure. In CVPR, vol. 1, pp. 666–673, 2000.
  57. 57.D. Morris and T. Kanade. Image-consistent surface triangulation. In CVPR, vol. 1, pp. 332–338, 2000.
  58. 58.C. J. Taylor. Surface reconstruction from feature based stereo. In ICCV, pp. 184–190, 2003.
  59. 59.D. Wood, D. Azuma, K. Aldinger, B. Curless, T. Duchamp, D. Salesin, and W. Stuetzle. Surface light fields for 3D photography. In SIGGRAPH, pp. 287–296, 1996.
  60. 60.W.-C. Chen, J.-Y. Bouguet, M. Chu, and R. Grzeszczuk. Light field mapping: Efficient representation and hardware rendering of surface light fields. ACM Trans. Graphics, 21(3):447–456, 2002.
  61. 61.J.-Y. Bouguet. Camera calibration toolbox for Matlab. http://www.vision.caltech.edu/bouguetj/calib doc/.
  62. 62.B. Curless and M. Levoy. A volumetric method for building complex models from range images. In SIGGRAPH, pp. 303–312, 1996.
  63. 63.M. Goesele, B. Curless, and S. Seitz. Multi-view stereo revisited. In CVPR, 2006. To appear.

Citation

MLA
Seitz, S. M., et al. “A Comparison and Evaluation of Multi-View Stereo Reconstruction Algorithms”. 2006 IEEE Computer Society Conference on Computer Vision and Pattern Recognition - Volume 1 (CVPR'06), vol. 1, 2006, pp. 519–28, https://doi.org/10.1109/CVPR.2006.19.
APA
Seitz, S. M., Curless, B., Diebel, J., Scharstein, D., & Szeliski, R. (2006). A Comparison and Evaluation of Multi-View Stereo Reconstruction Algorithms. 2006 IEEE Computer Society Conference on Computer Vision and Pattern Recognition - Volume 1 (CVPR'06), 1, 519–528. https://doi.org/10.1109/CVPR.2006.19
Chicago
Seitz, S. M., B. Curless, J. Diebel, D. Scharstein, and R. Szeliski. 2006. “A Comparison and Evaluation of Multi-View Stereo Reconstruction Algorithms”. 2006 IEEE Computer Society Conference on Computer Vision and Pattern Recognition - Volume 1 (CVPR'06) 1: 519–28. https://doi.org/10.1109/CVPR.2006.19.
Harvard
Seitz, S.M. et al. (2006) “A Comparison and Evaluation of Multi-View Stereo Reconstruction Algorithms”, 2006 IEEE Computer Society Conference on Computer Vision and Pattern Recognition - Volume 1 (CVPR'06). IEEE, pp. 519–528. Available at: https://doi.org/10.1109/CVPR.2006.19.
Vancouver
1. Seitz SM, Curless B, Diebel J, Scharstein D, Szeliski R (2006) A Comparison and Evaluation of Multi-View Stereo Reconstruction Algorithms. In: 2006 IEEE Computer Society Conference on Computer Vision and Pattern Recognition - Volume 1 (CVPR'06). IEEE, pp 519–528

BibTeX

@inproceedings{Seitz, title={A Comparison and Evaluation of Multi-View Stereo Reconstruction Algorithms}, volume={1}, url={http://dx.doi.org/10.1109/CVPR.2006.19}, DOI={10.1109/cvpr.2006.19}, booktitle={2006 IEEE Computer Society Conference on Computer Vision and Pattern Recognition - Volume 1 (CVPR′06)}, publisher={IEEE}, author={Seitz, S.M. and Curless, B. and Diebel, J. and Scharstein, D. and Szeliski, R.}, pages={519–528} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF

License: IEEE