Modeling and rendering architecture from photographs: a hybrid geometry- and image-based approach

Paul E. DebevecCamillo J. TaylorJitendra Malik

article1996SIGGRAPH2,259 citations

Presents a hybrid modeling framework that reconstructs photorealistic 3D architectural scenes from a sparse set of photographs by combining interactive block-based photogrammetry, model-based stereo, and view-dependent texture mapping.

Listen

Digital 3D modeling of real-world architecture is vital for virtual exploration and visual simulation, but conventional methods face significant practical barriers. Traditional geometry-based computer-aided design requires tedious manual labor, expensive surveying equipment, or difficult-to-verify floor plans, yet still yields synthetic-looking results. Conversely, automated image-based systems rely heavily on computational stereo algorithms that fail when camera viewpoints are far apart, demanding thousands of closely spaced photographs and intensive manual corrections to cover even a single city block.

The article demonstrates a hybrid geometry- and image-based modeling and rendering framework that reconstructs photorealistic 3D architectural scenes from a sparse set of still photographs. By dividing the workload between an interactive human interface and robust computer algorithms, the system evaluates how geometric constraints and approximate structural models can overcome standard limitations in stereo correspondence and image rendering.

The approach consists of three integrated components. First, users assemble an approximate, hierarchical block model of a building in an interactive program called Façade, marking visible edges that are automatically matched to photograph lines to solve for camera and model parameters. Second, a view-dependent texture-mapping algorithm composites projected images onto this base model, dynamically weighting source images based on the virtual viewing angle. Third, an automated model-based stereo algorithm projects widely spaced image pairs onto the basic model to remove distortion and compute fine geometric deviations, such as ornate carvings and facades.

Key findings show that architectural geometric constraints reduce the mathematical parameters needed for reconstruction by orders of magnitude—lowering a representative tower model from nearly 3,000 independent line parameters to just 33 variables. As a result, non-linear optimization converges in fewer than ten iterations with projected model edges matching photograph features within sub-pixel accuracy. Additionally, projecting widely spaced photographs onto the base geometry eliminates severe foreshortening differences, allowing stereo algorithms to accurately compute depth from images taken far apart without dense sampling. Entire building exteriors were fully modeled and rendered into 360-degree virtual fly-arounds using as few as 12 standard photographs in roughly four hours of modeling time.

These results demonstrate that organizations can generate high-fidelity virtual environments rapidly without specialized scanning instruments, CAD blueprints, or intrusive physical access. This significantly reduces data capture costs, operational timelines, and computational overhead for large-scale architectural simulation. Next steps supported by the article include automating the recovery of curved architectural surfaces (such as domes), developing algorithms to estimate surface material reflectances across varying lighting conditions, and integrating these models into real-time rendering pipelines.

While highly effective for polyhedral structures in consistent daylight, the framework's primary limitations include manual intervention for curved geometry, sensitivity to harsh lighting variations across photographs, and occasional visual seam artifacts during image blending. Confidence in the underlying method is high, as the recovered camera and structural parameters routinely align to within one pixel of source imagery.

Cover for Modeling and rendering architecture from photographs: a hybrid geometry- and image-based approach

Abstract

We present a new approach for modeling and rendering existing architectural scenes from a sparse set of still photographs. Our modeling approach, which combines both geometry-based and image-based techniques, has two components. The first component is a photogrammetric modeling method which facilitates the recovery of the basic geometry of the photographed scene. Our photogrammetric modeling approach is effective, convenient, and robust because it exploits the constraints that are characteristic of architectural scenes. The second component is a model-based stereo algorithm, which recovers how the real scene deviates from the basic model. By making use of the model, our stereo technique robustly recovers accurate depth from widely-spaced image pairs. Consequently, our approach can model large architectural environments with far fewer photographs than current image-based modeling approaches. For producing renderings, we present view-dependent texture mapping, a method of compositing multiple views of a scene that better simulates geometric detail on basic models. Our approach can be used to recover models for use in either geometry-based or image-based rendering systems. We present results that demonstrate our approach’s ability to create realistic renderings of architectural scenes from viewpoints far from the original photographs.

Table of Contents

  • 1 INTRODUCTION
  • 1.1 Background and Related Work
  • 1.1.1 Camera Calibration
  • 1.1.2 Structure from Motion
  • 1.1.3 Stereo Correspondence
  • 1.1.4 Image-Based Rendering
  • 1.2 Overview
  • 2 Photogrammetric Modeling
  • 2.1 The User's View
  • 2.2 Model Representation
  • 2.4 Computing an Initial Estimate
  • 2.5 Results
  • 3 View-Dependent Texture-Mapping
  • 3.1 Projecting a Single Image
  • 3.2 Compositing Multiple Images
  • 4 Model-Based Stereopsis
  • 4.1 Model-Based Epipolar Geometry
  • 4.2 Stereo Results and Rerendering
  • 5 Conclusion and Future Work
  • Acknowledgments
  • References

Knowls

  1. Knowl 1 — Model-Based Stereopsis via Warped Offset Images

    model/method

    Model-based stereopsis recovers fine geometric surface relief and depth deviations from wide-baseline image pairs by utilizing an approximate 3D geometric proxy model recovered through photogrammetry.

    Given a key image and an offset image taken from widely separated viewpoints, conventional correlation-based stereo fails due to significant differences in perspective foreshortening, disparity range, and occlusion. In model-based stereopsis, the offset image is projected onto the approximate geometric proxy model and then reprojected onto the image plane of the key camera. This generated image is termed the warped offset image.

    Matching is performed by comparing pixel correlation neighborhoods between the key image and the warped offset image rather than the raw offset image. The warped offset representation provides several key mathematical properties:

    • Points where the true scene coincides with the approximate model have exactly zero disparity.
    • Unequal perspective foreshortening between wide-baseline views is eliminated across surfaces modeled by the proxy.
    • Regions where the proxy model occludes itself relative to the key camera viewpoint can be predicted and masked prior to matching.
    • Disparities computed between the key and warped offset image map directly to true scene depth relative to the key camera.
  2. Knowl 2 — Epipolar Line Coincidence Under Model Proxy Warping

    theoretical result

    In model-based stereopsis, projecting an offset image onto a 3D geometric proxy model and reprojecting it into the key camera view preserves a linear epipolar search geometry that is identical to the key camera's original epipolar lines.

    Let PP be a point in the 3D scene, and let CkC_k and CoC_o be the optical centers of the key and offset cameras, respectively. The unique epipolar plane passes through PP, CkC_k, and CoC_o, intersecting the key and offset image planes in linear epipolar lines eke_k and eoe_o. The projection of PP into the offset image plane is pop_o. When pop_o is backprojected from CoC_o onto the proxy model surface, it intersects the proxy at point QQ, which lies strictly within the same epipolar plane. Reprojecting QQ from CkC_k onto the key camera's image plane produces a point qkq_k in the warped offset image.

    Because CkC_k, PP, QQ, and CoC_o are coplanar within the epipolar plane, the point qkq_k necessarily lies on the key image epipolar line eke_k. Thus, any correspondence for a key image feature pkp_k along eke_k remains strictly restricted to the identical 1D line eke_k in the warped offset image, preserving the 1D search space and enabling standard dynamic programming ordering constraints despite non-uniform geometric image warping.

  3. Knowl 3 — Photogrammetric Edge Disparity Objective Function

    equation

    The geometric error ErriErr_i between an observed 2D edge segment marked in an image and the corresponding 3D model line segment projected onto the image plane is defined as the integral of the squared perpendicular distance from points on the observed segment to the projected line.

    Let a 3D line be defined by a direction vector v\mathbf{v} and a point d\mathbf{d}. For a camera with rotation matrix Rj\mathbf{R}_j, translation vector tj\mathbf{t}_j, and focal length ff, the normal m=(mx,my,mz)T\mathbf{m} = (m_x, m_y, m_z)^T to the projection plane passing through the camera center and the line is given by: m=Rj(v×(d−tj))\mathbf{m} = \mathbf{R}_j (\mathbf{v} \times (\mathbf{d} - \mathbf{t}_j)) The projected line on the image plane located at z=−fz = -f is described by mxx+myy−mzf=0m_x x + m_y y - m_z f = 0.

    Let the observed image edge segment have endpoints (x1,y1)(x_1, y_1) and (x2,y2)(x_2, y_2) and length ll. Parameterizing points along the segment by s∈[0,l]s \in [0, l] with shortest distance h(s)h(s) to the projected line, the total disparity error is: Erri=∫0lh2(s) ds=l3(h12+h1h2+h22)=mT(ATBA)mErr_i = \int_0^l h^2(s)\,ds = \frac{l}{3}\left(h_1^2 + h_1 h_2 + h_2^2\right) = \mathbf{m}^T (\mathbf{A}^T \mathbf{B} \mathbf{A}) \mathbf{m} where h1h_1 and h2h_2 are the perpendicular distances from (x1,y1)(x_1, y_1) and (x2,y2)(x_2, y_2) to the projected line, and matrices A\mathbf{A} and B\mathbf{B} are defined as: A=(x1y11x2y21),B=l3(mx2+my2)(10.50.51)\mathbf{A} = \begin{pmatrix} x_1 & y_1 & 1 \\ x_2 & y_2 & 1 \end{pmatrix}, \quad \mathbf{B} = \frac{l}{3(m_x^2 + m_y^2)} \begin{pmatrix} 1 & 0.5 \\ 0.5 & 1 \end{pmatrix} The total objective function minimized across all correspondences is O=∑iErri\mathcal{O} = \sum_i Err_i.

  4. Knowl 4 — Decoupled Two-Stage Parameter Initialization for Photogrammetric Reconstruction

    algorithm

    To prevent non-linear optimization of the camera poses and scene model from falling into local minima, initial estimates of camera rotations, camera translations, and model dimensions are computed sequentially using decoupled linear/quadratic subproblems.

    Input: Marked 2D edge segments {(x1_i, y1_i), (x2_i, y2_i)}, camera focal lengths f_j, and hierarchical model definitions with principal axis alignments
    Output: Initial camera rotations R_j, camera translations t_j, and model parameters X
    // Stage 1: Estimate camera rotations independently of translation and scale
    for each camera j do
        Compute measured normal vectors m'_i = (x1_i, y1_i, -f_j)^T x (x2_i, y2_i, -f_j)^T for observed edges corresponding to known principal directions v_i in {x_hat, y_hat, z_hat}
        Find R_j that minimizes O_1(R_j) = sum_i (m'_i^T R_j v_i)^2
    end for
    // Stage 2: Estimate camera translations and linear model parameters
    Form objective function O_2(X, {t_j}) = sum_i [ (m_i^T R_j (P_i(X) - t_j))^2 + (m_i^T R_j (Q_i(X) - t_j))^2 ]
    where P_i(X) and Q_i(X) are expressions for the 3D endpoint vertices of model edge i
    Solve the resulting linear system / quadratic form for translation vectors t_j and parameter vector X with fixed rotations R_j
    return {R_j}, {t_j}, X
  5. Knowl 5 — View-Dependent Texture Mapping (VDTM)

    model/method

    View-dependent texture mapping (VDTM) renders novel views of a scene by dynamically projecting and compositing original calibrated photographs onto a 3D geometric proxy model based on the position of the virtual viewpoint.

    Each input photograph is mapped onto the 3D model using projective texture mapping from its recovered camera pose. Visibility and self-shadowing with respect to the camera are computed using an image-space shadow map algorithm via z-buffering.

    When multiple photographs project onto the same visible model surface, their contributions at a given pixel (or polygon face) are blended via a weighted average. The weight wkw_k assigned to photograph kk is inversely proportional to the angle αk\alpha_k between the viewing ray of the virtual camera and the optical ray from the original camera kk to the surface point (wk∝1/αkw_k \propto 1 / \alpha_k). To avoid visible seams:

    • Weights wkw_k are smoothly ramped down to zero toward the image boundaries of each projected photograph.
    • Occluding objects (such as pedestrians or vehicles) are masked out in the input photos, setting their projection weight to zero, and any unobserved regions across all photos are filled using depth-based hole-filling algorithms.
    • Face-constant weighting (computed at polygon centers) or coarsely subdivided polygon blending is used for real-time graphics hardware pipeline execution.
  6. Knowl 6 — Constrained Hierarchical Polyhedral Representation

    model/method

    Architectural scenes are represented as a hierarchical tree of parameterized polyhedral primitives (called blocks, such as boxes, wedges, prisms, and surfaces of revolution) connected by spatial relations.

    Each block's vertex coordinates P(X)\mathbf{P}(\mathbf{X}) are linear combinations of its internal shape parameters (e.g., width, height, length) within its local coordinate frame. The transformation from a block to the world coordinate system through its ancestral link chain g1(X),…,gn(X)g_1(\mathbf{X}), \dots, g_n(\mathbf{X}) is given by: Pw(X)=g1(X)⋯gn(X) P(X)\mathbf{P}_w(\mathbf{X}) = g_1(\mathbf{X}) \cdots g_n(\mathbf{X})\,\mathbf{P}(\mathbf{X}) vw(X)=g1(X)⋯gn(X) v(X)\mathbf{v}_w(\mathbf{X}) = g_1(\mathbf{X}) \cdots g_n(\mathbf{X})\,\mathbf{v}(\mathbf{X}) in homogeneous coordinates, where X\mathbf{X} is the vector of all model parameters.

    Spatial relations between parent and child blocks enforce structural constraints:

    • Rotations R\mathbf{R} can be unconstrained (3 DOF), single-axis constrained (1 DOF), or fixed/null (0 DOF).
    • Translation vectors t\mathbf{t} can be explicitly parameterized or constrained via bounding box alignments (e.g., setting the vertical translation ty=parentyMAX−childyMINt_y = \text{parent}_{y}^{\text{MAX}} - \text{child}_{y}^{\text{MIN}}).
    • Instantiated block parameters refer to symbolic named variables; multiple parameters referencing identical symbols enforce geometric symmetries and dimension equalities. This representation dramatically reduces the degrees of freedom (e.g., reducing a model from 2,896 independent line segment parameters or 240 block parameters down to 33 variables).
  7. Knowl 7 — Symbolic Newton-Raphson Optimization for Scene and Pose Recovery

    algorithm

    Following two-stage parameter initialization, the full non-linear photogrammetric error objective O=∑iErri\mathcal{O} = \sum_i Err_i is minimized over the combined vector of all unknown model dimensions X\mathbf{X} and external camera poses {(Rj,tj)}\{(\mathbf{R}_j, \mathbf{t}_j)\}.

    Because the world coordinates of points d(X)\mathbf{d}(\mathbf{X}) and directions v(X)\mathbf{v}(\mathbf{X}) are simple linear/hierarchical expressions of the symbolic block parameters, analytical expressions for line normal vectors m=Rj(v×(d−tj))\mathbf{m} = \mathbf{R}_j (\mathbf{v} \times (\mathbf{d} - \mathbf{t}_j)) and the error term Erri=mT(ATBA)mErr_i = \mathbf{m}^T (\mathbf{A}^T \mathbf{B} \mathbf{A}) \mathbf{m} are constructed symbolically.

    The gradient ∇O\nabla \mathcal{O} and Hessian H\mathbf{H} are evaluated symbolically and updated at each iteration using a variant of the Newton-Raphson method. From the initial estimates, the optimization typically converges in fewer than 10 iterations, adjusting parameters by only a few percent and achieving sub-pixel alignment (less than 1 pixel deviation) between projected model edges and marked photograph features.

  8. Knowl 8 — Hybrid Architecture Modeling and Rendering Pipeline

    model/method

    The hybrid photogrammetric and image-based architecture pipeline integrates geometry-based recovery, model-based stereo, and view-dependent rendering:

    1. Data Acquisition & Calibration: A sparse set of calibrated photographs is acquired from widely spaced positions.
    2. Photogrammetric Proxy Modeling: The user interactively defines a hierarchical block model and marks corresponding 2D edge segments. The system solves decoupled initializations and refines camera poses and block parameters via symbolic non-linear optimization.
    3. Model-Based Stereo Refinement: Offset images are projected onto the recovered proxy geometry and reprojected into the key camera view. Correlation matching along epipolar lines recovers dense sub-pixel disparity and depth maps for unmodeled architectural relief.
    4. View-Dependent Rendering: Novel views are synthesized by warping key images according to recovered depth maps and compositing multiple projected views with view-angle dependent weighting and boundary ramping.
  9. Knowl 9 — Reconstruction Accuracy on Architectural Benchmarks

    empirical result

    The photogrammetric and model-based stereo pipeline demonstrated precise geometric recovery and novel view synthesis across real-world architectural benchmarks:

    • Berkeley Campanile (Clock Tower): Reconstructed from a single photograph by exploiting vertical symmetry and hierarchical axial constraints. The resulting 33-parameter model allowed synthetic aerial rendering from 250 feet above ground, reprojecting onto the original image within 1-pixel accuracy.
    • University High School (Urbana, IL): Entire exterior modeled using 12 photographs digitized at 1536×10241536 \times 1024 and downsampled to 768×512768 \times 512 after lens distortion correction (initial geometry recovered from 5 images). The model supported a continuous 360∘360^\circ virtual fly-around animation.
    • Peterhouse Chapel (Cambridge, UK): Fine unmodeled relief (doorway mouldings and arch ornaments) was recovered from wide-baseline views using a two-plane proxy model via model-based stereopsis, producing unedited disparity maps and continuous rotation renderings without manual point seeding.

    Each complete model required approximately 4 hours of interactive modeling time.

  10. Knowl 10 — Limitations in Material Photometry and Curved Surface Recovery

    limitation

    The modeling and rendering pipeline has several identified limitations:

    • Lambertian Assumption & Photometry: Model-based stereo and view-dependent blending assume predominantly Lambertian reflectance and consistent illumination across photographs. Non-Lambertian specularities, varying times of day, or differing weather conditions cause composite seams or matching errors unless surface BRDFs are explicitly estimated.
    • Curved Surface Recovery: The photogrammetric optimization algorithm natively solves only for polyhedral parameters; non-polyhedral structures (such as domes or columns) require manual parameter adjustment or separate image contour recovery algorithms.
    • Resampling and Sparse View Selection: The angular weighting function wk∝1/αkw_k \propto 1/\alpha_k does not account for image resampling artifacts or scale mismatches when original views are taken at vastly different distances or arbitrary viewpoints.

Coverage note — None was omitted. All key contributions—including hierarchical parametric block representation, the closed-form edge disparity metric, decoupled initialization, model-based stereo with warped offset images, epipolar invariance proof, view-dependent texture mapping, and experimental results—are fully represented.

References

  1. 1.Ali Azarbayejani and Alex Pentland. Recursive estimation of motion, structure, and focal length. IEEE Trans. Pattern Anal. Machine Intell., 17(6):562–575, June 1995.
  2. 2.H. H. Baker and T. O. Binford. Depth from edge and intensity based stereo. In Proceedings of the Seventh IJCAI, Vancouver, BC, pages 631–636, 1981.
  3. 3.Paul E. Debevec, Camillo J. Taylor, and Jitendra Malik. Modeling and rendering architecture from photographs: A hybrid geometry- and image-based approach. Technical Report UCB//CSD-96-893, U.C. Berkeley, CS Division, January 1996.
  4. 4.D.J. Fleet, A.D. Jepson, and M.R.M. Jenkin. Phase-based disparity measurement. CVGIP: Image Understanding, 53(2):198–210, 1991.
  5. 5.Oliver Faugeras and Giorgio Toscani. The calibration problem for stereo. In Proceedings IEEE CVPR 86, pages 15–20, 1986.
  6. 6.Olivier Faugeras. Three-Dimensional Computer Vision. MIT Press, 1993.
  7. 7.Olivier Faugeras, Stephane Laveau, Luc Robert, Gabriella Csurka, and Cyril Zeller. 3-d reconstruction of urban scenes from sequences of images. Technical Report 2572, INRIA, June 1995.
  8. 8.W. E. L. Grimson. From Images to Surface. MIT Press, 1981.
  9. 9.D. Jones and J. Malik. Computational framework for determining stereo correspondence from a set of linear spatial filters. Image and Vision Computing, 10(10):699–708, December 1992.
  10. 10.E. Kruppa. Zur ermittlung eines objectes aus zwei perspektiven mit innerer orientierung. Sitz.-Ber. Akad. Wiss., Wien, Math. Naturw. Kl., Abt. Ila., 122:1939–1948, 1913.
  11. 11.H.C. Longuet-Higgins. A computer algorithm for reconstructing a scene from two projections. Nature, 293:133–135, September 1981.
  12. 12.D. Marr and T. Poggio. A computational theory of human stereo vision. Proceedings of the Royal Society of London, 204:301–328, 1979.
  13. 13.Leonard McMillan and Gary Bishop. Plenoptic modeling: An image-based rendering system. In SIGGRAPH ’95, 1995.
  14. 14.Eric N. Mortensen and William A. Barrett. Intelligent scissors for image composition. In SIGGRAPH ’95, 1995.
  15. 15.S. B. Pollard, J. E. W. Mayhew, and J. P. Frisby. A stereo correspondence algorithm using a disparity gradient limit. Perception, 14:449–470, 1985.
  16. 16.R. Szeliski. Image mosaicing for tele-reality applications. In IEEE Computer Graphics and Applications, 1996.
  17. 17.Camillo J. Taylor and David J. Kriegman. Structure and motion from line segments in multiple images. IEEE Trans. Pattern Anal. Machine Intell., 17(11), November 1995.
  18. 18.S. J. Teller, Celeste Fowler, Thomas Funkhouser, and Pat Hanrahan. Partitioning and ordering large radiosity computations. In SIGGRAPH ’94, pages 443–450, 1994.
  19. 19.Carlo Tomasi and Takeo Kanade. Shape and motion from image streams under orthography: a factorization method. International Journal of Computer Vision, 9(2):137–154, November 1992.
  20. 20.Roger Tsai. A versatile camera calibration technique for high accuracy 3d machine vision metrology using off-the-shelf tv cameras and lenses. IEEE Journal of Robotics and Automation, 3(4):323–344, August 1987.
  21. 21.S. Ullman. The Interpretation of Visual Motion. The MIT Press, Cambridge, MA, 1979.
  22. 22.L Williams. Casting curved shadows on curved surfaces. In SIGGRAPH ’78, pages 270–274, 1978.
  23. 23.Lance Williams and Eric Chen. View interpolation for image synthesis. In SIGGRAPH ’93, 1993.
  24. 24.Mourad Zerroug and Ramakant Nevatia. Segmentation and recovery of shgcs from a real intensity image. In European Conference on Computer Vision, pages 319–330, 1994.

Citation

MLA
Debevec, P. E., et al. “Modeling and Rendering Architecture from Photographs”. Proceedings of the 23rd Annual Conference on Computer Graphics and Interactive Techniques, 1996, pp. 11–20, https://doi.org/10.1145/237170.237191.
APA
Debevec, P. E., Taylor, C. J., & Malik, J. (1996). Modeling and rendering architecture from photographs. Proceedings of the 23rd Annual Conference on Computer Graphics and Interactive Techniques, 11–20. https://doi.org/10.1145/237170.237191
Chicago
Debevec, P. E., C. J. Taylor, and J. Malik. 1996. “Modeling and Rendering Architecture from Photographs”. Proceedings of the 23rd Annual Conference on Computer Graphics and Interactive Techniques, 11–20. https://doi.org/10.1145/237170.237191.
Harvard
Debevec, P.E., Taylor, C.J. and Malik, J. (1996) “Modeling and rendering architecture from photographs”, Proceedings of the 23rd annual conference on Computer graphics and interactive techniques. ACM, pp. 11–20. Available at: https://doi.org/10.1145/237170.237191.
Vancouver
1. Debevec PE, Taylor CJ, Malik J (1996) Modeling and rendering architecture from photographs. In: Proceedings of the 23rd annual conference on Computer graphics and interactive techniques. ACM, pp 11–20

BibTeX

@inproceedings{Debevec_1996, series={SIGGRAPH96}, title={Modeling and rendering architecture from photographs: a hybrid geometry- and image-based approach}, url={http://dx.doi.org/10.1145/237170.237191}, DOI={10.1145/237170.237191}, booktitle={Proceedings of the 23rd annual conference on Computer graphics and interactive techniques}, publisher={ACM}, author={Debevec, Paul E. and Taylor, Camillo J. and Malik, Jitendra}, year={1996}, month=Aug, pages={11–20}, collection={SIGGRAPH96} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF