Shape from Focus

S. NayarY. Nakagawa

article1994TPAMI1,341 citations

Proposes a high-precision shape recovery framework that pairs a Sum-Modified-Laplacian focus operator with Gaussian interpolation across varying focus levels to accurately reconstruct dense 3D depth maps of microscopic and rough-textured industrial surfaces.

Listen

Recovering dense and accurate three-dimensional surface shapes of rough or microscopic objects presents a major challenge in automated visual inspection. Conventional optical techniques, such as stereo vision, structured light, and shape from shading, frequently fail when applied to microscopic surfaces because local texture variations and complex reflections cause unpredictable image intensity fluctuations.

The article demonstrates a fully automated "shape-from-focus" method designed to accurately reconstruct dense depth maps of rough and textured microscopic surfaces. It evaluates whether analyzing focus variations across a sequence of images can reliably determine surface depth without relying on regular surface geometry or complex lighting models.

The approach translates an object through the focused plane of an optical microscope in controlled step increments, acquiring a series of images at fixed optical magnification. A specialized focus operator—termed the sum-modified-Laplacian—calculates the degree of focus within local image windows. Because the operator's response peaks in a bell-shaped (Gaussian) curve near the point of sharpest focus, the system interpolates depth from just three focus measurements around the peak rather than requiring hundreds of finely spaced images. Precomputed look-up tables are used to perform this interpolation quickly during operation.

Experimental testing on microscopic industrial specimens yielded three primary findings. First, the Gaussian interpolation algorithm reduced the mean absolute depth error by more than 50% compared to coarse, non-interpolated peak matching (reducing mean error from 30.32 microns to 13.82 microns on a calibrated steel sphere). Second, the method maintained high depth accuracy across varied surface orientations without systematic bias. Third, the fully automated setup—tested on microelectronic manufacturing defects such as improperly filled via-holes on ceramic substrates—successfully reconstructed both 3D depth topographies and fully focused composite images.

These findings indicate that shape-from-focus provides a practical, robust solution for micro-scale quality control and defect detection where alternative vision systems struggle. In industrial environments, this technique can help prevent electrical faults and structural defects in microelectronics manufacturing without requiring complex calibration for varying surface textures.

To move toward production deployment, organizations should consider implementing the algorithm on dedicated hardware. While the workstation-based prototype required approximately 40 seconds to process a sequence of 10 images, custom hardware can reduce processing time to under 1 second, enabling real-time visual inspection on assembly lines.

Confidence in the method's accuracy is high for textured materials, but limitations remain. The system requires sufficient local surface texture or roughness to evaluate focus; completely smooth, featureless regions can produce depth errors that require secondary filtering. In addition, step sizes and focus evaluation window sizes must be selected to balance computational speed with the resolution of fine surface details.

Cover for Shape from Focus

Abstract

The shape from focus method presented here uses different focus levels to obtain a sequence of object images. The sum-modified-Laplacian (SML) operator is developed to provide local measures of the quality of image focus. The operator is applied to the image sequence to determine a set of focus measures at each image point. A depth estimation algorithm interpolates a small number of focus measure values to obtain accurate depth estimates. A fully automated shape from focus system has been implemented using an optical microscope and tested on a variety of industrial samples. Experimental results are presented that demonstrate the accuracy and robustness of the proposed method. These results suggest shape from focus to be an effective approach for a variety of challenging visual inspection tasks.

Table of Contents

  • I. INTRODUCTION
  • II. FOCUSED AND DEFOCUSED IMAGES
  • III. SHAPE FROM FOCUS
  • IV. A FOCUS MEASURE OPERATOR
  • V. EVALUATING THE FOCUS MEASURE
  • VI. DEPTH ESTIMATION
  • VII. AUTOMATED SHAPE FROM FOCUS SYSTEM
  • VIII. DISCUSSION
  • ACKNOWLEDGMENT
  • REFERENCES

Knowls

  1. Knowl 1 — Sum-Modified-Laplacian Focus Measure Operator

    model/method

    In textured images, second derivatives along orthogonal directions often have opposite signs and cancel each other out when evaluated using the standard Laplacian ∇2I=∂2I∂x2+∂2I∂y2\nabla^2 I = \frac{\partial^2 I}{\partial x^2} + \frac{\partial^2 I}{\partial y^2}. To prevent this cancellation, the Modified Laplacian (MLML) at pixel coordinates (x,y)(x,y) computes the sum of absolute second derivatives using a variable difference step SS:

    ML(x,y)=∣2I(x,y)−I(x−S,y)−I(x+S,y)∣+∣2I(x,y)−I(x,y−S)−I(x,y+S)∣ML(x,y) = |2I(x,y) - I(x - S, y) - I(x + S, y)| + |2I(x,y) - I(x, y - S) - I(x, y + S)|

    where I(x,y)I(x,y) denotes image intensity at pixel (x,y)(x,y), and SS is a discrete step size adjusted to match the scale of surface texture elements (typically S=1S=1).

    The Sum-Modified-Laplacian (SMLSML) focus measure F(i,j)F(i,j) over a local window of size (2N+1)×(2N+1)(2N+1) \times (2N+1) centered at pixel (i,j)(i,j) sums all Modified Laplacian responses that exceed a noise threshold T1T_1:

    F(i,j)=∑x=i−Ni+N∑y=j−Nj+NML(x,y)for ML(x,y)≥T1F(i,j) = \sum_{x=i-N}^{i+N} \sum_{y=j-N}^{j+N} ML(x,y) \quad \text{for } ML(x,y) \ge T_1

    where NN specifies window radius (typically N=1N=1 for 3×33\times 3 or N=2N=2 for 5×55\times 5). Pixel responses below T1T_1 are discarded to suppress sensor noise.

  2. Knowl 2 — Gaussian Model of the Focus Measure Profile

    model/method

    For a surface element imaged across a sequence of discrete focus positions, the focus measure response F(d)F(d) as a function of translation stage displacement dd can be modeled locally near its maximum by a continuous Gaussian distribution:

    F(d)=Fpexp⁡(−12(d−dˉσF)2)F(d) = F_p \exp\left( -\frac{1}{2} \left(\frac{d - \bar{d}}{\sigma_F}\right)^2 \right)

    where:

    • dd is the translation stage displacement along the optical axis,
    • dˉ\bar{d} is the displacement corresponding to the peak focus position, representing the true depth of the surface point,
    • FpF_p is the peak focus measure value, which scales with the high-frequency texture strength of the local surface element,
    • σF\sigma_F is the Gaussian spread parameter, determined by the depth of field of the imaging system and the spatial frequency distribution of the surface texture.

    Applying the natural logarithm transforms the model into a quadratic relationship:

    ln⁡F(d)=ln⁡Fp−12(d−dˉσF)2\ln F(d) = \ln F_p - \frac{1}{2}\left(\frac{d - \bar{d}}{\sigma_F}\right)^2

  3. Knowl 3 — Three-Point Closed-Form Gaussian Interpolation for Depth Estimation

    equation

    Let dmd_m be the discrete stage displacement yielding the maximum focus measure FmF_m among a sequence of measurements, and let Fm−1F_{m-1} and Fm+1F_{m+1} be the focus measures at adjacent stage positions dm−1d_{m-1} and dm+1d_{m+1} such that Fm≥Fm−1F_m \ge F_{m-1} and Fm≥Fm+1F_m \ge F_{m+1}. Assuming a constant stage increment Δd=dm−dm−1=dm+1−dm\Delta d = d_m - d_{m-1} = d_{m+1} - d_m, fitting the Gaussian profile F(d)=Fpexp⁡(−12(d−dˉσF)2)F(d) = F_p \exp\left(-\frac{1}{2}\left(\frac{d - \bar{d}}{\sigma_F}\right)^2\right) through these three points yields the closed-form depth estimate dˉ\bar{d}:

    dˉ=(ln⁡Fm−ln⁡Fm+1)(dm2−dm−12)−(ln⁡Fm−ln⁡Fm−1)(dm2−dm+12)2Δd[(ln⁡Fm−ln⁡Fm−1)+(ln⁡Fm−ln⁡Fm+1)]\bar{d} = \frac{(\ln F_m - \ln F_{m+1})(d_m^2 - d_{m-1}^2) - (\ln F_m - \ln F_{m-1})(d_m^2 - d_{m+1}^2)}{2\Delta d \left[(\ln F_m - \ln F_{m-1}) + (\ln F_m - \ln F_{m+1})\right]}

    The corresponding spread parameter σF\sigma_F and peak value FpF_p are given by:

    σF2=−(dm2−dm−12)+(dm2−dm+12)2[(ln⁡Fm−ln⁡Fm−1)+(ln⁡Fm−ln⁡Fm+1)]\sigma_F^2 = -\frac{(d_m^2 - d_{m-1}^2) + (d_m^2 - d_{m+1}^2)}{2\left[(\ln F_m - \ln F_{m-1}) + (\ln F_m - \ln F_{m+1})\right]}

    Fp=Fm/exp⁡(−12(dm−dˉσF)2)F_p = F_m \Big/ \exp\left( -\frac{1}{2}\left(\frac{d_m - \bar{d}}{\sigma_F}\right)^2 \right)

  4. Knowl 4 — Look-Up Table Acceleration for Gaussian Depth Interpolation

    algorithm

    To eliminate online evaluation of logarithms, divisions, and transcendental functions during depth estimation, the interpolation equations are parameterized using normalized focus measures and displacements:

    F^m−1=Fm−1Fm,F^m=1,F^m+1=Fm+1Fm\hat{F}_{m-1} = \frac{F_{m-1}}{F_m}, \quad \hat{F}_m = 1, \quad \hat{F}_{m+1} = \frac{F_{m+1}}{F_m} d^m−1=−1,d^m=0,d^m+1=1\hat{d}_{m-1} = -1, \quad \hat{d}_m = 0, \quad \hat{d}_{m+1} = 1

    Two-dimensional look-up tables (LUTs) precomputed over (F^m−1,F^m+1)∈(0,1]2(\hat{F}_{m-1}, \hat{F}_{m+1}) \in (0, 1]^2 store the normalized depth d^\hat{d}, spread σ^F\hat{\sigma}_F, and peak value F^p\hat{F}_p.

    Input: Discrete focus measures F1,F2,…,FMF_1, F_2, \dots, F_M at stage positions di=i⋅Δdd_i = i \cdot \Delta d
    Output: Continuous depth estimate dˉ\bar{d}, spread σF\sigma_F, peak value FpF_p
    Find index mm such that Fm=max⁡iFiF_m = \max_i F_i
    F^m−1←Fm−1/Fm\hat{F}_{m-1} \leftarrow F_{m-1} / F_m
    F^m+1←Fm+1/Fm\hat{F}_{m+1} \leftarrow F_{m+1} / F_m
    d^←LUTd[F^m−1,F^m+1]\hat{d} \leftarrow \text{LUT}_d[\hat{F}_{m-1}, \hat{F}_{m+1}]
    σ^F←LUTσ[F^m−1,F^m+1]\hat{\sigma}_F \leftarrow \text{LUT}_\sigma[\hat{F}_{m-1}, \hat{F}_{m+1}]
    F^p←LUTp[F^m−1,F^m+1]\hat{F}_p \leftarrow \text{LUT}_p[\hat{F}_{m-1}, \hat{F}_{m+1}]
    dˉ←dm+d^⋅Δd\bar{d} \leftarrow d_m + \hat{d} \cdot \Delta d
    σF←σ^F⋅Δd\sigma_F \leftarrow \hat{\sigma}_F \cdot \Delta d
    Fp←Fm⋅F^pF_p \leftarrow F_m \cdot \hat{F}_p
    return dˉ,σF,Fp\bar{d}, \sigma_F, F_p
  5. Knowl 5 — Shape from Focus System Pipeline

    model/method

    The automated Shape from Focus pipeline reconstructs dense 3D depth maps of microscopic objects through the following stages:

    1. Image Acquisition via Stage Translation: The object is placed on a computer-controlled translation stage (step resolution and accuracy of 0.02 μm0.02\,\mu\text{m}) and stepped along the optical axis by a fixed increment Δd\Delta d. At each position did_i, an image is captured through a fixed microscope and CCD camera (512×480512\times 480 pixels).
    2. Focus Measure Computation: The Sum-Modified-Laplacian (SML) operator is applied at every pixel across all images in the sequence, producing a sequence of focus measure values F(di)F(d_i) per pixel.
    3. Coarse Peak Detection: For each pixel, the stage position dmd_m corresponding to the discrete maximum focus measure Fm=max⁡iF(di)F_m = \max_i F(d_i) is identified.
    4. Sub-Step Interpolation: A Gaussian curve is fitted to the triplet (Fm−1,Fm,Fm+1)(F_{m-1}, F_m, F_{m+1}) using the closed-form interpolation formula or precomputed look-up tables to determine the continuous depth estimate dˉ\bar{d}.
    5. Post-Filtering: A 5×55\times 5 median filter is applied to the depth map to eliminate spurious depth values resulting from smooth or weakly textured surface areas.
    6. All-in-Focus Reconstruction: An all-in-focus image is generated by assigning to each pixel (x,y)(x,y) the intensity from the image frame whose displacement did_i is closest to dˉ(x,y)\bar{d}(x,y).
  6. Knowl 6 — Constant Magnification via Object Stage Translation

    model/method

    Varying optical focus by shifting the sensor plane or moving the lens alters the magnification of the image, causing image coordinates of scene points to shift and brightness to vary across frames. Shape from Focus eliminates magnification variations for focused points by fixing the optical system (lens and sensor) and translating the object on a motorized stage through the stationary focused plane.

    Because the focused plane is fixed in space, any surface element on the object that passes through this plane is imaged with identical magnification. For small stage displacements Δd\Delta d, magnification variations for points located near the focused plane remain negligible, allowing focus measures across consecutive images to be evaluated within a fixed image window.

  7. Knowl 7 — Equivalence of Focus Measurement and Defocused High-Pass Filtering

    theoretical result

    Let If(x,y)I_f(x,y) be the focused image of a planar surface, h(x,y)h(x,y) be a linear shift-invariant defocus blur kernel, and o(x,y)o(x,y) be a linear focus measure operator kernel. The observed defocused image is Id(x,y)=h(x,y)∗If(x,y)I_d(x,y) = h(x,y) * I_f(x,y). Applying the focus operator to Id(x,y)I_d(x,y) produces the output:

    r(x,y)=o(x,y)∗(h(x,y)∗If(x,y))r(x,y) = o(x,y) * (h(x,y) * I_f(x,y))

    By the commutativity and associativity of convolution, this is equivalent to:

    r(x,y)=h(x,y)∗(o(x,y)∗If(x,y))r(x,y) = h(x,y) * (o(x,y) * I_f(x,y))

    This demonstrates that applying a linear focus measure operator to a defocused image is equivalent to applying optical defocus blur to an image pre-filtered by the operator. Consequently, an effective focus measure operator must function as a high-pass filter to ensure that optical low-pass defocus attenuation produces a sharp, detectable peak at the true in-focus displacement.

  8. Knowl 8 — Experimental Depth Accuracy on a Spherical Test Sample

    data/table

    A steel test sphere of known diameter 1590 μm1590\,\mu\text{m} with a rough, textured surface was translated through the focused plane of an optical microscope in steps of Δd=100 μm\Delta d = 100\,\mu\text{m} across 13 images. The Sum-Modified-Laplacian (SML) operator was applied with a 5×55\times 5 window (N=2N=2), threshold T1=7T_1 = 7, and step S=1S = 1. The resulting depth maps were quantitatively evaluated against the known sphere geometry:

    Metric Coarse Interpolation Gaussian Interpolation
    Diameter of Test Sphere 1590 μm1590\,\mu\text{m} 1590 μm1590\,\mu\text{m}
    Number of Points 22682 23257
    Mean Error (μm\mu\text{m}) 7.861 3.857
    Mean Absolute Error (μm\mu\text{m}) 30.32 13.815
    Maximum Absolute Error (μm\mu\text{m}) 187.80 175.82

    Gaussian three-point interpolation reduced the mean absolute depth error from 30.32 μm30.32\,\mu\text{m} (obtained by discrete peak selection) to 13.815 μm13.815\,\mu\text{m}, representing an error reduction of more than 54%54\%.

  9. Knowl 9 — Operating Assumptions and Limitations of Shape from Focus

    limitation

    The accuracy and validity of the Shape from Focus method depend on several operational constraints:

    1. Surface Texture Requirement: The surface must possess microscopic roughness or reflectance variations. On smooth, textureless, or specular surfaces, the Modified Laplacian remains below the threshold T1T_1, resulting in zero focus measure values and unreliable depth estimates that require spatial filtering.
    2. Local Gaussian Validity: The focus measure profile is well-approximated by a Gaussian only near its primary peak. Outer fringes deviate from Gaussian behavior and are sensitive to image noise and slight magnification changes.
    3. Window Size Trade-Off: The SML window size (2N+1)(2N+1) introduces a trade-off: large windows enhance signal-to-noise ratio in weakly textured regions but smooth out high-frequency surface topography and depth boundaries; small windows retain localized depth variations but are sensitive to noise.
    4. Displacement Step Size: The stage step Δd\Delta d must be sufficiently small to ensure that the three highest focus measurements (Fm−1,Fm,Fm+1F_{m-1}, F_m, F_{m+1}) lie within the primary Gaussian mode.

Coverage note — The qualitative reconstruction experiment on the ceramic substrate via-hole filling (Section VII) was omitted as its methodological principles and quantitative conclusions duplicate those of the steel sphere benchmark.

References

  1. 1.M. Born and E. Wolf, Principles of Optics. London: Pergamon, 1965.
  2. 2.T. Darrell and K. Wohn, "Pyramid based depth from focus," in Proc. CVPR, 1988, pp. 504–509.
  3. 3.J. Ens and P. Lawrence, "A matrix based method for determining depth from focus," in Proc. CVPR, 1991, pp. 600–606.
  4. 4.P. Grossmann, "Depth from focus," Pattern Recognit. Lett., vol. 5, pp. 63–69, 1987.
  5. 5.B. K. P. Horn, "Focusing," MIT Artificial Intell. Lab., Memo no. 160, May, 1968
  6. 6.---, Robot Vision. Cambridge, MA: MIT Press, 1986.
  7. 7.R. A. Jarvis, "Focus optimization criteria for computer image processing," Microscope, vol. 24, no. 2, pp. 163–180, 1976.
  8. 8.J. F. Koretz and G. H. Handelman, "How the human eye focuses," Scientific American, pp. 92–99, July 1988.
  9. 9.E. Krotkov, "Focusing," Int. J. of Comput. Vision, vol. 1, pp. 223–237 1987.
  10. 10.G. Ligthart and F. Groen, "A comparison of different autofocus algorithms," in Proc. of ICPR, 1982, pp. 597–600.
  11. 11.H. N. Nair and C. V. Stewart, "Robust focus ranging," in Proc. CVPR, 1992, pp. 309–314.
  12. 12.S. K. Nayar and Y. Nakagawa, "Shape from focus," Tech. Rep., Dept. of Comput. Sci., Columbia Univ., CUCS 058-92, Nov. 1992.
  13. 13.A. Pentland, "A new sense for depth of field," IEEE Trans. Pattern Analysis and Machine Intell., vol. 9, no. 4, pp. 523–531, July 1987.
  14. 14.J. F. Schlag, A. C. Sanderson, C. P. Neumann, and F. C. Wimberly, "Implementation of automatic focusing algorithms for a computer vision system with camera control," Tech. Rep., Carnegie Mellon Univ., CMU-RI-TR-83-14, Aug. 1983.
  15. 15.M. Subbarao, "Efficient depth recovery through inverse optics," Machine Vision for Inspection and Measurement, H. Freeman, Ed.. New York: Academic Press, 1989.
  16. 16.M. Subbarao and G. Surya, "Depth from defocus: A spatial domain approach," Tech. Rep. 92.12.03, SUNY, Stony Brook, Dec. 1992.
  17. 17.J. M. Tenenbaum, "Accommodation in computer vision," Ph.D. thesis, Stanford Univ., 1970.
  18. 18.R. G. Willson and S. A. Shafer, "Dynamic lens compensation for active color imaging and constant magnification focusing," Tech. Rep., Carnegie Mellon Univ., CMU-RI-TR-91, Oct. 1991.

Citation

MLA
Nayar, S. K., and Y. Nakagawa. “Shape from Focus”. IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 16, no. 8, 1994, pp. 824–31, https://doi.org/10.1109/34.308479.
APA
Nayar, S. K., & Nakagawa, Y. (1994). Shape from focus. IEEE Transactions on Pattern Analysis and Machine Intelligence, 16(8), 824–831. https://doi.org/10.1109/34.308479
Chicago
Nayar, S. K., and Y. Nakagawa. 1994. “Shape from Focus”. IEEE Transactions on Pattern Analysis and Machine Intelligence 16 (8): 824–31. https://doi.org/10.1109/34.308479.
Harvard
Nayar, S.K. and Nakagawa, Y. (1994) “Shape from focus”, IEEE Transactions on Pattern Analysis and Machine Intelligence, 16(8), pp. 824–831. Available at: https://doi.org/10.1109/34.308479.
Vancouver
1. Nayar SK, Nakagawa Y (1994) Shape from focus. IEEE Transactions on Pattern Analysis and Machine Intelligence 16:824–831

BibTeX

@article{Nayar_1994, title={Shape from focus}, volume={16}, ISSN={0162-8828}, url={http://dx.doi.org/10.1109/34.308479}, DOI={10.1109/34.308479}, number={8}, journal={IEEE Transactions on Pattern Analysis and Machine Intelligence}, publisher={Institute of Electrical and Electronics Engineers (IEEE)}, author={Nayar, S.K. and Nakagawa, Y.}, year={1994}, pages={824–831} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF