Camera Calibration with Distortion Models and Accuracy Evaluation

Juyang WengPaul R. CohenMarc Herniou

article1992TPAMI2,002 citations

Develops a two-step camera calibration framework incorporating radial, decentering, and thin prism distortions through closed-form initialization followed by full nonlinear optimization, complemented by a normalized error metric to quantitatively benchmark calibration accuracy across different optical setups.

Listen

Accurate camera calibration is essential for computer vision applications requiring precise three-dimensional measurements, such as robotics, quality control inspection, and stereoscopic depth estimation. Standard commercial video cameras and lenses contain notable imperfections, including radial distortion, optical misalignment, and non-orthogonal sensor planes. Existing calibration methods either ignore these distortions, handle only simple radial effects, or suffer from numerical instability when solving for all parameters simultaneously. The article aims to introduce and validate a comprehensive calibration model that simultaneously handles radial and tangential lens distortions, paired with a stable two-step estimation procedure and an objective performance metric tailored for digital imaging systems.

The authors evaluated their approach using both synthetic simulations across 50 trials and real-world laboratory experiments on stereo camera setups. The experimental hardware included telephoto lenses and wide-angle lenses observing a high-precision calibration pattern printed on an ultraflat optical glass plate. In the first phase of the calibration routine, a closed-form direct calculation determines initial estimates of internal and external camera parameters using only points near the image center, where distortion is minimal. The second phase refines these estimates and calculates five distinct distortion parameters through nonlinear optimization, decoupling distortion terms from standard geometric parameters to avoid false solutions. Calibration quality was assessed using a newly introduced metric that normalizes measurement errors against the camera's theoretical digital pixel resolution.

The findings demonstrate substantial improvements in measurement accuracy when accounting for comprehensive lens distortions. In wide-angle lens configurations, which exhibit severe warping, applying the complete distortion model reduced the normalized error from 3.98 in a distortion-free model and 1.41 in a radial-only model down to 1.16, successfully capturing radial distortion as large as 23 pixels alongside 3 pixels of tangential distortion. For telephoto lenses, the complete model brought the normalized error down from 1.75 to 1.06, aligning real-world accuracy within roughly 6% of the theoretical digital resolution limit. Furthermore, initializing the optimization with central image points proved vital, decreasing parameter estimation error by a factor of 10 in linear calculations and preventing optimization divergence.

These results demonstrate that ignoring lens distortion or correcting only radial components leaves substantial geometric errors in digital vision systems. Organizations deploying computer vision for precise spatial measurements can achieve performance near the theoretical limits of their sensors without investing in expensive, specialized metric cameras. Because the two-step procedure decouples standard and distortion parameters, it provides a stable and mathematically rigorous optimization path that circumvents common convergence failures in automated systems.

For engineering teams and decision-makers implementing machine vision solutions, adoption of complete distortion modeling is strongly recommended, especially when utilizing wide-angle optics. The calibration routine can be performed once offline during manufacturing or setup, imposing no computational penalty during real-time runtime operations. Future implementations should rely on standardized normalized error metrics rather than raw distance errors to benchmark vision hardware, though practitioners should maintain high-precision physical calibration targets to ensure that manufacturing defects in the target do not degrade calibration quality.

  • Paper: A Flexible New Technique for Camera Calibration, Zhengyou Zhang (2000). It extends parametric camera calibration by introducing a flexible, practical method using planar patterns and closed-form homography initialization followed by nonlinear refinement of distortion.
  • Paper: An efficient solution to the five-point relative pose problem, D. Nistér (2004). It builds directly on calibrated camera models to provide an exact, minimal five-point relative pose solver essential for multi-view geometry and structure-from-motion.
  • Paper: Modeling the World from Internet Photo Collections, Noah Snavely et al. (2008). It incorporates intrinsic calibration and lens distortion modeling within large-scale structure-from-motion pipelines using bundle adjustment across unconstrained photo collections.
  • Paper: Structure-from-Motion Revisited, Johannes L. Schönberger et al. (2016). It modernizes and scales structure-from-motion and multi-view triangulation pipelines that rely heavily on accurate camera calibration and distortion compensation.
  • Paper: A Comparison and Evaluation of Multi-View Stereo Reconstruction Algorithms, S. Seitz et al. (2006). It establishes standardized evaluation frameworks and benchmarks for multi-view 3D reconstruction algorithms that assume calibrated camera systems.
  • Paper: MonoSLAM: Real-Time Single Camera SLAM, Andrew J. Davison et al. (2007). It demonstrates how real-time monocular visual SLAM relies on explicit camera calibration and distortion modeling to accurately track 3D landmarks and camera trajectories.
Cover for Camera Calibration with Distortion Models and Accuracy Evaluation

Abstract

The objective of stereo camera calibration is to estimate the internal and external parameters of each camera. Using these parameters, the 3-D position of a point in the scene, which is identified and matched in two stereo images, can be determined by the method of triangulation. In this paper, we present a camera model that accounts for major sources of camera distortion, namely, radial, decentering, and thin prism distortions. The proposed calibration procedure consists of two steps. In the first step, the calibration parameters are estimated using a closed-form solution based on a distortion-free camera model. In the second step, the parameters estimated in the first step are improved iteratively through a nonlinear optimization, taking into account camera distortions. According to minimum variance estimation, the objective function to be minimized is the mean-square discrepancy between the observed image points and their inferred image projections computed with the estimated calibration parameters. We introduce a type of measure that can be used to directly evaluate the performance of calibration and compare calibrations among different systems. The validity and performance of our calibration procedure are tested with both synthetic data and real images taken by tele- and wide-angle lenses. The results consistently show significant improvements over less complete camera models.

Table of Contents

  • I. INTRODUCTION
  • 11. CAMERA MODELS
  • A. Distortion-Free Camera Model
  • B. Geometrical Distortion
  • C. Our Complete Camera Model
  • 111. RESOLUTION STRATEGIES
  • A. Procedure to Compute the Optimal Solution
  • B. Estimation of m with d Fixed
  • C. Estimation of d with m Fixed
  • IV. EVALUATION OF CALIBRATION ACCURACY
  • v. EXPERIMENTAL RESULTS AND DISCUSSIONS
  • A. Simulations with Synthetic Data
  • B. Experiments with Real Images
  • VI. SUMMARY AND CONCLUSIONS
  • APPENDIX
  • ACKNOWLEDGMENT
  • REFERENCES

Knowls

  1. Knowl 1 — Complete Camera Model with Radial, Decentering, and Thin-Prism Distortions

    model/method

    A comprehensive camera model maps a 3D point (x,y,z)(x, y, z) in the world coordinate system to observed pixel coordinates (r′,c′)(r', c') in a digitized image, accounting for internal geometry, external pose, and geometric lens distortions up to third-order terms.

    Let R=(ri,j)3×3R = (r_{i,j})_{3 \times 3} be the rotation matrix and T=(t1,t2,t3)TT = (t_1, t_2, t_3)^T be the translation vector relating world coordinates (x,y,z)(x,y,z) to camera-centered coordinates (xc,yc,zc)T=R(x,y,z)T+T(x_c, y_c, z_c)^T = R (x, y, z)^T + T. In a distortion-free pinhole model, the normalized image plane coordinates (u,v)=(xc/zc,yc/zc)(u, v) = (x_c / z_c, y_c / z_c) correspond to pixel row and column coordinates (r,c)(r, c) via: u=r−r0fu,v=c−c0fvu = \frac{r - r_0}{f_u}, \quad v = \frac{c - c_0}{f_v} where (r0,c0)(r_0, c_0) is the pixel coordinate of the principal point, and fu,fvf_u, f_v denote the effective focal lengths scaled in row and column pixel units, respectively.

    When actual distorted pixel coordinates (r′,c′)(r', c') are observed, intermediate normalized coordinates (uˉ,vˉ)(\bar{u}, \bar{v}) are defined as: uˉ=r′−r0fu,vˉ=c′−c0fv\bar{u} = \frac{r' - r_0}{f_u}, \quad \bar{v} = \frac{c' - c_0}{f_v} The combined effects of symmetrical radial distortion (parameter k1k_1), decentering distortion (due to non-collinear lens centers), and thin prism distortion (due to lens tilt or sensor array misalignment) are parameterized by five coefficients d=(k1,g1,g2,g3,g4)Td = (k_1, g_1, g_2, g_3, g_4)^T. The relationship between the world coordinates and observed distorted coordinates is governed by the two equations: r1,1x+r1,2y+r1,3z+t1r3,1x+r3,2y+r3,3z+t3=uˉ+(g1+g3)uˉ2+g4uˉvˉ+g1vˉ2+k1uˉ(uˉ2+vˉ2)\frac{r_{1,1}x + r_{1,2}y + r_{1,3}z + t_1}{r_{3,1}x + r_{3,2}y + r_{3,3}z + t_3} = \bar{u} + (g_1 + g_3)\bar{u}^2 + g_4 \bar{u}\bar{v} + g_1 \bar{v}^2 + k_1 \bar{u}(\bar{u}^2 + \bar{v}^2) r2,1x+r2,2y+r2,3z+t2r3,1x+r3,2y+r3,3z+t3=vˉ+g2uˉ2+g3uˉvˉ+(g2+g4)vˉ2+k1vˉ(uˉ2+vˉ2)\frac{r_{2,1}x + r_{2,2}y + r_{2,3}z + t_2}{r_{3,1}x + r_{3,2}y + r_{3,3}z + t_3} = \bar{v} + g_2 \bar{u}^2 + g_3 \bar{u}\bar{v} + (g_2 + g_4)\bar{v}^2 + k_1 \bar{v}(\bar{u}^2 + \bar{v}^2)

    The right-hand sides are linear with respect to the distortion parameter vector dd, which enables linear least-squares solving for dd whenever the extrinsic and non-distortion intrinsic parameters m=(r0,c0,fu,fv,T,R)m = (r_0, c_0, f_u, f_v, T, R) are fixed.

  2. Knowl 2 — Normalized Stereo Calibration Error and Normalized Calibration Error

    definition

    The Normalized Stereo Calibration Error (NSCE\text{NSCE}) is an objective performance metric for evaluating stereo camera calibration accuracy that is invariant to baseline length, field of view, digital image resolution, and object-to-camera distance. It measures the ratio of the lateral triangulation error to the standard deviation of the spatial pixel digitization noise at the object depth.

    For an image with row and column focal lengths fu,fvf_u, f_v (in pixel units), the spatial back-projection of a single pixel onto an orthogonal plane at depth ziz_i defines a rectangle of dimensions a=zi/∣fu∣a = z_i / |f_u| and b=zi/fvb = z_i / f_v. Under a uniform distribution of discretization error across the pixel area, the lateral spatial variance induced by digitization is: σi2=a2+b212=zi2(fu−2+fv−2)12\sigma_i^2 = \frac{a^2 + b^2}{12} = \frac{z_i^2 (f_u^{-2} + f_v^{-2})}{12}

    Given nn 3D control points with true camera-centered coordinates (xi,yi,zi)(x_i, y_i, z_i) and triangulated 3D positions (x^i,y^i,z^i)(\hat{x}_i, \hat{y}_i, \hat{z}_i), the mean NSCE\text{NSCE} is defined as: NSCE=1n∑i=1n[(x^i−xi)2+(y^i−yi)2zi2(fu−2+fv−2)12]1/2\text{NSCE} = \frac{1}{n} \sum_{i=1}^n \left[ \frac{(\hat{x}_i - x_i)^2 + (\hat{y}_i - y_i)^2}{\frac{z_i^2 (f_u^{-2} + f_v^{-2})}{12}} \right]^{1/2} Alternatively, the root-mean-squared NSCE\text{NSCE} is defined as: NSCERMS=[1n∑i=1n(x^i−xi)2+(y^i−yi)2zi2(fu−2+fv−2)12]1/2\text{NSCE}_{\text{RMS}} = \left[ \frac{1}{n} \sum_{i=1}^n \frac{(\hat{x}_i - x_i)^2 + (\hat{y}_i - y_i)^2}{\frac{z_i^2 (f_u^{-2} + f_v^{-2})}{12}} \right]^{1/2}

    Interpretation:

    • NSCE≈1\text{NSCE} \approx 1: the residual calibration error matches the fundamental physical limit imposed by pixel spatial digitization noise.
    • NSCE<1\text{NSCE} < 1: subpixel feature localization has successfully reduced triangulation error below nominal single-pixel quantization noise.
    • NSCE≫1\text{NSCE} \gg 1: the calibration is deficient, typically because significant optical distortion remains unmodeled.

    For single-camera calibration, the Normalized Calibration Error (NCE\text{NCE}) is defined identically, with (x^i,y^i,z^i)(\hat{x}_i, \hat{y}_i, \hat{z}_i) computed as the intersection of the back-projected optical ray of the sensed point with the ground-truth depth plane z=ziz = z_i.

  3. Knowl 3 — Two-Step Decoupled Camera Calibration Procedure

    algorithm

    To prevent harmful coupling and numerical divergence between non-distortion parameters m=(r0,c0,fu,fv,T,R)m = (r_0, c_0, f_u, f_v, T, R) and lens distortion parameters d=(k1,g1,g2,g3,g4)Td = (k_1, g_1, g_2, g_3, g_4)^T, parameter estimation is partitioned into a closed-form initialization on central points followed by alternating iterative optimization.

    Input: 3D world coordinates of nn control points (xi,yi,zi)i=1n{(x_i, y_i, z_i)}_{i=1}^n, corresponding observed pixel locations (ri′,ci′)i=1n{(r'_i, c'_i)}_{i=1}^n, subset of central points Ω1⊂{1,…,n}\Omega_1 \subset \{1, \dots, n\} whose image projections lie within the central image region (radius ≤1/4\le 1/4 image width).
    Output: Calibrated non-distortion parameter set m∗=(r0,c0,fu,fv,T,R)m^* = (r_0, c_0, f_u, f_v, T, R) and distortion parameter vector d∗=(k1,g1,g2,g3,g4)Td^* = (k_1, g_1, g_2, g_3, g_4)^T.
    1. Initialize distortion parameters d←0d \leftarrow 0.
    2. Compute initial non-distortion parameters mˉ\bar{m} on central points Ω1\Omega_1 using the closed-form linear estimation algorithm under d=0d = 0.
    3. Enforce orthonormality on initial rotation matrix Rˉ\bar{R} via quaternion eigenvector decomposition to obtain refined initial estimate m~\tilde{m}.
    4. Perform nonlinear minimization of reprojection error over mm on central points Ω1\Omega_1 with d=0d = 0 fixed, starting from m~\tilde{m}, to obtain mm.
    5. repeat
        a. Compute distortion vector dd in closed form via linear least-squares on all nn points with mm fixed:
           d←arg⁡min⁡d∥Qd+C∥2d \leftarrow \arg\min_d \|Q d + C\|^2
        b. Compute non-distortion parameters mm via nonlinear optimization of reprojection error on all nn points with dd fixed:
           m←arg⁡min⁡m∑i=1n[(ri′−ri(m,d))2+(ci′−ci(m,d))2]m \leftarrow \arg\min_m \sum_{i=1}^n [(r'_i - r_i(m, d))^2 + (c'_i - c_i(m, d))^2]
    until convergence or maximum iterations reached.
    6. return m∗=mm^* = m, d∗=dd^* = d.

    Restricting the initial linear and non-linear estimation of mm to central points prevents peripheral lens distortions from corrupting the initial pose and intrinsic estimates. Alternating between mm and dd ensures that the high parameter sensitivity of dd does not lead to spurious local minima in mm.

  4. Knowl 4 — Closed-Form Linear Estimation of Distortion-Free Parameters

    algorithm

    For an uncalibrated camera under a distortion-free model, the internal parameters (r0,c0,fu,fv)(r_0, c_0, f_u, f_v) and external parameters (R,T)(R, T) can be computed in closed form from n≥6n \ge 6 non-coplanar 3D control points (xi,yi,zi)(x_i, y_i, z_i) and their observed pixel coordinates (ri′,ci′)(r'_i, c'_i).

    Input: 3D control points (xi,yi,zi)i=1n{(x_i, y_i, z_i)}_{i=1}^n and observed pixel coordinates (ri′,ci′)i=1n{(r'_i, c'_i)}_{i=1}^n (restricted to central image points).
    Output: Initial estimates of (r0,c0,fu,fv)(r_0, c_0, f_u, f_v), translation T=(t1,t2,t3)TT = (t_1, t_2, t_3)^T, and rotation matrix row vectors R1,R2,R3R_1, R_2, R_3.
    1. For each point i∈{1,…,n}i \in \{1, \dots, n\}, construct two rows of a 2n×122n \times 12 matrix AA:
       Row 2i−12i-1: [−xi,−yi,−zi,0,0,0,ri′xi,ri′yi,ri′zi,−1,0,ri′][-x_i, -y_i, -z_i, 0, 0, 0, r'_i x_i, r'_i y_i, r'_i z_i, -1, 0, r'_i]
       Row 2i2i: [0,0,0,−xi,−yi,−zi,ci′xi,ci′yi,ci′zi,0,−1,ci′][0, 0, 0, -x_i, -y_i, -z_i, c'_i x_i, c'_i y_i, c'_i z_i, 0, -1, c'_i]
    2. Partition A=[A′∣B′]A = [A' \mid B'], where A′A' contains the first 11 columns and B′B' is the 12th column of AA.
    3. Solve the linear least-squares system A′W′+B′=0A' W' + B' = 0 for the 11-dimensional vector W′W':
       W′=−(A′TA′)−1A′TB′W' = -(A'^T A')^{-1} A'^T B'
    4. Form the 12-dimensional vector W=(W1T,W2T,W3T,w4,w5,w6)T=(W′T,1)TW = (W_1^T, W_2^T, W_3^T, w_4, w_5, w_6)^T = (W'^T, 1)^T, where W1,W2,W3∈R3W_1, W_2, W_3 \in \mathbb{R}^3 and w4,w5,w6∈Rw_4, w_5, w_6 \in \mathbb{R}.
    5. Compute the intermediate vector S=(s1T,s2T,s3T,s4,s5,s6)T=±1∥W3∥WS = (s_1^T, s_2^T, s_3^T, s_4, s_5, s_6)^T = \pm \frac{1}{\|W_3\|} W, where the sign is chosen positive if the camera is in front of the world coordinate z=0z=0 plane and negative otherwise.
    6. Compute internal and external parameters:
       r0=s1Ts3r_0 = s_1^T s_3
       c0=s2Ts3c_0 = s_2^T s_3
       fu=−∥s1−r0s3∥f_u = -\|s_1 - r_0 s_3\|
       fv=∥s2−c0s3∥f_v = \|s_2 - c_0 s_3\|
       t3=s6t_3 = s_6
       t1=(s4−r0s6)/fut_1 = (s_4 - r_0 s_6) / f_u
       t2=(s5−c0s6)/fvt_2 = (s_5 - c_0 s_6) / f_v
       R1=(s1−r0s3)/fuR_1 = (s_1 - r_0 s_3) / f_u
       R2=(s2−c0s3)/fvR_2 = (s_2 - c_0 s_3) / f_v
       R3=s3R_3 = s_3
    7. return (r0,c0,fu,fv,T,R)(r_0, c_0, f_u, f_v, T, R).
  5. Knowl 5 — Linear Least-Squares Estimation of Distortion Parameters with Fixed Camera Extrinsics and Intrinsics

    model/method

    When camera external parameters R=(ri,j)3×3,T=(t1,t2,t3)TR = (r_{i,j})_{3 \times 3}, T = (t_1, t_2, t_3)^T and non-distortion internal parameters (r0,c0,fu,fv)(r_0, c_0, f_u, f_v) are held fixed, the estimation of the five distortion parameters d=(k1,g1,g2,g3,g4)Td = (k_1, g_1, g_2, g_3, g_4)^T reduces to an unconstrained linear least-squares problem ∥Qd+C∥2→min⁡\|Q d + C\|^2 \to \min.

    For each 3D point (xi,yi,zi)(x_i, y_i, z_i) with observed pixel coordinate (ri′,ci′)(r'_i, c'_i), define the intermediate normalized coordinates: uˉi=ri′−r0fu,vˉi=ci′−c0fv\bar{u}_i = \frac{r'_i - r_0}{f_u}, \quad \bar{v}_i = \frac{c'_i - c_0}{f_v} and the projected coordinates predicted from geometry: u^i=r1,1xi+r1,2yi+r1,3zi+t1r3,1xi+r3,2yi+r3,3zi+t3,v^i=r2,1xi+r2,2yi+r2,3zi+t2r3,1xi+r3,2yi+r3,3zi+t3\hat{u}_i = \frac{r_{1,1}x_i + r_{1,2}y_i + r_{1,3}z_i + t_1}{r_{3,1}x_i + r_{3,2}y_i + r_{3,3}z_i + t_3}, \quad \hat{v}_i = \frac{r_{2,1}x_i + r_{2,2}y_i + r_{2,3}z_i + t_2}{r_{3,1}x_i + r_{3,2}y_i + r_{3,3}z_i + t_3}

    The residual discrepancy in row and column pixel coordinates yields two linear equations per point: fuuˉi(uˉi2+vˉi2)k1+fu(uˉi2+vˉi2)g1+0⋅g2+fuuˉi2g3+fuuˉivˉig4+fu(uˉi−u^i)=0f_u \bar{u}_i (\bar{u}_i^2 + \bar{v}_i^2) k_1 + f_u (\bar{u}_i^2 + \bar{v}_i^2) g_1 + 0 \cdot g_2 + f_u \bar{u}_i^2 g_3 + f_u \bar{u}_i \bar{v}_i g_4 + f_u (\bar{u}_i - \hat{u}_i) = 0 fvvˉi(uˉi2+vˉi2)k1+0⋅g1+fv(uˉi2+vˉi2)g2+fvuˉivˉig3+fvvˉi2g4+fv(vˉi−v^i)=0f_v \bar{v}_i (\bar{u}_i^2 + \bar{v}_i^2) k_1 + 0 \cdot g_1 + f_v (\bar{u}_i^2 + \bar{v}_i^2) g_2 + f_v \bar{u}_i \bar{v}_i g_3 + f_v \bar{v}_i^2 g_4 + f_v (\bar{v}_i - \hat{v}_i) = 0

    For nn control points, stacking these equations forms a 2n×52n \times 5 matrix QQ and a 2n2n-dimensional vector CC. The optimal distortion parameter vector dd is determined in closed form by: d=−(QTQ)−1QTCd = -(Q^T Q)^{-1} Q^T C This non-iterative linear solution directly minimizes the sum of squared image plane discrepancies without requiring numerical gradient descent.

  6. Knowl 6 — Closed-Form Rotation Matrix Orthogonalization via Quaternion Eigen-Decomposition

    algorithm

    Given an unconstrained 3×33 \times 3 matrix Rˉ\bar{R} (or point correspondences C=[C1,…,Cn]C = [C_1, \dots, C_n] and D=[D1,…,Dn]D = [D_1, \dots, D_n]), the closest valid rotation matrix R∈SO(3)R \in \text{SO}(3) minimizing ∥RC−D∥2\|R C - D\|^2 or ∥R−Rˉ∥F2\|R - \bar{R}\|_F^2 is solved in closed form using a unit quaternion representation.

    Input: 3D point correspondences {Ci}i=1n\{C_i\}_{i=1}^n and {Di}i=1n\{D_i\}_{i=1}^n (for matrix orthogonalization of Rˉ\bar{R}, C=I3×3C = I_{3\times 3} and D=RˉD = \bar{R}).
    Output: Optimal rotation matrix R∈SO(3)R \in \text{SO}(3) minimizing ∑i=1n∥RCi−Di∥2\sum_{i=1}^n \|R C_i - D_i\|^2.
    1. For any 3D vector v=(x,y,z)Tv = (x, y, z)^T, define the skew-symmetric cross-product matrix:
       [v]×=(0−zyz0−x−yx0)[v]_\times = \begin{pmatrix} 0 & -z & y \\ z & 0 & -x \\ -y & x & 0 \end{pmatrix}
    2. For each pair (Ci,Di)(C_i, D_i), construct the 4×44 \times 4 block matrix BiB_i:
       Bi=(0(Ci−Di)TDi−Ci[Di+Ci]×)B_i = \begin{pmatrix} 0 & (C_i - D_i)^T \\ D_i - C_i & [D_i + C_i]_\times \end{pmatrix}
    3. Compute the symmetric 4×44 \times 4 matrix BB:
       B=∑i=1nBiTBiB = \sum_{i=1}^n B_i^T B_i
    4. Compute the unit eigenvector q=(q0,q1,q2,q3)Tq = (q_0, q_1, q_2, q_3)^T of BB associated with its smallest eigenvalue.
    5. Form the optimal rotation matrix RR:
       R=(q02+q12−q22−q322(q1q2−q0q3)2(q1q3+q0q2)2(q2q1+q0q3)q02−q12+q22−q322(q2q3−q0q1)2(q3q1−q0q2)2(q3q2+q0q1)q02−q12−q22+q32)R = \begin{pmatrix} q_0^2 + q_1^2 - q_2^2 - q_3^2 & 2(q_1 q_2 - q_0 q_3) & 2(q_1 q_3 + q_0 q_2) \\ 2(q_2 q_1 + q_0 q_3) & q_0^2 - q_1^2 + q_2^2 - q_3^2 & 2(q_2 q_3 - q_0 q_1) \\ 2(q_3 q_1 - q_0 q_2) & 2(q_3 q_2 + q_0 q_1) & q_0^2 - q_1^2 - q_2^2 + q_3^2 \end{pmatrix}
    6. return RR.
  7. Knowl 7 — Minimum Variance Reprojection Objective Function for Camera Calibration

    equation

    Under the assumption that image feature detection errors are zero-mean, uncorrelated Gaussian noise proportional to pixel spacing along row and column directions, the minimum variance estimation of camera parameters (m,d)(m, d) corresponds to minimizing the sum of squared discrepancies between observed and predicted pixel locations: F(Ω,ω,m,d)=∑i=1n[(ri′−ri(m,d))2+(ci′−ci(m,d))2]F(\Omega, \omega, m, d) = \sum_{i=1}^n \left[ (r'_i - r_i(m, d))^2 + (c'_i - c_i(m, d))^2 \right] where:

    • Ω={(xi,yi,zi)}i=1n\Omega = \{(x_i, y_i, z_i)\}_{i=1}^n is the set of 3D control points in the world coordinate system.
    • ω={(ri′,ci′)}i=1n\omega = \{(r'_i, c'_i)\}_{i=1}^n is the set of observed digitized pixel coordinates (measured with subpixel accuracy).
    • m=(r0,c0,fu,fv,t1,t2,t3,α,β,γ)m = (r_0, c_0, f_u, f_v, t_1, t_2, t_3, \alpha, \beta, \gamma) denotes the 10 internal and external non-distortion parameters, with rotation angles α,β,γ\alpha, \beta, \gamma.
    • d=(k1,g1,g2,g3,g4)Td = (k_1, g_1, g_2, g_3, g_4)^T represents the 5 radial, decentering, and thin prism distortion parameters.
    • ri(m,d)r_i(m, d) and ci(m,d)c_i(m, d) are the predicted row and column pixel coordinates obtained by mapping (xi,yi,zi)(x_i, y_i, z_i) through the complete distortion camera model.
  8. Knowl 8 — Central Control Point Constraint for Stable Non-Distortion Parameter Estimation

    model/method

    In uncalibrated cameras with optical distortions, fitting a distortion-free linear model across the entire field of view leads to severe biases in estimated focal lengths and camera pose. In non-linear optimization, starting from an unconstrained linear fit across all points frequently causes divergence or traps the solver in false local minima.

    To provide a high-quality initial estimate:

    1. A subset of control points Ω1\Omega_1 is selected such that their image projections lie near the optical center (within a circle or rectangle with radius/side length equal to one-quarter of the image dimension).
    2. Because radial and tangential optical distortions vanish towards the principal point (δu,δv→0\delta_u, \delta_v \to 0 as u,v→0u, v \to 0), the assumption d=0d = 0 holds accurately on Ω1\Omega_1.
    3. Closed-form linear calibration and the first step of non-linear optimization are performed exclusively on Ω1\Omega_1.

    Synthetic simulations demonstrate that applying the near-center constraint reduces the initial linear residual positional error μ′\mu' by a factor of 10 (from 0.001693 to 0.000174) and halves the post-nonlinear optimization residual μ′\mu' (from 0.000330 to 0.000146), preventing divergence in iterative refinement.

  9. Knowl 9 — Calibration Accuracy Comparison Across Distortion Models on Telephoto and Wide-Angle Lenses

    data/table

    Experiments using a pair of CCD cameras (512×480512 \times 480 pixels) calibrated with a precision target plate at multiple depths demonstrate the comparative triangulation accuracy of distortion-free, radial-only, and complete (radial + tangential) distortion models for a telephoto lens (f=25 mmf = 25\text{ mm}, diagonal FOV ≈23∘\approx 23^\circ) and a wide-angle lens (f=8.5 mmf = 8.5\text{ mm}, diagonal FOV ≈64∘\approx 64^\circ).

    Lens Model Setup Mean 3D Error M1M_1 (mm) Lateral Error M2M_2 (mm) Depth Ratio M3M_3 NSCE
    f=25 mmf = 25\text{ mm} Telelens
    Linear (Distortion-Free) 0.6436 0.2391 1653.4 1.7478
    Nonlinear (Complete Distortion Model) 0.4365 0.2054 2256.9 1.0598
    f=8.5 mmf = 8.5\text{ mm} Wide-Angle
    Linear (Distortion-Free) 3.7574 0.6773 87.34 2.9952
    Nonlinear (Distortion-Free, d=0d=0) 2.1910 0.9098 170.30 3.9793
    Nonlinear (Radial Distortion Only) 0.5811 0.3149 722.86 1.4081
    Nonlinear (Complete Distortion Model) 0.4653 0.2599 897.46 1.1627

    Here, M1M_1 is the average 3D position error of triangulated test points, M2M_2 is the average lateral error in the xx-yy plane, M3M_3 is the ratio of average depth to absolute depth error, and NSCE\text{NSCE} is the Normalized Stereo Calibration Error.

    Key conclusions:

    1. For the f=25 mmf = 25\text{ mm} telelens, incorporating the complete distortion model reduces NSCE\text{NSCE} from 1.74781.7478 to 1.05981.0598 (within 6% of the theoretical digitization noise limit).
    2. For the f=8.5 mmf = 8.5\text{ mm} wide-angle lens, attempting nonlinear optimization without distortion parameters causes severe parameter distortion (NSCE=3.9793\text{NSCE} = 3.9793). Symmetrical radial distortion correction reduces NSCE\text{NSCE} to 1.40811.4081, while the addition of tangential/decentering/thin-prism distortion further reduces NSCE\text{NSCE} to 1.16271.1627.
  10. Knowl 10 — Independence and Relative Magnitude of Radial and Tangential Distortion Components

    empirical result

    In real wide-angle camera systems (f=8.5 mmf = 8.5\text{ mm}, 512×480512 \times 480 pixel CCD sensors), optical distortions exhibit distinct magnitudes and decoupled parameter behavior between radial and tangential terms:

    1. Disparity in magnitude: Symmetrical radial distortion causes a maximum displacement of approximately 23 pixels at the image borders, whereas decentering and thin-prism distortions (which produce tangential displacement) account for a maximum displacement of approximately 3 pixels (a factor of 5 to 7 smaller).
    2. Parameter stability across models: The estimated radial distortion coefficient k1k_1 remains virtually unchanged between a radial-only model (k1≈0.1742k_1 \approx 0.1742 for camera 1, 0.16900.1690 for camera 2) and the complete model (k1≈0.1731k_1 \approx 0.1731 for camera 1, 0.17200.1720 for camera 2). This demonstrates that decentering and thin-prism parameters (g1,…,g4g_1, \dots, g_4) do not numerically compensate for radial distortion, confirming their optical orthogonality.
    3. Significance of tangential correction: Despite its smaller pixel magnitude (3 pixels vs. 23 pixels), modeling tangential distortion provides a necessary correction that reduces the Normalized Stereo Calibration Error (NSCE\text{NSCE}) from 1.40811.4081 to 1.16271.1627, bringing the calibration error close to the nominal pixel discretization limit (NSCE≈1\text{NSCE} \approx 1).

Coverage note — None was omitted; all key contributions—distortion formulation, decoupled optimization strategy, closed-form linear solvers, rotation matrix orthogonalization, NSCE error metric, and empirical experimental results on synthetic and real tele/wide-angle lenses—are covered.

References

  1. 1.Y. I. Abdel-Aziz and M. M. Karara, ‘‘Direct linear transformation into object space coordinates in close-range photogrammetry,’’ in Proc. Symp. Close-Range Photogrammetry (Urbana, IL), Jan. 1971, pp. 1-18.
  2. 2.
    1. Bottema and B. Roth, Theoretical Kinematics. Amsterdam: North-Holland, 1979.
  3. 3.D. C. Brown, ‘‘Decentering distortion of lenses,’’ Photogrammetric Eng. Remote Sensing, May 1966, pp. 444-462.
  4. 4.W. Faig, ‘‘Calibration of close-range photogrammetric systems: Mathematical formulation,’’ Photogrammetric Eng. Remote Sensing, vol. 41, no. 12, pp. 1479-1486, Dec. 1975.
  5. 5.
    1. D. Faugeras and M. Hebert, ‘‘A 3-D recognition and positioning algorithm using geometrical matching between primitive surfaces,’’ in Proc. 8th Int. Joint Conf. Artificial Intell. (Karlsruhe, W. Germany), Aug. 1983, pp. 996-1002.
  6. 6.
    1. D. Faugeras and G. Toscani, ‘‘Calibration problem for stereo,’’ in Proc. Int. Conf. Comput. Vision Patt. Recogn. (Miami Beach, FL), June 1986, pp. 15-20.
  7. 7.D. B. Gennery, ‘‘Stereo-camera calibration,’’ in Proc. 10th Image Understanding Workshop, 1979, pp. 101-108.
  8. 8.S. Ganapathy, ‘‘Decomposition of transformation matrices for robot vision,’’ in Proc. IEEE Int. Conf. Robotics Automat. (Atlanta), Mar. 1984, pp. 130-139.
  9. 9.A. Isaguirre, P. Pu, and J. Summers, ‘‘A new development in camera calibration calibrating a pair of mobile cameras,’’ in Proc. IEEE Int. Conf. Robotics Automat. (St. Louis), Mar. 1985, pp. 74-79.
  10. 10.R. K. Lenz and R. Y. Tsai, ‘‘Techniques for calibration of the scale factor and image center for high accuracy 3D machine vision metrology,’’ in Proc. IEEE Int. Conf. Robotics Automat. (Raleigh, NC), Mar. 1987, pp. 68-75.
  11. 11.Manual of Photogrammetry. Amer. Soc. Photogrammetry, 1980, 4th ed.
  12. 12.M. D. Shuster, ‘‘Approximate algorithms for fast optimal attitude computation,’’ in Proc. AIAA Guidance Contr. Spec. Conf. (Palo Alto, CA), Aug. 1978, pp. 88-95.
  13. 13.R. Y. Tsai, ‘‘A versatile camera calibration technique for high-accuracy 3D machine vision metrology using off-the-shelf TV cameras and lenses,’’ IEEE J. Robotics Automat., vol. RA-3, no. 4, pp. 323-344, Aug. 1987.
  14. 14.J. Weng, T. S. Huang, and N. Ahuja, ‘‘Motion and structure from two perspective views: Algorithm, error analysis and error estimation,’’ IEEE Trans. Patt. Anal. Machine Intell., vol. 11, no. 5, pp. 451-476, May. 1989.
  15. 15.K. W. Wong, ‘‘Mathematical formulation and digital analysis in close-range photogrammetry,’’ Photogrammetric Eng. Remote Sensing, vol. 41, no. 11, pp. 1355-1373, Nov. 1975.

Citation

MLA
Weng, J., et al. “Camera Calibration with Distortion Models and Accuracy Evaluation”. IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 14, no. 10, 1992, pp. 965–80, https://doi.org/10.1109/34.159901.
APA
Weng, J., Cohen, P., & Herniou, M. (1992). Camera calibration with distortion models and accuracy evaluation. IEEE Transactions on Pattern Analysis and Machine Intelligence, 14(10), 965–980. https://doi.org/10.1109/34.159901
Chicago
Weng, J., P. Cohen, and M. Herniou. 1992. “Camera Calibration with Distortion Models and Accuracy Evaluation”. IEEE Transactions on Pattern Analysis and Machine Intelligence 14 (10): 965–80. https://doi.org/10.1109/34.159901.
Harvard
Weng, J., Cohen, P. and Herniou, M. (1992) “Camera calibration with distortion models and accuracy evaluation”, IEEE Transactions on Pattern Analysis and Machine Intelligence, 14(10), pp. 965–980. Available at: https://doi.org/10.1109/34.159901.
Vancouver
1. Weng J, Cohen P, Herniou M (1992) Camera calibration with distortion models and accuracy evaluation. IEEE Transactions on Pattern Analysis and Machine Intelligence 14:965–980

BibTeX

@article{Weng_1992, title={Camera calibration with distortion models and accuracy evaluation}, volume={14}, ISSN={0162-8828}, url={http://dx.doi.org/10.1109/34.159901}, DOI={10.1109/34.159901}, number={10}, journal={IEEE Transactions on Pattern Analysis and Machine Intelligence}, publisher={Institute of Electrical and Electronics Engineers (IEEE)}, author={Weng, J. and Cohen, P. and Herniou, M.}, year={1992}, pages={965–980} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF