Feature Detection with Automatic Scale Selection

Tony Lindeberg

article1998IJCV3,106 citations

Develops a foundational scale-space framework for automatic scale selection by identifying local extrema of gamma-normalized Gaussian derivatives, enabling vision algorithms to adaptively detect and localize image features such as blobs, corners, and edges across varying scales without manual parameter tuning.

Listen

The research paper develops a general method for automatic scale selection in image analysis, addressing the core challenge that real-world objects and image structures manifest differently depending on the observation scale. Without explicit mechanisms to choose appropriate local scales, multi-scale representations such as scale-space expand data without indicating which structures or scales matter, limiting the performance of low-level vision modules in autonomous systems operating on unknown scenes.

The work sets out to demonstrate that local maxima over scales of suitably normalized differential descriptors can serve as reliable indicators of characteristic lengths in image data, thereby providing a unified computational mechanism for detecting features such as blobs, junctions, edges, and ridges while adapting processing scales to local structure.

The approach combines theoretical analysis of scaling properties under image rescaling with extensive experiments on both synthetic model patterns and real-world images. Normalized Gaussian derivative operators are applied across a range of scales, and local maxima are detected simultaneously in space and scale; the resulting scale-space extrema are evaluated for several differential invariants, including the normalized Laplacian and Hessian determinants for blobs and the rescaled level-curve curvature for junctions. Discrete implementations ensure consistency with continuous scale-space properties, and two-stage pipelines separate coarse detection from refined localization.

The central findings are that maxima over scales of γ-normalized derivatives reliably reflect the intrinsic size of image structures, with the selected scale proportional to the wavelength or spatial extent of the underlying pattern; that this property holds for a broad class of homogeneous differential invariants and yields stable, intuitively plausible features on complex natural images; that a subsequent localization stage minimizing a normalized residual over scales markedly improves spatial accuracy, especially under noise or diffuse boundaries; and that the method produces scale-invariant strength measures and natural support regions around detected features.

These results imply that vision systems can now operate without externally supplied scale parameters, reducing bias toward structures of particular sizes and enabling more robust handling of size variations, noise, and perspective effects in uncontrolled environments. The approach therefore removes a longstanding bottleneck in constructing fully autonomous feature detectors.

The paper recommends incorporating the scale-selection stage into existing feature pipelines and extending the framework to additional descriptors such as edge and ridge detectors. Further integration with higher-level matching, tracking, or verification processes is needed before decisions about feature significance can be made automatically.

The main limitations are that the method relies on local differential models that may break down near interfering structures or at extreme noise levels, and that ranking of candidates still requires an external threshold on the number of features retained. Confidence in the core scaling property and practical utility is high, given the combination of necessity proofs for the normalization form, closed-form verification on model patterns, and consistent performance across dozens of real and synthetic examples.

No sufficiently relevant recommendations were found.

  • Paper: Distinctive Image Features from Scale-Invariant Keypoints, David G. Lowe (2004). Builds directly upon the source's automatic scale selection methodology to formulate the Difference-of-Gaussian scale-space keypoint detector in SIFT.
  • Paper: SURF: Speeded Up Robust Features, Herbert Bay et al. (2006). Extends the source's principles of scale-space analysis and automated scale detection to construct faster box-filtered Hessian interest points.
Cover for Feature Detection with Automatic Scale Selection

Abstract

The fact that objects in the world appear in different ways depending on the scale of observation has important implications if one aims at describing them. It shows that the notion of scale is of utmost importance when processing unknown measurement data by automatic methods. In their seminal works, Witkin (1983) and Koenderink (1984) proposed to approach this problem by representing image structures at different scales in a so-called scale-space representation. Traditional scale-space theory building on this work, however, does not address the problem of how to select local appropriate scales for further analysis.

This article proposes a systematic methodology for dealing with this problem. A framework is proposed for generating hypotheses about interesting scale levels in image data, based on a general principle stating that local extrema over scales of different combinations of γ-normalized derivatives are likely candidates to correspond to interesting structures. Specifically, it is shown how this idea can be used as a major mechanism in algorithms for automatic scale selection, which adapt the local scales of processing to the local image structure.

Support for the proposed approach is given in terms of a general theoretical investigation of the behaviour of the scale selection method under rescalings of the input pattern and by experiments on real-world and synthetic data. Support is also given by a detailed analysis of how different types of feature detectors perform when integrated with a scale selection mechanism and then applied to characteristic model patterns. Specifically, it is described in detail how the proposed methodology applies to the problems of blob detection, junction detection, edge detection, ridge detection and local frequency estimation.

In many computer vision applications, the poor performance of the low-level vision modules constitutes a major bottle-neck. It will be argued that the inclusion of mechanisms for automatic scale selection is essential if we are to construct vision systems to analyse complex unknown environments.

Table of Contents

  • 1 Introduction
  • 1.1 Outline of the presentation
  • 2 Scale-space representation: Review
  • 3 Normalized derivatives and intuitive idea for scale selection
  • 4 Proposed methodology for scale selection
  • 4.1 General scaling property of local maxima over scales
  • 4.2 *The scale selection mechanism in practice*
  • 4.3 Experiments: Scale-space signatures from real data
  • 4.4 Simultaneous detection of interesting points and scales
  • 5 Blob detection with automatic scale selection
  • 5.1 Analysis of scale-space maxima for idealized model patterns
  • 5.2 Comparisons with fixed-scale blob detection
  • 5.3 Applications of blob detection with automatic scale selection
  • 6 Junction detection with automatic scale selection
  • 6.1 Selection of detection scales from normalized scale-space maxima
  • 6.2 Analysis of scale-space maxima for diffuse junction models
  • 6.3 Experiments: Scale-space signatures in junction detection
  • 7 Feature localization with automatic scale selection
  • 7.1 Corner localization by local consistency
  • 7.2 Automatic selection of localization scales
  • 7.3 Experiments: Choice of localization scale
  • 7.4 Composed scheme for junction detection and localization
  • 7.5 Further experiments
  • 7.6 *Applications of corner detection with automatic scale selection*
  • 7.7 *Extensions of the junction detection method*
  • 7.8 Extensions to edge detection
  • 8 Dense frequency estimation
  • 9 Analysis and interpretation of normalized derivatives
  • 9.1 Interpretation of γ\gamma-normalized derivatives in terms of LpL_p-norms
  • 9.2 Interpretation in terms of self-similar Fourier spectrum
  • 9.3 *Relations to previous work*
  • 10 Summary and discussion
  • 10.1 Technical contributions
  • A Appendix
  • A.1 Necessity of the form of the γ\gamma-parameterized derivative operator
  • A.2 LpL_p-normalization interpretation of γ\gamma-normalized derivatives
  • A.3 Normalized derivative responses to self-similar power spectra
  • A.4 Discrete implementation of the scale selection mechanisms
  • A.4.1 Computing discrete derivative approximations
  • A.4.2 Normalization in the discrete case
  • A.4.3 Detection of scale-space maxima in discrete data
  • A.4.4 Implementing junction localization
  • References

Knowls

  1. Knowl 1 — Normalized Derivatives in Linear Scale-Space

    definition

    Given a continuous signal f:RDRf: \mathbb{R}^D \to \mathbb{R}, its linear scale-space representation L:RD×R+RL: \mathbb{R}^D \times \mathbb{R}_+ \to \mathbb{R} is defined as the solution to the diffusion equation tL=122L=12i=1DxixiL\partial_t L = \frac{1}{2} \nabla^2 L = \frac{1}{2} \sum_{i=1}^D \partial_{x_i x_i} L with initial condition L(;0)=f()L(\cdot; 0) = f(\cdot), or equivalently by convolution with Gaussian kernels L(;t)=g(;t)f()L(\cdot; t) = g(\cdot; t) * f(\cdot) where g(x;t)=1(2πt)D/2e(x12++xD2)/(2t)g(x; t) = \frac{1}{(2\pi t)^{D/2}} e^{-(x_1^2 + \dots + x_D^2)/(2t)} and t0t \ge 0 is the scale parameter (with standard deviation σ=t\sigma = \sqrt{t}). Spatial derivatives of order α=(α1,,αD)\alpha = (\alpha_1, \dots, \alpha_D) are denoted by Lxα=xαLL_{x^\alpha} = \partial_{x^\alpha} L.

    Because derivative amplitudes decrease under scale-space smoothing, the γ\gamma-normalized derivative operator is defined by ξi=tγ/2xi\partial_{\xi_i} = t^{\gamma/2} \partial_{x_i} corresponding to the coordinate transformation ξi=xi/tγ/2\xi_i = x_i / t^{\gamma/2}, where γ[0,1]\gamma \in [0, 1] is a scale normalization parameter. For a spatial derivative of order m=α=i=1Dαim = |\alpha| = \sum_{i=1}^D \alpha_i, the γ\gamma-normalized derivative is Lξα=tmγ/2Lxα.L_{\xi^\alpha} = t^{m\gamma / 2} L_{x^\alpha}. When γ=1\gamma = 1, the normalized derivative operator is dimensionless and perfectly scale invariant.

  2. Knowl 2 — Scale Covariance of Homogeneous Normalized Differential Invariants

    theoretical result

    Let two continuous signals ff and ff' be related by a spatial scaling factor s>0s > 0 such that f(x)=f(sx)f(x) = f'(sx), with scale-space representations L(x;t)L(x; t) and L(x;t)L'(x'; t') related under x=sxx' = sx and t=s2tt' = s^2 t by L(x;t)=L(sx;s2t)L(x; t) = L'(sx; s^2 t). For any homogeneous polynomial differential invariant DLD L of total differentiation order MM defined as DL=i=1Icij=1JLxαij,where j=1Jαij=Mi,D L = \sum_{i=1}^I c_i \prod_{j=1}^J L_{x^{\alpha_{ij}}}, \quad \text{where } \sum_{j=1}^J |\alpha_{ij}| = M \quad \forall i, the corresponding γ\gamma-normalized differential expression Dγ-normL=tMγ/2DLD_{\gamma\text{-norm}} L = t^{M\gamma/2} D L transforms under spatial rescaling as Dγ-normL(x;t)=sM(1γ)Dγ-normL(sx;s2t).D_{\gamma\text{-norm}} L(x; t) = s^{M(1-\gamma)} D'_{\gamma\text{-norm}} L'(sx; s^2 t).

    Local extrema over scale are preserved under this scaling transformation for any γ\gamma: t(Dγ-normL)(x0;t0)=0    t(Dγ-normL)(sx0;s2t0)=0.\partial_t (D_{\gamma\text{-norm}} L)(x_0; t_0) = 0 \iff \partial_{t'} (D'_{\gamma\text{-norm}} L')(s x_0; s^2 t_0) = 0. When γ1\gamma \neq 1, a scale-compensated magnitude measure invariant to signal scaling is obtained by multiplying by tM(1γ)/2t^{M(1-\gamma)/2}: Mγ-normL=tM(1γ)/2Dγ-normL=tM/2DL,M_{\gamma\text{-norm}} L = t^{M(1-\gamma)/2} D_{\gamma\text{-norm}} L = t^{M/2} D L, which satisfies Mγ-normL(x;t)=Mγ-normL(sx;s2t)M_{\gamma\text{-norm}} L(x; t) = M'_{\gamma\text{-norm}} L'(sx; s^2 t).

  3. Knowl 3 — Principle of Automatic Scale Selection via Scale-Space Extrema

    model/method

    The scale selection principle states that in the absence of prior knowledge, a scale level t0t_0 at which a γ\gamma-normalized differential descriptor DnormLD_{\text{norm}} L assumes a local maximum over scale reflects the characteristic spatial extent (or characteristic wavelength) of the corresponding structure in the signal.

    To simultaneously identify feature points and their characteristic scales, interest points are defined as normalized scale-space extrema (x0;t0)RD×R+(x_0; t_0) \in \mathbb{R}^D \times \mathbb{R}_+, which satisfy the simultaneous spatial and scale extremum conditions: (DnormL)(x0;t0)=0,t(DnormL)(x0;t0)=0.\nabla (D_{\text{norm}} L)(x_0; t_0) = 0, \quad \partial_t (D_{\text{norm}} L)(x_0; t_0) = 0. Under an input rescaling f(x)=f(sx)f(x) = f'(sx), any scale-space extremum (x0;t0)(x_0; t_0) in LL maps directly to (sx0;s2t0)(s x_0; s^2 t_0) in LL', ensuring that feature detection and scale selection commute with uniform size variations.

  4. Knowl 4 — Necessity of Power-Law Normalization for Scale-Invariant Scale Selection

    theoretical result

    Let an mm-th order normalized derivative operator at scale t>0t > 0 be defined as ξm=ϕ(t)xm\partial_{\xi^m} = \phi(t) \partial_{x^m}, where ϕ:R+R\phi: \mathbb{R}_+ \to \mathbb{R} is a smooth multiplicative factor depending solely on scale. If the scale selection mechanism based on detecting local maxima over scale is required to commute with spatial rescalings—such that for f(x)=f(sx)f(x) = f'(sx), a scale maximum at (x0;t0)(x_0; t_0) in ff guarantees a corresponding scale maximum at (sx0;s2t0)(sx_0; s^2 t_0) in ff'—then ϕ(t)\phi(t) must necessarily be a power function: ϕ(t)=AtB\phi(t) = A t^B for constants AA and BB. Setting B=mγ/2B = m\gamma / 2 yields the γ\gamma-normalized derivative formulation ξm=tmγ/2xm\partial_{\xi^m} = t^{m\gamma/2} \partial_{x^m}.

  5. Knowl 5 — Blob Detection with Automatic Scale Selection

    model/method

    Two-dimensional blob features are detected as points (x0,y0;t0)(x_0, y_0; t_0) that are simultaneous local maxima in space and scale of normalized differential invariants with γ=1\gamma = 1:

    1. Normalized Laplacian of Gaussian: norm2L=t2L=t(Lxx+Lyy).\nabla^2_{\text{norm}} L = t \nabla^2 L = t (L_{xx} + L_{yy}). For an isotropic 2D Gaussian blob f(x1,x2)=g(x1;t0)g(x2;t0)f(x_1, x_2) = g(x_1; t_0) g(x_2; t_0), the amplitude norm2L(0,0;t)=t(2t0+2t)2π(t0+t)3|\nabla^2_{\text{norm}} L(0, 0; t)| = \frac{t(2t_0 + 2t)}{2\pi (t_0 + t)^3} attains its unique maximum over scale at t=t0t = t_0. For a 2D sum of perpendicular sine waves f(x,y)=sin(ω0x)+sin(ω0y)f(x, y) = \sin(\omega_0 x) + \sin(\omega_0 y), the maximum over scale is attained at t=2/ω02t = 2 / \omega_0^2.

    2. Normalized Determinant of the Hessian Matrix: detHnormL=t2(LxxLyyLxy2).\det \mathcal{H}_{\text{norm}} L = t^2 (L_{xx} L_{yy} - L_{xy}^2). For an elliptical Gaussian blob f(x1,x2)=g(x1;t1)g(x2;t2)f(x_1, x_2) = g(x_1; t_1) g(x_2; t_2), the maximum over scale of detHnormL(0,0;t)=t24π2(t1+t)2(t2+t)2|\det \mathcal{H}_{\text{norm}} L(0, 0; t)| = \frac{t^2}{4\pi^2 (t_1 + t)^2 (t_2 + t)^2} occurs at t=t1t2t = \sqrt{t_1 t_2}. For perpendicular sine waves f(x,y)=sin(ω1x)+sin(ω2y)f(x, y) = \sin(\omega_1 x) + \sin(\omega_2 y), the maximum over scale occurs at t=4ω12+ω22t = \frac{4}{\omega_1^2 + \omega_2^2}.

    The characteristic radius of the detected blob is proportional to σ0=t0\sigma_0 = \sqrt{t_0}.

  6. Knowl 6 — Junction Detection via Normalized Rescaled Level Curve Curvature

    model/method

    Junctions (corners) are detected in 2D image scale-space as scale-space maxima of the normalized rescaled level curve curvature κ~norm2\tilde{\kappa}_{\text{norm}}^2, with γ=1\gamma = 1: κ~norm=t2γκ~=t2(Lx22Lx1x12Lx1Lx2Lx1x2+Lx12Lx2x2).\tilde{\kappa}_{\text{norm}} = t^{2\gamma} \tilde{\kappa} = t^2 (L_{x_2}^2 L_{x_1 x_1} - 2 L_{x_1} L_{x_2} L_{x_1 x_2} + L_{x_1}^2 L_{x_2 x_2}). Here κ~=Lv2Luu\tilde{\kappa} = L_v^2 L_{uu} corresponds to the level curve curvature multiplied by the gradient magnitude cubed, which constitutes the lowest-order polynomial differential invariant for corners whose spatial extrema are affine invariant.

    For a diffuse L-junction model f(x1,x2)=Φ(x1;t0)Φ(x2;t0)f(x_1, x_2) = \Phi(x_1; t_0) \Phi(x_2; t_0) (where Φ(xi;t0)=xig(x;t0)dx\Phi(x_i; t_0) = \int_{-\infty}^{x_i} g(x'; t_0) dx' has edge diffuseness t0t_0) analyzed with γ(0,1)\gamma \in (0, 1), the central scale response κ~norm(0,0;t)=t2γ8π2(t0+t)2|\tilde{\kappa}_{\text{norm}}(0, 0; t)| = \frac{t^{2\gamma}}{8\pi^2 (t_0 + t)^2} attains a unique maximum over scale at tκ~=γ1γt0t_{\tilde{\kappa}} = \frac{\gamma}{1-\gamma} t_0. For a finite-extent junction or non-uniform Gaussian blob L(x1,x2;t)=g(x1;t0+t)g(x2;t0+t)L(x_1, x_2; t) = g(x_1; t_0+t) g(x_2; t_0+t) evaluated with γ=1\gamma = 1, the spatial maximum over scale is attained at tκ~=23t0t_{\tilde{\kappa}} = \frac{2}{3} t_0.

    Candidate junction points (x0,y0;t0)(x_0, y_0; t_0) extracted as 3D scale-space maxima capture the spatial extent and diffuseness of corner structures.

  7. Knowl 7 — Composed Two-Stage Junction Detection and Localization Algorithm

    algorithm

    Junction detection and localization is formulated as a two-stage procedure where detection at coarse scales determines feature existence and region of support, followed by fine-scale feature localization via minimization of a normalized residual over scale:

    Input: Discrete 2D image ff, scale range [tmin,tmax][t_{\min}, t_{\max}], number of scale levels KK, candidate count NN, max iterations Imax=3I_{\max} = 3
    Output: Set of localized junction coordinates xx^* with detection scales t0t_0 and localization scales tloct_{\text{loc}}
    1. Scale-Space Representation:
       Compute scale-space slices L(;tk)L(\cdot; t_k) for k=1,,Kk = 1, \dots, K with logarithmic spacing tk+1/tk=constt_{k+1}/t_k = \text{const}.
       Compute normalized derivatives and form κ~norm2(x;tk)=tk4(Lx22Lx1x12Lx1Lx2Lx1x2+Lx12Lx2x2)2\tilde{\kappa}_{\text{norm}}^2(x; t_k) = t_k^4 (L_{x_2}^2 L_{x_1 x_1} - 2 L_{x_1} L_{x_2} L_{x_1 x_2} + L_{x_1}^2 L_{x_2 x_2})^2.
    2. Stage 1 - Detection:
       Detect 3D local maxima (x0,y0;t0)(x_0, y_0; t_0) of κ~norm2\tilde{\kappa}_{\text{norm}}^2 over 26 space-scale neighbors.
       Sort by saliency t0κ~norm2t_0 \cdot \tilde{\kappa}_{\text{norm}}^2 and retain the NN strongest candidates.
    3. Stage 2 - Localization and Scale Selection:
       for each candidate (x0;t0)(x_0; t_0) do
          Initialize position estimate x^x0\hat{x} \leftarrow x_0.
          Set Gaussian window wx0(x)=g(xx0;t0)w_{x_0}(x') = g(x' - x_0; t_0) with integration scale t0t_0.
          for iteration =1= 1 to ImaxI_{\max} do
             for a set of localization scales t[tlow,t0]t \in [t_{\text{low}}, t_0] (with tlow0.01t_{\text{low}} \approx 0.01) do
                Compute gradient statistics at scale tt within window wx0w_{x_0}:
                   A(t)=R2(L)(x;t)(L)T(x;t)wx0(x)dxA(t) = \int_{\mathbb{R}^2} (\nabla L)(x'; t) (\nabla L)^T(x'; t) w_{x_0}(x') dx'
                   b(t)=R2(L)(x;t)(L)T(x;t)xwx0(x)dxb(t) = \int_{\mathbb{R}^2} (\nabla L)(x'; t) (\nabla L)^T(x'; t) x' w_{x_0}(x') dx'
                   c(t)=R2xT(L)(x;t)(L)T(x;t)xwx0(x)dxc(t) = \int_{\mathbb{R}^2} x'^T (\nabla L)(x'; t) (\nabla L)^T(x'; t) x' w_{x_0}(x') dx'
                Compute normalized residual:
                   d~min(t)=c(t)b(t)TA(t)1b(t)trace(A(t))\tilde{d}_{\min}(t) = \frac{c(t) - b(t)^T A(t)^{-1} b(t)}{\text{trace}(A(t))}
             Select optimal localization scale tloc=argmintd~min(t)t_{\text{loc}} = \arg\min_t \tilde{d}_{\min}(t).
             Update position estimate: xnewA(tloc)1b(tloc)x_{\text{new}} \leftarrow A(t_{\text{loc}})^{-1} b(t_{\text{loc}}).
             if x^xnew<1\|\hat{x} - x_{\text{new}}\| < 1 pixel then
                x^xnew\hat{x} \leftarrow x_{\text{new}}
                break
             x^xnew\hat{x} \leftarrow x_{\text{new}}
          if x^x0>t0\|\hat{x} - x_0\| > \sqrt{t_0} then
             Discard candidate as diverged
          else
             Record localized junction x=x^x^* = \hat{x} with detection scale t0t_0 and localization scale tloct_{\text{loc}}
  8. Knowl 8 — Dense Frequency Estimation via Quasi-Quadrature Differential Invariant

    model/method

    To perform dense, phase-independent local frequency estimation within the scale-space framework without non-local Hilbert transforms, a normalized quasi-quadrature entity QLQL (with γ=1\gamma = 1) is defined.

    In 1D, the entity is defined by QL=Lξ2+CLξξ2=tLx2+Ct2Lxx2QL = L_\xi^2 + C L_{\xi\xi}^2 = t L_x^2 + C t^2 L_{xx}^2 where CC is a weighting constant. For a sinusoidal signal f(x)=sin(ω0x)f(x) = \sin(\omega_0 x), (QL)(x;t)=tω02eω02t(1+(Ctω021)sin2(ω0x)).(QL)(x; t) = t \omega_0^2 e^{-\omega_0^2 t} \left( 1 + (C t \omega_0^2 - 1) \sin^2(\omega_0 x) \right). Spatial oscillations in QLQL vanish when t1/(Cω02)t \to 1/(C \omega_0^2):

    • C=2/30.6667C = 2/3 \approx 0.6667 produces the most symmetric spatial variation of the scale maximizing QLQL, where tQLt_{QL} varies between 1/ω021/\omega_0^2 at ω0x=0\omega_0 x = 0 and 2/ω022/\omega_0^2 at ω0x=π/2\omega_0 x = \pi/2, satisfying tQLω0x=π/4=12(tQL0+tQLπ/2)t_{QL}|_{\omega_0 x = \pi/4} = \frac{1}{2}(t_{QL}|_0 + t_{QL}|_{\pi/2}).
    • C=e/40.6796C = e/4 \approx 0.6796 minimizes ripple in the maximum response magnitude across spatial positions.

    In 2D, the rotationally symmetric quasi-quadrature entity is defined by QL=t(Lx2+Ly2)+Ct2(Lxx2+2Lxy2+Lyy2).QL = t (L_x^2 + L_y^2) + C t^2 (L_{xx}^2 + 2 L_{xy}^2 + L_{yy}^2). Selecting local maxima over scale of QL(x,y;t)QL(x, y; t) provides a dense estimate of local dominant frequency ω01/tmax\omega_0 \approx 1/\sqrt{t_{\max}}.

  9. Knowl 9 — Lp-Norm Interpretation and Spectral Neutrality of Normalized Derivatives

    theoretical result

    The γ\gamma-normalized derivative concept ξ=tγ/2x\partial_\xi = t^{\gamma/2} \partial_x relates to kernel LpL_p-norms and self-similar power spectra as follows:

    1. LpL_p-Norm Invariance Across Scale: For an mm-th order normalized Gaussian derivative kernel gξm(;t)=tmγ/2xmg(;t)g_{\xi^m}(\cdot; t) = t^{m\gamma/2} \partial_{x^m} g(\cdot; t) in DD dimensions, its LpL_p-norm satisfies gξm(;t)p=t12[m(γ1)+D(1/p1)]gξm(;1)p.\|g_{\xi^m}(\cdot; t)\|_p = t^{\frac{1}{2}[m(\gamma - 1) + D(1/p - 1)]} \|g_{\xi^m}(\cdot; 1)\|_p. This norm is constant over scale tt if and only if p=11+mD(1γ).p = \frac{1}{1 + \frac{m}{D}(1 - \gamma)}. When γ=1\gamma = 1, p=1p = 1 for all derivative orders mm and dimensions DD, proving that γ=1\gamma = 1 normalization is uniquely equivalent to L1L_1-normalization of Gaussian derivative kernels.

    2. Neutrality to Self-Similar Spectra: For a DD-dimensional signal with self-similar power spectrum Sf(ω)=ω2βS_f(\omega) = |\omega|^{-2\beta}, the total differential energy of mm-th order γ\gamma-normalized derivatives Em(t)=RDα=mtmγLxα2dxE_m(t) = \int_{\mathbb{R}^D} \sum_{|\alpha|=m} t^{m\gamma} |L_{x^\alpha}|^2 dx scales as Em(t)tβD/2m(1γ).E_m(t) \sim t^{\beta - D/2 - m(1-\gamma)}. Em(t)E_m(t) is scale-invariant if and only if β=D2+m(1γ)\beta = \frac{D}{2} + m(1 - \gamma). For natural images with scale-invariant power spectra Sf(ω)ωDS_f(\omega) \sim |\omega|^{-D} in D=2D = 2 dimensions, the energy of γ=1\gamma = 1 normalized derivatives is scale-neutral for all orders mm.

  10. Knowl 10 — Discrete Implementation and l1-Normalization of Scale-Space Derivatives

    model/method

    To preserve scale-space properties on discrete grids without artificial zero responses at scale t=0t = 0:

    1. Discrete Scale-Space Smoothing: A discrete image f(x,y)f(x, y) is smoothed using the discrete scale-space kernel T(x,y;t)=T1(x;t)T1(y;t)T(x, y; t) = T_1(x; t) T_1(y; t), where T1(n;t)=etIn(t)T_1(n; t) = e^{-t} I_n(t) and In(t)I_n(t) is the modified Bessel function of integer order nn.

    2. Discrete Difference Operators: Spatial derivatives are approximated by central differences: δxL(x,y;t)=12(L(x+1,y;t)L(x1,y;t))\delta_x L(x, y; t) = \frac{1}{2}(L(x+1, y; t) - L(x-1, y; t)) δxxL(x,y;t)=L(x+1,y;t)2L(x,y;t)+L(x1,y;t).\delta_{xx} L(x, y; t) = L(x+1, y; t) - 2L(x, y; t) + L(x-1, y; t).

    3. Discrete l1l_1-Normalization: Discrete first-order derivatives are multiplied by the normalization factor α1(t)=2π(T1(0;t)+T1(1;t)).\alpha_1(t) = \frac{\sqrt{2}}{\sqrt{\pi}(T_1(0; t) + T_1(1; t))}. As tt \to \infty, α1(t)t\alpha_1(t) \to \sqrt{t}, matching the continuous factor. At t=0t = 0, α1(0)=2/π0.798>0\alpha_1(0) = \sqrt{2/\pi} \approx 0.798 > 0, avoiding the zero-derivative artifact of the continuous factor t\sqrt{t} on unsmoothed discrete data.

    4. Space-Scale Extremum Detection: 3D scale-space maxima are identified on a logarithmically distributed set of scale levels tkt_k by comparing each grid point (x,y,tk)(x, y, t_k) to all 26 adjacent space-scale neighbors.

  11. Knowl 11 — Junction Localization Accuracy Under Additive Noise

    data/table

    Applying scale selection via minimization of the normalized residual d~min(t)=cbTA1btraceA\tilde{d}_{\min}(t) = \frac{c - b^T A^{-1} b}{\text{trace} A} to localize a synthetic TT-junction (with 9090^\circ opening angles) initialized with a 10-pixel spatial offset demonstrates how the selected localization scale adapts to additive white Gaussian noise:

    Noise level Selected scale Normalized residual Absolute error Reference scale Absolute error at ref.
    0.01 1.2 1.0 0.9 1.0 0.9
    0.03 2.8 2.3 1.4 2.1 1.4
    0.10 8.1 6.7 5.1 6.3 4.9
    0.30 18.8 13.9 16.6 13.4 14.7
    1.00 50.3 28.5 116.0 28.2 96.3

    Here, "Noise level" is the standard deviation of the added noise, "Selected scale" is tdmin=argmintd~min(t)t_{d_{\min}} = \arg\min_t \tilde{d}_{\min}(t), "Normalized residual" is d~min(tdmin)\tilde{d}_{\min}(t_{d_{\min}}), "Absolute error" is the Euclidean distance from the true junction after one iteration at tdmint_{d_{\min}}, "Reference scale" is the scale tabst_{\text{abs}} achieving the minimum absolute error across all evaluated scales, and "Absolute error at ref." is the absolute error at tabst_{\text{abs}}.

    The data show that as the noise level increases, both the selected scale and the optimal localization scale increase monotonically to filter out noise, and the error obtained at the automatically selected scale closely matches the minimal achievable error at the reference scale.

Coverage note — Detailed algorithms and theoretical models for edge detection and ridge detection with automatic scale selection are summarized in Table 2 and Section 7.8 but their full development is explicitly deferred to companion papers (Lindeberg 1996a, 1996b).

References

  1. 1.M. Abramowitz and I. A. Stegun, editors. Handbook of Mathematical Functions. Applied Mathematics Series. National Bureau of Standards, 55 edition, 1964.
  2. 2.A. Almansa and T. Lindeberg. “Enhancement of Fingerprint Images by Shape-Adapted Scale-Space Operators”. In J. Sporring; M. Nielsen; L. Florack, and P. Johansen, editors, Gaussian Scale-Space Theory: Proc. PhD School on Scale-Space Theory, Copenhagen, Denmark, May. 1996. Kluwer Academic Publishers.
  3. 3.A. Almansa and T. Lindeberg. “Fingerprint Enhancement by Shape Adaptation of Scale-Space Operators with Automatic Scale-Selection”. Technical Report ISRN KTH/NA/P--98/03--SE, Dept. of Numerical Analysis and Computing Science, KTH, Stockholm, Sweden, Apr. 1998.
  4. 4.J. Babaud; A. P. Witkin; M. Baudin, and R. O. Duda. “Uniqueness of the Gaussian Kernel for Scale-Space Filtering”. IEEE Trans. Pattern Analysis and Machine Intell., 8(1):26–33, 1986.
  5. 5.J. Blom. Topological and Geometrical Aspects of Image Structure. PhD thesis. , Dept. Med. Phys. Physics, Univ. Utrecht, NL-3508 Utrecht, Netherlands, 1992.
  6. 6.D. Blostein and N. Ahuja. “A Multiscale Region Detector”. Computer Vision, Graphics, and Image Processing, 45:22–41, 1989.
  7. 7.L. Bretzner and T. Lindeberg. “Feature Tracking with Automatic Selection of Spatial Scales”. Technical Report ISRN KTH/NA/P--96/21--SE, Dept. of Numerical Analysis and Computing Science, KTH, Stockholm, Sweden, Jun. 1996.
  8. 8.L. Bretzner and T. Lindeberg. “On the Handling of Spatial and Temporal Scales in Feature Tracking”. In B. M. ter Haar Romeny; L. M. J. Florack; J. J. Koenderink, and M. A. Viergever, editors, Scale-Space Theory in Computer Vision: Proc. First Int. Conf. Scale-Space’97, volume 1252 of Lecture Notes in Computer Science, pages 128–139, Utrecht, The Netherlands, July 1997. Springer Verlag, New York.
  9. 9.L. Bretzner and T. Lindeberg. “Feature tracking with automatic selection of spatial scales”. Computer Vision and Image Understanding, 1998. (To appear).
  10. 10.K. Brunnström; T. Lindeberg, and J.-O. Eklundh. “Active detection and classification of junctions by foveation with a head-eye system guided by the scale-space primal sketch”. In G. Sandini, editor, Proc. 2nd European Conf. on Computer Vision, volume 588 of Lecture Notes in Computer Science, pages 701–709, Santa Margherita Ligure, Italy, May. 1992. Springer-Verlag.
  11. 11.P. J. Burt. “Fast Filter Transforms for Image Processing”. Computer Vision, Graphics, and Image Processing, 16:20–51, 1981.
  12. 12.J. L. Crowley and A. C. Parker. “A Representation for Shape Based on Peaks and Ridges in the Difference of Low-Pass Transform”. IEEE Trans. Pattern Analysis and Machine Intell., 6(2):156–170, 1984.
  13. 13.J. L. Crowley. A Representation for Visual Information. PhD thesis. , Carnegie-Mellon University, Robotics Institute, Pittsburgh, Pennsylvania, 1981.
  14. 14.R. Deriche and G. Giraudon. “Accurate Corner Detection: An Analytical Study”. In Proc. 3rd Int. Conf. on Computer Vision, pages 66–70, Osaka, Japan, 1990.
  15. 15.L. Dreschler and H.-H. Nagel. “Volumetric Model and 3D-Trajectory of a Moving Car Derived from Monocular TV-Frame Sequences of a Street Scene”. Computer Vision, Graphics, and Image Processing, 20(3):199–228, 1982.
  16. 16.D. J. Field. “Relations between the statistics of natural images and the response properties of cortical cells”. J. of the Optical Society of America, 4:2379–2394, 1987.
  17. 17.L. M. J. Florack. The Syntactical Structure of Scalar Images. PhD thesis. , Dept. Med. Phys. Physics, Univ. Utrecht, NL-3508 Utrecht, Netherlands, 1993.
  18. 18.L. M. J. Florack; B. M. ter Haar Romeny; J. J. Koenderink, and M. A. Viergever. “Scale and the Differential Structure of Images”. Image and Vision Computing, 10(6):376–388, Jul. 1992.
  19. 19.L. M. J. Florack; B. M. ter Haar Romeny; J. J. Koenderink, and M. A. Viergever. “Linear scale-space”. J. of Mathematical Imaging and Vision, 4(4):325–351, 1994.
  20. 20.W. A. Förstner and E. Gülch. “A Fast Operator for Detection and Precise Location of Distinct Points, Corners and Centers of Circular Features”. In Proc. Intercommission Workshop of the Int. Soc. for Photogrammetry and Remote Sensing, Interlaken, Switzerland, 1987.
  21. 21.J. Gårding and T. Lindeberg. “Direct computation of shape cues using scale-adapted spatial derivative operators”. Int. J. of Computer Vision, 17(2):163–191, 1996.
  22. 22.B. ter Haar Romeny, editor. Geometry-Driven Diffusion in Computer Vision. Series in Mathematical Imaging and Vision. Kluwer Academic Publishers, Dordrecht, Netherlands, 1994.
  23. 23.P. Johansen; S. Skelboe; K. Grue, and J. D. Andersen. “Representing Signals by Their Top Points in Scale-Space”. In Proc. 8:th Int. Conf. on Pattern Recognition, pages 215–217, Paris, France, Oct. 1986.
  24. 24.L. Kitchen and A. Rosenfeld. “Gray-Level Corner Detection”. Pattern Recognition Letters, 1(2):95–102, 1982.
  25. 25.J. J. Koenderink and W. Richards. “Two-Dimensional Curvature Operators”. J. of the Optical Society of America, 5:7:1136–1141, 1988.
  26. 26.J. J. Koenderink and A. J. van Doorn. “Receptive Field Families”. Biological Cybernetics, 63:291–298, 1990.
  27. 27.J. J. Koenderink and A. J. van Doorn. “Generic neighborhood operators”. IEEE Trans. Pattern Analysis and Machine Intell., 14(6):597–605, Jun. 1992.
  28. 28.J. J. Koenderink. “The structure of images”. Biological Cybernetics, 50:363–370, 1984.
  29. 29.A. F. Korn. “Toward a Symbolic Representation of Intensity Changes in Images”. IEEE Trans. Pattern Analysis and Machine Intell., 10(5):610–625, 1988.
  30. 30.T. Lindeberg and J. Gårding. “Shape from Texture from a Multi-Scale Perspective”. In H.-H. Nagel et. al., editor, Proc. 4th Int. Conf. on Computer Vision, pages 683–691, Berlin, Germany, May. 1993. IEEE Computer Society Press.
  31. 31.T. Lindeberg and J. Gårding. “Shape-adapted smoothing in estimation of 3-D depth cues from affine distortions of local 2-D structure”. Image and Vision Computing, 15:415–434, 1997.
  32. 32.T. Lindeberg and M. Li. “Segmentation and classification of edges using minimum description length approximation and complementary junction cues”. In G. Borgefors, editor, Proc. 9th Scandinavian Conference on Image Analysis, pages 767–776, Uppsala, Sweden, June 1995. Swedish Society for Automated Image Processing.
  33. 33.T. Lindeberg and M. Li. “Segmentation and classification of edges using minimum description length approximation and complementary junction cues”. Computer Vision and Image Understanding, 67(1):88–98, 1997.
  34. 34.T. Lindeberg and G. Olofsson. “The Aspect Feature Graph in Recognition by Parts”. draft manuscript, 1995.
  35. 35.T. Lindeberg. “Scale-Space for Discrete Signals”. IEEE Trans. Pattern Analysis and Machine Intell., 12(3):234–254, Mar. 1990.
  36. 36.T. Lindeberg. Discrete Scale-Space Theory and the Scale-Space Primal Sketch. Ph. D. dissertation. Ph. D. dissertation, Dept. of Numerical Analysis and Computing Science, KTH, Stockholm, Sweden, May. 1991. ISRN KTH/NA/P--91/08--SE. An extended and revised version published as book ”Scale-Space Theory in Computer Vision” in The Kluwer International Series in Engineering and Computer Science.
  37. 37.T. Lindeberg. “Detecting salient blob-like image structures and their scales with a scale-space primal sketch: A method for focus-of-attention”. Int. J. of Computer Vision, 11(3):283–318, Dec. 1993.
  38. 38.T. Lindeberg. “On Scale Selection for Differential Operators”. In K. Heia K. A. Høgdra, B. Braathen, editor, Proc. 8th Scandinavian Conf. on Image Analysis, pages 857–866, Tromsø, Norway, May. 1993. Norwegian Society for Image Processing and Pattern Recognition.
  39. 39.T. Lindeberg. “Junction detection with automatic selection of detection scales and localization scales”. In Proc. 1st International Conference on Image Processing, volume I, pages 924–928, Austin, Texas, Nov. 1994. IEEE Computer Society Press.
  40. 40.T. Lindeberg. “On the Axiomatic Foundations of Linear Scale-Space: Combining Semi-Group Structure with Causality vs. Scale Invariance”. Technical Report ISRN KTH/NA/P--94/20--SE, Dept. of Numerical Analysis and Computing Science, KTH, Stockholm, Sweden, Aug. 1994. Extended version to appear in J. Sporring and M. Nielsen and L. Florack and P. Johansen (eds.) Gaussian Scale-Space Theory: Proc. PhD School on Scale-Space Theory, Copenhagen, Denmark, Kluwer Academic Publishers, May 1996.
  41. 41.T. Lindeberg. “Scale Selection for Differential Operators”. Technical Report ISRN KTH/NA/P--94/03--SE, Dept. of Numerical Analysis and Computing Science, KTH, Stockholm, Sweden, Jan. 1994. Extended version in Int. J. of Computer Vision, vol 30, number 2, 1998. (In press).
  42. 42.T. Lindeberg. Scale-Space Theory in Computer Vision. The Kluwer International Series in Engineering and Computer Science. Kluwer Academic Publishers, Dordrecht, Netherlands, 1994.
  43. 43.T. Lindeberg. “Direct Estimation of Affine Deformations of Brightness Patterns Using Visual Front-End Operators with Automatic Scale Selection”. In Proc. 5th International Conference on Computer Vision, pages 134–141, Cambridge, MA, June 1995.
  44. 44.T. Lindeberg. “On scale selection in subsampled (hybrid) multi-scale representations”. 1995. draft manuscript.
  45. 45.T. Lindeberg. “Edge detection and ridge detection with automatic scale selection”. 1996. (Submitted).
  46. 46.T. Lindeberg. “Edge detection and ridge detection with automatic scale selection”. In Proc. IEEE Comp. Soc. Conf. on Computer Vision and Pattern Recognition, 1996, pages 465–470, San Francisco, California, June 1996. IEEE Computer Society Press.
  47. 47.T. Lindeberg. “Edge Detection and Ridge Detection with Automatic Scale Selection”. Technical Report ISRN KTH/NA/P--96/06--SE, Dept. of Numerical Analysis and Computing Science, KTH, Stockholm, Sweden, Jan. 1996. Also in Proc. IEEE Comp. Soc. Conf. on Computer Vision and Pattern Recognition, 1996. Extended version in Int. J. of Computer Vision, vol 30, number 2, 1998. (In press).
  48. 48.T. Lindeberg. “A Scale Selection Principle for Estimating Image Deformations”. Technical Report ISRN KTH/NA/P--96/16--SE, Dept. of Numerical Analysis and Computing Science, KTH, Stockholm, Sweden, Apr. 1996. to appear in Image and Vision Computing.
  49. 49.T. Lindeberg. “On Automatic Selection of Temporal Scales in Time-Casual Scale-Space”. In G. Sommer and J. J. Koenderink, editors, Proc. AFPAC’97: Algebraic Frames for the Perception-Action Cycle, volume 1315 of Lecture Notes in Computer Science, pages 94–113, Kiel, Germany, September 1997. Springer Verlag, Berlin.
  50. 50.S. G. Mallat and W. L. Hwang. “Singularity Detection and Processing with Wavelets”. IEEE Trans. Information Theory, 38(2):617–643, 1992.
  51. 51.S. G. Mallat and S. Zhong. “Characterization of Signals from Multi-Scale Edges”. IEEE Trans. Pattern Analysis and Machine Intell., 14(7):710–723, 1992.
  52. 52.D. Marr. Vision. W.H. Freeman, New York, 1982.
  53. 53.J. A. Noble. “Finding Corners”. Image and Vision Computing, 6(2):121–128, 1988.
  54. 54.E. J. Pauwels; P. Fiddelaers; T. Moons, and L. J. van Gool. “An extended class of scale-invariant and recursive scale-space filters”. IEEE Trans. Pattern Analysis and Machine Intell., 17(7):691–701, 1995.
  55. 55.H. Voorhees and T. Poggio. “Detecting Textons and Texture Boundaries in Natural Images”. In Proc. 1st Int. Conf. on Computer Vision, London, England, 1987.
  56. 56.K. Wiltschi; A. Pinz, and T. Lindeberg. “Classification of Carbide Distributions using Scale Selection and Directional Distributions”. In Proc. 4th International Conference on Image Processing, volume II, pages 122–125, Santa Barbara, California, October 1997. IEEE.
  57. 57.A. P. Witkin. “Scale-space filtering”. In Proc. 8th Int. Joint Conf. Art. Intell., pages 1019–1022, Karlsruhe, West Germany, Aug. 1983.
  58. 58.R. A. Young. “The Gaussian Derivative Theory of Spatial Vision: Analysis of Cortical Cell Receptive Field Line-Weighting Profiles”. Technical Report GMR-4920, Computer Science Department, General Motors Research Lab., Warren, Michigan, 1985.
  59. 59.R. A. Young. “The Gaussian derivative model for spatial vision: I. Retinal mechanisms”. Spatial Vision, 2:273–293, 1987.
  60. 60.A. L. Yuille and T. A. Poggio. “Scaling Theorems for Zero-Crossings”. IEEE Trans. Pattern Analysis and Machine Intell., 8:15–25, 1986.
  61. 61.W. Zhang and F. Bergholm. “An Extension of Marr’s Signature Based Edge Classification and Other Methods for Determination of Diffuseness and Height of Edges, as Well as Line Width”. In H.-H. Nagel et. al., editor, Proc. 4th Int. Conf. on Computer Vision, pages 183–191, Berlin, Germany, May. 1993. IEEE Computer Society Press.

Citation

MLA
Lindeberg, T. “Feature Detection with Automatic Scale Selection”. International Journal of Computer Vision, vol. 30, no. 2, 1998, pp. 79–116, https://doi.org/10.1023/A:1008045108935.
APA
Lindeberg, T. (1998). Feature Detection with Automatic Scale Selection. International Journal of Computer Vision, 30(2), 79–116. https://doi.org/10.1023/A:1008045108935
Chicago
Lindeberg, T. 1998. “Feature Detection with Automatic Scale Selection”. International Journal of Computer Vision 30 (2): 79–116. https://doi.org/10.1023/A:1008045108935.
Harvard
Lindeberg, T. (1998) “Feature Detection with Automatic Scale Selection”, International Journal of Computer Vision, 30(2), pp. 79–116. Available at: https://doi.org/10.1023/A:1008045108935.
Vancouver
1. Lindeberg T (1998) Feature Detection with Automatic Scale Selection. International Journal of Computer Vision 30:79–116

BibTeX

@article{Lindeberg_1998, title={Feature Detection with Automatic Scale Selection}, volume={30}, ISSN={1573-1405}, url={http://dx.doi.org/10.1023/A:1008045108935}, DOI={10.1023/a:1008045108935}, number={2}, journal={International Journal of Computer Vision}, publisher={Springer Science and Business Media LLC}, author={Lindeberg, Tony}, year={1998}, month=Nov, pages={79–116} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF