A Comparison of Affine Region Detectors

K. MikolajczykT. TuytelaarsCordelia SchmidAndrew ZissermanJiri MatasF. SchaffalitzkyT. KadirL. Gool

article2005IJCV3,522 citations

Establishes the standard benchmark dataset and quantitative evaluation protocol for assessing affine covariant region detectors across diverse imaging transformations, providing systematic performance comparisons of leading methods like MSER, Harris-Affine, and Hessian-Affine.

Listen

This paper evaluates and compares six affine covariant region detectors developed for robust local feature matching in computer vision applications such as wide-baseline stereo, object recognition, image retrieval, and robot localization. These detectors identify image regions whose shape adapts automatically to viewpoint and illumination changes so that the same physical surface patch can be described consistently across images. The work addresses the lack of standardized performance data by testing the detectors on a shared benchmark of planar scenes under controlled transformations.

The study set out to measure how well each detector finds repeatable and distinctive regions across changes in viewpoint, scale, blur, compression, and illumination, while also establishing public test images and evaluation code for future comparisons.

The authors implemented the detectors using code supplied by their creators and ran them on eight image sequences containing both structured scenes with distinct edges and textured scenes with repeated patterns. Repeatability was quantified by projecting detected elliptical regions through known homographies and measuring normalized overlap error. Distinctiveness was assessed by attaching 128-dimensional SIFT descriptors to the regions and counting correct nearest-neighbor matches.

MSER and Hessian-Affine detectors achieved the highest repeatability scores on most sequences, often exceeding 60 percent for moderate viewpoint or scale changes and remaining above 40 percent for larger deformations; Harris-Affine followed closely, while edge-based, intensity-extrema, and salient-region detectors performed well only on particular scene types. Absolute numbers of repeatable regions also varied widely, with Hessian-Affine and Harris-Affine typically returning several times more correspondences than the others. Matching scores using SIFT descriptors tracked the repeatability results for most transformations, indicating that the regions are sufficiently distinctive for automatic correspondence, although performance degraded noticeably on highly repetitive textures or under strong blur.

These results matter because practitioners can now select or combine detectors according to scene characteristics and application constraints rather than relying on anecdotal evidence. The public benchmark further enables objective progress tracking as new detectors appear.

The authors recommend running several complementary detectors in parallel to increase coverage and robustness, then filtering matches with geometric consistency checks. They also note that future detectors should be evaluated on the same data and protocol.

The evaluation is restricted to planar scenes or fixed-camera sequences related by homographies, uses default parameter settings, and measures performance only on the supplied medium-resolution images; results on non-planar 3D scenes or very different image statistics may differ. Overall confidence in the reported ordering of detectors on the tested conditions is high because the protocol is transparent, the data are public, and multiple performance metrics were examined.

  • Paper: A performance evaluation of local descriptors, Krystian Mikolajczyk et al. (2005). Reading this evaluation of local descriptors first provides the necessary background on how feature representations are paired with the affine region detectors evaluated in the source paper.
  • Paper: SURF: Speeded Up Robust Features, Herbert Bay et al. (2006). This paper builds directly upon the foundational affine and scale-invariant detectors benchmarked in the source by introducing faster approximate detector-descriptor designs.
  • Paper: BRIEF: Binary Robust Independent Elementary Features, Michael Calonder et al. (2010). This paper extends the local feature paradigm established by the source and its contemporaries into ultra-fast binary descriptors suitable for resource-constrained platforms.
Cover for A Comparison of Affine Region Detectors

Abstract

The paper gives a snapshot of the state of the art in affine covariant region detectors, and compares their performance on a set of test images under varying imaging conditions. Six types of detectors are included: detectors based on affine normalization around Harris (Mikolajczyk and Schmid, 2002; Schaffalitzky and Zisserman, 2002) and Hessian points (Mikolajczyk and Schmid, 2002), a detector ofmaximally stable extremal regions’, proposed by Matas et al. (2002); an edge-based region detector (Tuytelaars and Van Gool, 1999) and a detector based on intensity extrema (Tuytelaars and Van Gool, 2000), and a detector ofsalient regions’,

Table of Contents

  • A Comparison of Affine Region Detectors
  • 1. Introduction
  • 2. Affine Covariant Detectors
  • 2.1. Detectors Based on Affine Normalization—Harris-Affine & Hessian-Affine
  • 2.2. An Edge-Based Region Detector
  • 2.3. Intensity Extrema-Based Region Detector
  • 2.4. Maximally Stable Extremal Region Detector
  • 2.5. Salient Region Detector
  • 3. The Image Data Set
  • 3.1. Discussion
  • 4. Overlap Comparison Using Homographies
  • 4.1. Repeatability Measure
  • 4.2. Repeatability Under Various Transformations
  • 4.3. More Detailed Tests
  • 5. Matching Experiments
  • 5.1. Matching Score
  • 6. Conclusions

Knowls

  1. Knowl 1 — Overlap Error and Repeatability Evaluation Metric for Affine Region Detectors

    definition

    To quantitatively evaluate affine covariant region detectors independently of specific descriptor matching algorithms, detector accuracy is defined through the geometric overlap of detected regions under a ground truth planar homography.

    An elliptical affine region in an image is defined as: Rμ={xR2xTμx1}R_\mu = \{ \mathbf{x} \in \mathbb{R}^2 \mid \mathbf{x}^T \mu \mathbf{x} \le 1 \} where μ\mu is a 2×22 \times 2 symmetric positive-definite matrix representing the shape and size of the ellipse. Given a ground truth homography HH that maps coordinates from the reference image AA to target image BB, an ellipse μb\mu_b detected in image BB projects onto image AA as HTμbHH^T \mu_b H.

    The overlap error ϵO\epsilon_O between an ellipse RμaR_{\mu_a} detected in image AA and an ellipse RμbR_{\mu_b} detected in image BB (projected onto AA) is: ϵO=1Area(RμaRHTμbH)Area(RμaRHTμbH)\epsilon_O = 1 - \frac{\text{Area}(R_{\mu_a} \cap R_{H^T \mu_b H})}{\text{Area}(R_{\mu_a} \cup R_{H^T \mu_b H})}

    Two regions are deemed to correspond if their overlap error is below a predefined threshold, ϵO<ϵ0\epsilon_O < \epsilon_0 (typically ϵ0=0.40\epsilon_0 = 0.40, which corresponds to an approximate 20%20\% difference in region radii).

    To prevent detectors with intrinsically larger regions from receiving artificially higher overlap scores due to spatial scaling effects, both the reference ellipse RμaR_{\mu_a} and the mapped ellipse RHTμbHR_{H^T \mu_b H} are normalized by a uniform scaling factor ss such that the reference ellipse has an equivalent circular radius of r0=30r_0 = 30 pixels prior to numerical calculation of area union and intersection.

    The repeatability score for an image pair is defined as the ratio of valid correspondences to the maximum possible number of correspondences in the shared field of view: Repeatability=Ncorrmin(Na,Nb)\text{Repeatability} = \frac{N_{\text{corr}}}{\min(N_a, N_b)} where NcorrN_{\text{corr}} is the number of mutually corresponding region pairs, and Na,NbN_a, N_b are the counts of regions detected in the shared visible scene area of images AA and BB, respectively.

  2. Knowl 2 — SIFT-Based Matching Score Metric for Affine Covariant Regions

    experimental setup

    To evaluate the distinctiveness and matching capability of detected affine regions, local appearance within each region is normalized and encoded using the 128-dimensional Scale-Invariant Feature Transform (SIFT) descriptor.

    1. Affine Geometric and Photometric Normalization: Each detected elliptical region RμR_\mu is mapped to a canonical circular patch of 30×3030 \times 30 pixels using the matrix transformation μ1/2\mu^{1/2}. To capture surrounding contextual texture, the measurement region is scaled by a factor of 3 relative to the detected distinguished region size. The patch is then rotated to align with the dominant gradient orientation, achieving invariance to 2D affine deformations and rotations.

    2. Ground Truth Correspondence: A detected region pair (Rμa,Rμb)(R_{\mu_a}, R_{\mu_b}) is labeled as a ground truth match if its normalized overlap error under the known inter-image homography HH satisfies ϵO0.40\epsilon_O \le 0.40. Each region in image AA is restricted to at most one ground truth match.

    3. Descriptor Matching: Descriptors computed for regions in the target image are matched to their nearest neighbor in the reference image using Euclidean distance in the 128-dimensional SIFT space.

    4. Matching Score Definition: The matching score measures the proportion of detected regions that form geometrically correct nearest-neighbor matches: Matching Score=Ncorrectmin(Na,Nb)\text{Matching Score} = \frac{N_{\text{correct}}}{\min(N_a, N_b)} where NcorrectN_{\text{correct}} is the number of nearest-neighbor descriptor matches that satisfy ϵO0.40\epsilon_O \le 0.40, and Na,NbN_a, N_b are the number of detected regions lying in the common scene area.

  3. Knowl 3 — Harris-Affine and Hessian-Affine Detectors

    model/method

    The Harris-Affine and Hessian-Affine detectors identify multi-scale interest points and iteratively adapt their local elliptical shape to achieve affine covariance.

    1. Multi-Scale Point Detection:

      • Harris-Affine locates spatial extrema of the cornerness measure derived from the second moment matrix (auto-correlation matrix): M(x,σI,σD)=σD2g(σI)[Ix2(x,σD)IxIy(x,σD)IxIy(x,σD)Iy2(x,σD)]M(\mathbf{x}, \sigma_I, \sigma_D) = \sigma_D^2 \, g(\sigma_I) * \begin{bmatrix} I_x^2(\mathbf{x}, \sigma_D) & I_x I_y(\mathbf{x}, \sigma_D) \\ I_x I_y(\mathbf{x}, \sigma_D) & I_y^2(\mathbf{x}, \sigma_D) \end{bmatrix} where σD\sigma_D is the differentiation scale, σI\sigma_I is the integration scale, and g(σ)g(\sigma) is a Gaussian smoothing filter.
      • Hessian-Affine locates spatial extrema of the determinant of the Hessian matrix: H(x,σD)=[Ixx(x,σD)Ixy(x,σD)Ixy(x,σD)Iyy(x,σD)]H(\mathbf{x}, \sigma_D) = \begin{bmatrix} I_{xx}(\mathbf{x}, \sigma_D) & I_{xy}(\mathbf{x}, \sigma_D) \\ I_{xy}(\mathbf{x}, \sigma_D) & I_{yy}(\mathbf{x}, \sigma_D) \end{bmatrix} penalizing elongated ridge structures and selecting blob-like features.
    2. Characteristic Scale Selection: For each spatial point x\mathbf{x}, the characteristic scale σD\sigma_D is chosen where the scale-normalized Laplacian σD2Ixx(x,σD)+Iyy(x,σD)\sigma_D^2 |I_{xx}(\mathbf{x}, \sigma_D) + I_{yy}(\mathbf{x}, \sigma_D)| attains a local extremum over continuous scale space.

    3. Iterative Affine Shape Adaptation: Given an initial point and scale, the local neighborhood is warped by M1/2M^{1/2} so that the eigenvalues of the second moment matrix become equal. In the normalized coordinate frame, the spatial position, integration scale, and differentiation scale are updated, and MM is recomputed. This process repeats iteratively until the ratio of eigenvalues of MM converges to 1 within a tolerance threshold, recovering the affine transformation up to an unknown rotation.

  4. Knowl 4 — Maximally Stable Extremal Region Detector

    model/method

    The Maximally Stable Extremal Region (MSER) detector extracts connected components of binarized image levels whose geometry remains stable across intensity thresholds.

    1. Extremal Regions: For an image I:ΩRI: \Omega \to \mathbb{R}, an extremal region QΩQ \subset \Omega is a contiguous connected component such that for all boundary pixels qQq \in \partial Q and all interior pixels pQp \in Q, either I(p)>I(q)I(p) > I(q) (bright extremal region) or I(p)<I(q)I(p) < I(q) (dark extremal region).

    2. Stability Criterion: Let Q(t)Q(t) denote the extremal connected component obtained by thresholding at intensity tt. Stability is quantified by the relative area growth rate across a threshold step Δ\Delta: q(t)=Q(t+Δ)Q(tΔ)Q(t)q(t) = \frac{|Q(t + \Delta) \setminus Q(t - \Delta)|}{|Q(t)|} An extremal region Q(t)Q(t^*) is defined as a maximally stable extremal region if q(t)q(t) achieves a local minimum at tt^*.

    3. Invariance Properties: Because the set of extremal regions depends solely on the relative sorting order of pixel intensities, MSERs are strictly invariant under any continuous monotonic photometric mapping I=f(I)I' = f(I) (including non-linear gamma corrections). Furthermore, MSERs preserve topology under continuous geometric transformations and homographies.

    4. Elliptical Representation: To enable standardized descriptor computation and benchmarking, each detected arbitrarily shaped region QQ is represented by an equivalent ellipse having the identical first-order centroid and second-order central spatial moments.

    5. Complexity: MSER detection is implemented in O(n)O(n) time for 8-bit pixel sorting (using bin sort) plus O(nloglogn)O(n \log \log n) time for component tree tracking using a Union-Find disjoint-set data structure.

  5. Knowl 5 — Edge-Based Region Detector

    model/method

    The Edge-Based Region (EBR) detector extracts affine covariant regions by exploiting the geometric constraints of corners and adjacent edge contours.

    1. Feature Extraction: Multi-scale Harris corner points p\mathbf{p} and Canny edge curves are detected in the image.

    2. 1D Affine Invariant Edge Traversal: Two points p1(s1)\mathbf{p}_1(s_1) and p2(s2)\mathbf{p}_2(s_2) traverse the edge away from corner p\mathbf{p} in opposite directions along edge parameter sis_i. Their traversal speeds are synchronized by equalizing the relative affine-invariant area parameters l1=l2=ll_1 = l_2 = l: li=pi(1)(si)×(ppi(si))dsil_i = \int \left| \mathbf{p}_i^{(1)}(s_i) \times (\mathbf{p} - \mathbf{p}_i(s_i)) \right| ds_i where pi(1)(si)\mathbf{p}_i^{(1)}(s_i) is the first derivative with respect to arc parameter sis_i, and ×|\cdot \times \cdot| denotes the 2D cross-product determinant.

    3. Parallelogram Selection via Photometric Invariants: For each parameter value ll, vectors (p1(l)p)(\mathbf{p}_1(l) - \mathbf{p}) and (p2(l)p)(\mathbf{p}_2(l) - \mathbf{p}) define a parallelogram Ω(l)\Omega(l). Photometric moment invariants are evaluated over Ω(l)\Omega(l): Inv1=(p1pg)×(p2pg)(pp1)×(pp2)M001M002M000(M001)2\text{Inv}_1 = \left| \frac{(\mathbf{p}_1 - \mathbf{p}_g) \times (\mathbf{p}_2 - \mathbf{p}_g)}{(\mathbf{p} - \mathbf{p}_1) \times (\mathbf{p} - \mathbf{p}_2)} \right| \frac{M_{00}^1}{\sqrt{M_{00}^2 M_{00}^0 - (M_{00}^1)^2}} Inv2=(ppg)×(qpg)(pp1)×(pp2)M001M002M000(M001)2\text{Inv}_2 = \left| \frac{(\mathbf{p} - \mathbf{p}_g) \times (\mathbf{q} - \mathbf{p}_g)}{(\mathbf{p} - \mathbf{p}_1) \times (\mathbf{p} - \mathbf{p}_2)} \right| \frac{M_{00}^1}{\sqrt{M_{00}^2 M_{00}^0 - (M_{00}^1)^2}} where Mpqn=Ω(l)In(x,y)xpyqdxdyM_{pq}^n = \iint_{\Omega(l)} I^n(x, y) x^p y^q \, dx dy, pg=(M101M001,M011M001)\mathbf{p}_g = \left(\frac{M_{10}^1}{M_{00}^1}, \frac{M_{01}^1}{M_{00}^1}\right) is the intensity-weighted center of gravity, and q\mathbf{q} is the corner vertex opposite to p\mathbf{p}. Parallelograms where these invariants attain local extrema are selected and fitted with their moment-equivalent ellipses.

  6. Knowl 6 — Intensity Extrema-Based Region Detector

    model/method

    The Intensity Extrema-Based Region (IBR) detector extracts affine covariant regions by analyzing 1D intensity profiles along rays emanating radially from local intensity extrema.

    1. Extrema Detection: Local intensity extrema x0\mathbf{x}_0 with intensity I0=I(x0)I_0 = I(\mathbf{x}_0) are detected across multiple scales.

    2. Radial Profile Invariant Function: Rays are projected outward from x0\mathbf{x}_0 in all directions. Along each ray parameterized by arc parameter tt, the following invariant function is computed: fI(t)=I(t)I0max(1t0tI(t)I0dt,d)f_I(t) = \frac{|I(t) - I_0|}{\max\left(\frac{1}{t} \int_0^t |I(t') - I_0| \, dt', \, d\right)} where d>0d > 0 is a small regularization constant preventing division by zero. The parameter tt where fI(t)f_I(t) reaches a local maximum identifies a sharp, affine-invariant boundary transition along that ray direction.

    3. Region Delineation and Ellipse Fitting: The boundary points corresponding to the extrema of fI(t)f_I(t) across all rays are linked to delineate an irregularly shaped closed region. The final region is represented by an ellipse having the identical first- and second-order moments as the delineated region, which preserves affine covariance.

  7. Knowl 7 — Salient Region Detector

    model/method

    The Salient Region detector identifies affine covariant regions by measuring the entropy of the local intensity probability distribution across scale and affine shape parameters.

    1. Local Intensity Entropy: For each pixel x\mathbf{x}, candidate elliptical windows E\mathcal{E} parameterized by scale ss (major axis length), orientation θ\theta, and axis ratio λ\lambda are constructed. The Shannon entropy of the intensity probability density function p(I;s,θ,λ)p(I; s, \theta, \lambda) within E\mathcal{E} is: H(s,θ,λ)=Ip(I;s,θ,λ)logp(I;s,θ,λ)\mathcal{H}(s, \theta, \lambda) = -\sum_I p(I; s, \theta, \lambda) \log p(I; s, \theta, \lambda)

    2. Scale Extrema Search: For every pixel and shape configuration (θ,λ)(\theta, \lambda), scale values ss^* where entropy H\mathcal{H} attains a local extremum over scale ss are selected as candidate salient regions.

    3. Saliency Weighting and Ranking: Each candidate region is assigned a saliency score Y=HW\mathcal{Y} = \mathcal{H} \cdot \mathcal{W}, where W\mathcal{W} weights the entropy by the magnitude of the scale derivative of the probability distribution: W=s22s1Ip(I;s,θ,λ)s\mathcal{W} = \frac{s^2}{2s - 1} \sum_I \left| \frac{\partial p(I; s, \theta, \lambda)}{\partial s} \right| All candidate regions across the image are ranked by their saliency Y\mathcal{Y}, and the top PP regions are selected.

  8. Knowl 8 — Computational Complexity and Runtime of Affine Region Detectors

    data/table

    Execution runtimes and region counts were measured on an 800×640800 \times 640 image on a 2.0 GHz Pentium 4 Linux PC using default author parameters:

    Detector Run time (min:sec) Number of regions
    Harris-Affine 0:01.43 1791
    Hessian-Affine 0:02.73 1649
    MSER 0:00.66 533
    IBR 0:10.82 679
    EBR 2:44.59 1265
    Salient Regions 33:33.89 513

    Theoretical computational complexities across detectors are:

    • MSER: O(n)O(n) for initial pixel sorting via bin sort plus O(nloglogn)O(n \log \log n) for connected component tracking with union-find, resulting in sub-second execution.
    • Harris-Affine and Hessian-Affine: O(n)O(n) for initial interest point extraction, followed by O((m+k)p)O((m + k)p) for automatic scale selection across mm scales and kk shape adaptation iterations over pp points.
    • IBR: O(n)O(n) for extrema detection plus O(p)O(p) for ray-based boundary construction over pp extrema.
    • EBR: O(n)O(n) for corner and Canny edge extraction plus O(pd)O(pd) for edge tracing across pp corners with an average of dd nearby edges.
    • Salient Regions: O(nl)O(n \cdot l) for testing ll discretized ellipse geometries per pixel, plus O(e)O(e) for ranking ee candidate extrema.
  9. Knowl 9 — Comparative Repeatability of Affine Detectors Under Image Transformations

    empirical result

    Empirical evaluation across test sequences spanning structured (planar surfaces with distinct contours) and textured scenes under five transformation types demonstrates:

    • Viewpoint Variation: Viewpoint tilt (from 2020^\circ to 6060^\circ) represents the most severe degradation, with repeatability declining from 40%–78% at 2020^\circ to 10%–46% at 6060^\circ. MSER achieves the highest percentage repeatability across both structured and textured scenes due to high boundary localization accuracy. Hessian-Affine and Harris-Affine produce the highest absolute number of correct correspondences (1200–1300 at 2020^\circ, dropping to 200–400 at 6060^\circ).
    • Scale Change and Zoom: For scale changes up to 4×4\times combined with in-plane rotation, Hessian-Affine achieves the highest repeatability, followed by MSER and Harris-Affine, confirming the robustness of Laplacian-based characteristic scale selection. EBR fails on textured scenes (<20%<20\% repeatability) due to a lack of stable long edges.
    • Image Blur: All detectors demonstrate high blur invariance with nearly horizontal repeatability curves, except MSER. MSER repeatability drops significantly under severe blur because blurred edges degrade the stability of intensity-thresholded component boundaries.
    • JPEG Compression and Illumination: Hessian-Affine and Harris-Affine achieve the highest repeatability under JPEG compression (up to 95%). All detectors maintain stable, near-constant repeatability across illumination changes induced by camera aperture variation, with MSER achieving the highest percentage score.
  10. Knowl 10 — Trade-offs in Region Distinctiveness and Matching Precision

    empirical result

    Matching experiments using 128-dimensional SIFT descriptors reveal systematic differences between affine detector types:

    1. High Precision at Low Match Counts (MSER and IBR): MSER and IBR detect sparse, distinct, non-overlapping regions. When matching candidate regions by nearest-neighbor distance in SIFT feature space, over 90%90\% of the top 100 returned matches are geometrically correct (overlap error ϵO0.40\epsilon_O \le 0.40). These detectors are optimal for applications requiring a minimal set of highly accurate correspondences (e.g., estimating two-view epipolar geometry without extensive filtering).

    2. High Region Density with Redundancy (Hessian-Affine and Harris-Affine): Hessian-Affine and Harris-Affine extract dense region sets (~1600–1800 detections per image), frequently producing multiple slightly shifted or differently scaled elliptical detections centered on the same local image structure. Because competing detections have similar descriptor representations, nearest-neighbor descriptor matching generates higher false positive rates at tight descriptor distance thresholds. However, they provide significantly more total correct matches (over 200 matches) when geometric filtering or RANSAC verification is applied, making them superior in cluttered or heavily occluded scenes.

Coverage note — None was omitted; all six affine region detector formulations, the benchmarking protocol, experimental parameters, complexity analyses, and empirical findings are fully covered.

References

  1. 1.Baumberg, A. 2000. Reliable feature matching across widely separated views. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Hilton Head Island, South Carolina, USA, pp. 774–781.
  2. 2.Brown, M. and Lowe, D. 2003. Recognizing panoramas. In Proceedings of the International Conference on Computer Vision, Nice, France, pp. 1218–1225.
  3. 3.Canny, J. 1986. A computational approach to edge detection. IEEE Transactions on Pattern Analysis and Machine Intelligence, 8: 679–698.
  4. 4.Csurka, G., Dance, C., Bray, C., and Fan, L. 2004. Visual categorization with bags of keypoints. In Proceedings Workshop on Statistical Learning in Computer Vision.
  5. 5.Dorko, G. and Schmid, C. 2003. Selection of scale invariant neighborhoods for object class recognition. In Proceedings International Conference on Computer Vision, Nice, France, pp. 634–640.
  6. 6.Fergus, R., Perona, P., and Zisserman, A. 2003. Object class recognition by unsupervised scale-invariant learning. In Proceedings IEEE Conference on Computer Vision and Pattern Recognition, Madison, Wisconsin, USA.
  7. 7.Ferrari, V., Tuytelaars, T., and Van Gool, L. 2001. Simultaneous object recognition and segmentation by image exploration. In Proceedings European Conference on Computer Vision, Prague, Czech Republic, pp. 40–54.
  8. 8.Ferrari, V., Tuytelaars, T., and Van Gool, L. 2005. Simultaneous object recognition and segmentation from single or multiple model views. International Journal of Computer Vision, to appear.
  9. 9.Goedeme, T., Tuytelaars, T., and Van Gool, L. 2004. Fast wide baseline matching for visual navigation. In Proceedings IEEE Conference on Computer Vision and Pattern Recognition, Washington, DC, USA, pp. 24–29.
  10. 10.Harris, C. and Stephens, M. 1988. A combined corner and edge detector. In Alvey Vision Conference, pp. 147–151.
  11. 11.Hartley, R.I. and Zisserman, A. 2004. Multiple View Geometry in Computer Vision, 2nd edition, Cambridge University Press, ISBN: 0521540518.
  12. 12.Kadir, T., Zisserman, A., and Brady, M. 2004. An affine invariant salient region detector. In Proceedings of the 8th European Conference on Computer Vision, Prague, Czech Republic, pp. 345–457.
  13. 13.Lazebnik, S., Schmid, C., and Ponce, J. 2003a. A sparse texture representation using affine-invariant regions. in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Madison, Wisconsin, USA, pp. 319–324.
  14. 14.Lazebnik, S., Schmid, C., and Ponce, J. 2003b. Affine-invariant local descriptors and neighborhood statistics for texture recognition. In Proceedings of the International Conference on Computer Vision, Nice, France, pp. 649–655.
  15. 15.Lazebnik, S., Schmid, C., and Ponce, J. 2005. A sparse texture representation using local affine regions. IEEE Transactions on Pattern Analysis and Machine Intelligence 27(8):1265–1278.
  16. 16.Lindeberg, T. and Garding, J. 1997. Shape-adapted smoothing in estimation of 3-D shape cues from affine deformations of local 2-D brightness structure. Image and Vision Computing 15(6):415–434.
  17. 17.Lindeberg, T. 1998. Feature detection with automatic scale selection. International Journal of Computer Vision 30(2):79–116.
  18. 18.Lowe, D. 1999. Object recognition from local scale-invariant features. In Proceedings of the 7th International Conference on Computer Vision, Kerkyra, Greece, pp. 1150–1157.
  19. 19.Lowe, D. 2004. Distinctive image features from scale-invariant keypoints. International Journal on Computer Vision 60(2):91–110.
  20. 20.Matas, J., Burianek, J., and Kittler, J. 2000. Object Recognition using the Invariant Pixel-Set Signature. In Proceedings of the British Machine Vision Conference, London, UK, pp. 606–615.
  21. 21.Matas, J. Chum, O., Urban, M., and Pajdla, T. 2002. Robust wide-baseline stereo from maximally stable extremal regions. In Proceedings of the British Machine Vision Conference, Cardiff, UK, pp. 384–393.
  22. 22.Matas, J., Chum, O., Urban, M., and Pajdla, T. 2004. Robust wide-baseline stereo from maximally stable extremal regions. Image and Vision Computing 22(10):761–767.
  23. 23.Mikolajczyk, K. and Schmid, C. 2001. Indexing based on scale invariant interest points. In Proceedings of the 8th International Conference on Computer Vision, Vancouver, Canada.
  24. 24.Mikolajczyk, K. and Schmid, C. 2002. An affine invariant interest point detector. In Proceedings of the 7th European Conference on Computer Vision, Copenhagen, Denmark.
  25. 25.Mikolajczyk, K. and Schmid, C. 2003. A performance evaluation of local descriptors. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition, Madison, Wisconsin, USA.
  26. 26.Mikolajczyk, K., Zisserman, A., and Schmid, C. 2003. Shape recognition with edge-based features. In Proceedings of the British Machine Vision Conference, Norwich, UK.
  27. 27.Mikolajczyk, K. and Schmid, C. 2004. Scale & affine invariant interest point detectors. International Journal on Computer Vision 60(1):63–86.
  28. 28.Mikolajczyk, K. and Schmid, C. 2005. A performance evaluation of local descriptors. IEEE Transactions on Pattern Analysis and Machine Intelligence 27(10):1615–1630.
  29. 29.Obdržalek, S. and Matas, J. 2002. Object recognition using local affine frames on distinguished regions. In Proceedings of the British Machine Vision Conference, Cardiff, UK, pp. 113–122.
  30. 30.Opelt, A., Fussenegger, M., Pinz, A., and Auer, P. 2004. Weak hypotheses and boosting for generic object detection and recognition. In Proceedings of European Conference on Computer Vision, Prague, Czech Republic, pp. 71–84.
  31. 31.Pritchett, P. and Zisserman, A. 1998. Wide baseline stereo matching. In Proceedings of the 6th International Conference on Computer Vision, Bombay, India, pp. 754–760.
  32. 32.Rothganger, F., Lazebnik, S., Schmid, C., and Ponce, J. 2003. 3D object modeling and recognition using affine-invariant patches and multi-view spatial constraints. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition, Madison, Wisconsin, USA, pp. 272–277.
  33. 33.Rothganger, F., Lazebnik, S., Schmid, C., and Ponce, J. 2005. Object modeling and recognition using local affine-invariant image descriptors and multi-view spatial consraints. International Journal of Computer Vision, to appear.
  34. 34.Schaffalitzky, F., and Zisserman, A. 2002. Multi-view matching for unordered image sets, or ‘‘How do I organize my holiday snaps?’’. In Proceedings of the 7th European Conference on Computer Vision, Copenhagen, Denmark, pp. 414–431.
  35. 35.Schaffalitzky, F. and Zisserman, A. 2003. Automated Location matching in movies. Computer Vision and Image Understanding, 92(2):236–264.
  36. 36.Schmid, C. and Mohr, R. 1997. Local grayvalue invariants for image retrieval. IEEE Transactions on Pattern Analysis and Machine Intelligence 19(5):530–535.
  37. 37.Se, S., Lowe, D., and Little, J. 2002. Mobile robot localization and mapping with uncertainty using scale-invariant visual landmarks. International Journal of Robotics Research 21(8):735–758.
  38. 38.Sedgewick, R. 1988. Algorithms, 2nd edition. Addison-Wesley.
  39. 39.Sivic, J., and Zisserman, A. 2003. Video google: A text retrieval approach to object matching in videos. In Proceedings of the International Conference on Computer Vision, Nice, France.
  40. 40.Sivic, J., Schaffalitzky, F., and Zisserman, A. 2004. Object level grouping for video shots. In Proceedings of the 8th European Conference on Computer Vision, Prague, Czech Republic, pp. 724–734.
  41. 41.Sivic, J., and Zisserman, A. 2004. Video data mining using configurations of viewpoint invariant regions. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Washington, DC, USA, pp. 488–495.
  42. 42.Tell, D. and Carlsson, S. 2000. Wide baseline point matching using affine invariants computed from intensity profiles. In Proceedings of the 6th European Conference on Computer Vision, Dublin, Ireland, pp. 814–828.
  43. 43.Tell, D. and Carlsson, S. 2002. Combining appearance and topology for wide baseline matching. In Proceedings of the 7th European Conference on Computer Vision, Copenhagen, Denmark, pp. 68–81.
  44. 44.Turina, A., Tuytelaars, T., and Van Gool, L. 2001. Efficient Grouping under perspective skew. In Proceedings IEEE Conference on Computer Vision and Pattern Recognition, Hawaii, USA, pp. 247–254.
  45. 45.Tuytelaars, T. and Van Gool, L. 1999. Content-based image retrieval based on local affinely invariant regions. In Int. Conf. on Visual Information Systems, pp. 493–500.
  46. 46.Tuytelaars, T., Van Gool, L., D’haene, L., and Koch, R. 1999. Matching of affinely invariant regions for visual servoing. In Int. Conference Robotics and Automation ICRA 99.
  47. 47.Tuytelaars, T. and Van Gool, L. 2000. Wide baseline stereo matching based on local, affinely invariant regions. In Proceedings of the 11th British Machine Vision Conference, Bristol, UK, pp. 412–425.
  48. 48.Tuytelaars, T. and Van Gool, L. 2004. Matching Widely Separated Views based on Affine Invariant Regions. International Journal on Computer Vision 59(1):61–85.

Citation

MLA
Mikolajczyk, K., et al. “A Comparison of Affine Region Detectors”. International Journal of Computer Vision, vol. 65, nos. 1-2, 2005, pp. 43–72, https://doi.org/10.1007/s11263-005-3848-x.
APA
Mikolajczyk, K., Tuytelaars, T., Schmid, C., Zisserman, A., Matas, J., Schaffalitzky, F., Kadir, T., & Gool, L. V. (2005). A Comparison of Affine Region Detectors. International Journal of Computer Vision, 65(1-2), 43–72. https://doi.org/10.1007/s11263-005-3848-x
Chicago
Mikolajczyk, K., T. Tuytelaars, C. Schmid, et al. 2005. “A Comparison of Affine Region Detectors”. International Journal of Computer Vision 65 (1-2): 43–72. https://doi.org/10.1007/s11263-005-3848-x.
Harvard
Mikolajczyk, K. et al. (2005) “A Comparison of Affine Region Detectors”, International Journal of Computer Vision, 65(1-2), pp. 43–72. Available at: https://doi.org/10.1007/s11263-005-3848-x.
Vancouver
1. Mikolajczyk K, Tuytelaars T, Schmid C, Zisserman A, Matas J, Schaffalitzky F, Kadir T, Gool LV (2005) A Comparison of Affine Region Detectors. International Journal of Computer Vision 65:43–72

BibTeX

@article{Mikolajczyk_2005, title={A Comparison of Affine Region Detectors}, volume={65}, ISSN={1573-1405}, url={http://dx.doi.org/10.1007/s11263-005-3848-x}, DOI={10.1007/s11263-005-3848-x}, number={1-2}, journal={International Journal of Computer Vision}, publisher={Springer Science and Business Media LLC}, author={Mikolajczyk, K. and Tuytelaars, T. and Schmid, C. and Zisserman, A. and Matas, J. and Schaffalitzky, F. and Kadir, T. and Gool, L. Van}, year={2005}, month=Oct, pages={43–72} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF