Fast Feature Pyramids for Object Detection

Piotr DollárRon AppelSerge J. BelongieP. Perona

article2014TPAMI2,020 citations

Demonstrates how to construct fast multi-scale feature pyramids by extrapolating intermediate levels from octave-spaced scales using natural image statistics, enabling real-time object and pedestrian detection with minimal accuracy loss.

Listen

Visual object detection systems, such as those used in autonomous driving, mobile devices, and robotics, require high accuracy while operating under strict real-time and low-power computational constraints. Over recent years, detection algorithms achieved dramatic improvements in accuracy by extracting rich visual representations across finely sampled scale pyramids, but this drastically increased processing time from real-time rates to multiple seconds per image. The article sets out to demonstrate that densely sampled visual feature pyramids can be rapidly estimated via mathematical extrapolation from coarsely sampled scales, eliminating the conventional trade-off between computational speed and detection accuracy.

The authors conducted a comprehensive statistical and empirical analysis grounded in the fractal properties of natural images. They evaluated a broad class of low-level visual features across thousands of natural and pedestrian images to model feature behavior across scale changes. They integrated this scaling model into three distinct detection systems: Aggregated Channel Features, Integral Channel Features, and Deformable Part Models. These systems were evaluated on standard benchmark datasets, including INRIA, Caltech, TUD-Brussels, and ETH for pedestrian detection, and PASCAL VOC for general multi-class object detection.

The article established several critical findings. First, low-level visual features scale predictably across image resolutions according to a consistent power law, allowing feature channels computed at sparse intervals of one octave to accurately predict intermediate scales. Second, feature computation costs are reduced by roughly an order of magnitude, requiring only 33% more computation than single-scale extraction rather than the dense multi-scale baseline. Third, the new Aggregated Channel Features detector achieved real-time speeds exceeding 30 frames per second on standard resolution imagery on a single computer processor, compared to 12 frames per second for exact multi-scale extraction. Fourth, this massive speedup caused negligible loss in accuracy, yielding an average miss rate across pedestrian datasets of 41% with fast feature pyramids compared to 40% with exact pyramids, while Deformable Part Models on general object detection saw only a 2% drop in mean average precision.

These findings show that computing redundant, high-resolution features across every scale is unnecessary. For technical leaders and engineering teams, this method enables real-time, high-accuracy computer vision on cost-effective, standard hardware without requiring expensive graphics processing units. Practitioners should adopt fast feature pyramid architectures in sliding-window visual pipelines, calibrate scaling exponents on representative domain data, and combine these pyramids with optimized classification cascades for maximum frame rate gains. Decision-makers should note that the underlying power-law assumptions apply specifically to natural scenes with broad visual spectra; caution is warranted when applying the method to synthetic environments, extreme magnifications, or narrow-band periodic textures where the approximation breaks down.

Cover for Fast Feature Pyramids for Object Detection

Abstract

Multi-resolution image features may be approximated via extrapolation from nearby scales, rather than being computed explicitly. This fundamental insight allows us to design object detection algorithms that are as accurate, and considerably faster, than the state-of-the-art. The computational bottleneck of many modern detectors is the computation of features at every scale of a finely-sampled image pyramid. Our key insight is that one may compute finely sampled feature pyramids at a fraction of the cost, without sacrificing performance: for a broad family of features we find that features computed at octave-spaced scale intervals are sufficient to approximate features on a finely-sampled pyramid. Extrapolation is inexpensive as compared to direct feature computation. As a result, our approximation yields considerable speedups with negligible loss in detection accuracy. We modify three diverse visual recognition systems to use fast feature pyramids and show results on both pedestrian detection (measured on the Caltech, INRIA, TUD-Brussels and ETH datasets) and general object detection (measured on the PASCAL VOC). The approach is general and is widely applicable to vision algorithms requiring fine-grained multi-scale analysis. Our approximation is valid for images with broad spectra (most natural images) and fails for images with narrow band-pass spectra (e.g. periodic textures).

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Multiscale Gradient Histograms
  • 3.1 Gradient Histograms in Upsampled Images
  • 3.2 Gradient Histograms in Downsampled Images
  • 3.3 Histograms of Normalized Gradients
  • 4 Statistics of Multiscale Features
  • 4.1 Power Law Governs Feature Scaling
  • 4.2 Estimating λ
  • 4.3 Deviation for Individual Images
  • 4.4 Miscellanea
  • 5 Fast Feature Pyramids
  • 5.1 Feature Channel Scaling
  • 5.2 Fast Feature Pyramids
  • 5.3 Complexity Analysis
  • 6 Applications to Object Detection
  • 6.1 Aggregated Channel Features (ACF)
  • 6.2 Integral Channel Features (ICF)
  • 6.3 Deformable Part Models (DPM)
  • 7 Conclusion
  • Acknowledgments
  • References

Knowls

  1. Knowl 1 — Power Law Governing Feature Channel Scaling

    theoretical result

    Let Ω\Omega be any shift-invariant feature mapping that transforms an image II into a multi-channel feature map C=Ω(I)C = \Omega(I) with spatial dimensions h×wh \times w and kk layers. Let IsI_s denote the image at scale ss (with dimensions sh×sws h \times s w), and let Cs=Ω(Is)C_s = \Omega(I_s) denote the channel computed at scale ss. The global mean feature response across all spatial locations and layers is defined as:

    fΩ(Is)=1s2hwk∑i=1sh∑j=1sw∑l=1kCs(i,j,l)f_\Omega(I_s) = \frac{1}{s^2 h w k} \sum_{i=1}^{s h} \sum_{j=1}^{s w} \sum_{l=1}^{k} C_s(i, j, l)

    Because natural images exhibit scale-invariant, fractal power spectra, the ratio of mean feature responses computed at two arbitrary scales s1s_1 and s2s_2 follows a power law:

    fΩ(Is1)fΩ(Is2)=(s1s2)−λΩ+E\frac{f_\Omega(I_{s_1})}{f_\Omega(I_{s_2})} = \left(\frac{s_1}{s_2}\right)^{-\lambda_\Omega} + \mathcal{E}

    where λΩ\lambda_\Omega is a channel-specific scaling exponent and E\mathcal{E} is a zero-mean random deviation (E[E]≈0\mathbb{E}[\mathcal{E}] \approx 0) whose variance increases gradually with the scale ratio ∣log⁡2(s1/s2)∣|\log_2(s_1/s_2)|. The same scaling relationship holds locally for sub-windows and local histograms.

  2. Knowl 2 — Feature Channel Approximation via Power-Law Rescaling

    equation

    Given a precomputed feature channel map Cs′=Ω(Is′)C_{s'} = \Omega(I_{s'}) at scale s′s', the feature channel map Cs=Ω(Is)C_s = \Omega(I_s) at a nearby scale ss is approximated directly without evaluating Ω(Is)\Omega(I_s) by:

    Cs≈R(Cs′,ss′)⋅(ss′)−λΩC_s \approx R\left(C_{s'}, \frac{s}{s'}\right) \cdot \left(\frac{s}{s'}\right)^{-\lambda_\Omega}

    where R(C,α)R(C, \alpha) denotes the spatial resampling (e.g., bilinear interpolation) of channel map CC by a scale factor of α\alpha, and λΩ\lambda_\Omega is the characteristic power-law scaling exponent for channel type Ω\Omega. The approximation error is minimized when CC is spatially aggregated (downsampled or smoothed) over local pixel blocks.

  3. Knowl 3 — Fast Feature Pyramid Construction

    algorithm

    The fast feature pyramid algorithm computes exact feature maps only at sparsely sampled octave intervals and approximates all intermediate scale levels via spatial resampling and power-law scaling.

    Input: Input image II, scale set S={s1,s2,…,sM}S = \{s_1, s_2, \dots, s_M\} sampled logarithmically with mm scales per octave, feature extraction operator Ω\Omega, scaling exponent λΩ\lambda_\Omega
    Output: Multiscale feature pyramid {Cs}s∈S\{C_s\}_{s \in S}
    OctaveScales ←{s∈S∣log⁡2(s)∈Z}\leftarrow \{s \in S \mid \log_2(s) \in \mathbb{Z}\}
    for each s′∈OctaveScaless' \in \text{OctaveScales} do
        Is′←R(I,s′)I_{s'} \leftarrow R(I, s')
        Cs′←Ω(Is′)C_{s'} \leftarrow \Omega(I_{s'})
    end for
    for each s∈Ss \in S do
        if s∈OctaveScaless \in \text{OctaveScales} then
            continue
        end if
        s∗←arg⁡min⁡u∈OctaveScales∣log⁡2(s)−log⁡2(u)∣s^* \leftarrow \arg\min_{u \in \text{OctaveScales}} |\log_2(s) - \log_2(u)|
        α←s/s∗\alpha \leftarrow s / s^*
        Cs←R(Cs∗,α)⋅α−λΩC_s \leftarrow R(C_{s^*}, \alpha) \cdot \alpha^{-\lambda_\Omega}
    end for
    return {Cs}s∈S\{C_s\}_{s \in S}
  4. Knowl 4 — Computational Complexity of Exact vs. Fast Feature Pyramids

    theoretical result

    Assuming the feature extractor Ω\Omega has a cost linear in the number of pixels of an n×nn \times n image, constructing an exact feature pyramid with mm scales per octave over an infinite geometric progression of scales has a computational cost of:

    ∑k=0∞n22−2k/m=n21−4−1/m≈mn2ln⁡4\sum_{k=0}^{\infty} n^2 2^{-2k/m} = \frac{n^2}{1 - 4^{-1/m}} \approx \frac{m n^2}{\ln 4}

    For typical sampling densities of m=8m = 8 to 1212 scales per octave, exact computation requires approximately 6n26 n^2 to 9n29 n^2 operations.

    In contrast, the fast feature pyramid evaluates Ω\Omega only once per octave (m=1m = 1), which requires:

    ∑k=0∞n24−k=43n2\sum_{k=0}^{\infty} n^2 4^{-k} = \frac{4}{3} n^2

    This represents only a 33%33\% overhead over evaluating features at the single base scale n2n^2. Approximating the intermediate m−1m-1 scales via 2D spatial resampling RR introduces minimal overhead, particularly when channels are spatially aggregated.

  5. Knowl 5 — Aggregated Channel Features (ACF) Detector

    model/method

    The Aggregated Channel Features (ACF) object detector operates on 10 channel maps computed from an input image II:

    1. 1 normalized gradient magnitude channel: M~(i,j)=M(i,j)/(Mˉ(i,j)+0.005)\widetilde{M}(i, j) = M(i, j) / (\bar{M}(i, j) + 0.005), where MM is gradient magnitude and Mˉ\bar{M} is obtained by convolving MM with an L1L_1-normalized 11×1111 \times 11 triangle filter.
    2. 6 orientation-quantized gradient histogram channels.
    3. 3 LUV color channels.

    Prior to channel computation, II is smoothed with a [1 2 1]/4[1\ 2\ 1]/4 filter. The 10 channels are divided into non-overlapping 4×44 \times 4 spatial blocks, summed within each block, and post-smoothed with a [1 2 1]/4[1\ 2\ 1]/4 filter.

    Features are single-pixel lookups into the aggregated channels. For a 128×64128 \times 64 detection window, this yields (128/4)×(64/4)×10=5120(128/4) \times (64/4) \times 10 = 5120 candidate features. A boosted classifier combining 2048 depth-two decision trees is trained using AdaBoost with multiple bootstrapping rounds. During sliding-window evaluation, the detector steps by 4 pixels with 8 scales per octave, running at over 30 frames per second on 640×480640 \times 480 images on a single CPU.

  6. Knowl 6 — Empirical Scaling Exponents Across Standard Visual Features

    empirical result

    Least-squares estimation of log⁡2(μs)=aΩ′−λΩlog⁡2(s)\log_2(\mu_s) = a'_\Omega - \lambda_\Omega \log_2(s) over 24 downsampled scales across 3 octaves yields distinct scaling exponents λΩ\lambda_\Omega and mean absolute fitting errors ∣E[E]∣|\mathbb{E}[\mathcal{E}]| on natural and pedestrian image ensembles:

    • Grayscale pixel intensities: λ=0.000\lambda = 0.000, ∣E[E]∣=0.000|\mathbb{E}[\mathcal{E}]| = 0.000
    • HOG (4×44 \times 4 spatial bins, 36 channels): λ=0.078\lambda = 0.078, ∣E[E]∣=0.004|\mathbb{E}[\mathcal{E}]| = 0.004 (natural), 0.0190.019 (pedestrian)
    • Normalized gradient histograms (6 bins): λ=0.101\lambda = 0.101, ∣E[E]∣=0.006|\mathbb{E}[\mathcal{E}]| = 0.006 (natural), 0.0260.026 (pedestrian)
    • Local standard deviation (5×55 \times 5 window): λ=0.358\lambda = 0.358, ∣E[E]∣=0.025|\mathbb{E}[\mathcal{E}]| = 0.025 (natural), 0.0440.044 (pedestrian)
    • Unnormalized gradient histograms (6 bins): λ=0.406\lambda = 0.406, ∣E[E]∣=0.018|\mathbb{E}[\mathcal{E}]| = 0.018 (natural), 0.0370.037 (pedestrian)
    • Difference of Gaussians (DoG, σin=0.71,σout=1.14\sigma_{in} = 0.71, \sigma_{out} = 1.14): λ=0.974\lambda = 0.974, ∣E[E]∣=0.013|\mathbb{E}[\mathcal{E}]| = 0.013 (natural), 0.0190.019 (pedestrian)

    For all tested channels, the standard deviation of single-image deviation at scale s=1/2s = 1/2 satisfies σ1/2<0.20\sigma_{1/2} < 0.20.

  7. Knowl 7 — Pedestrian Detection Benchmark Comparison of Exact vs. Fast Pyramids

    data/table

    Log-average miss rates (MR, in percent, lower is better) evaluated between 10−210^{-2} and 10010^0 false positives per image across four standard pedestrian datasets:

    Method INRIA Caltech TUD-Brussels ETH Mean
    Shapelet 82 91 95 91 90
    Viola-Jones (VJ) 72 95 95 90 88
    PoseInv 80 86 88 92 87
    HikSvm 43 73 83 72 68
    HOG 46 68 78 64 64
    HogLbp 39 68 82 55 61
    MF 36 68 73 60 59
    PLS 40 62 71 55 57
    MF+CSS 25 61 60 61 52
    LatSvmV2 20 63 70 51 51
    FPDW 21 57 63 60 50
    ChnFtrs 22 56 60 57 49
    Crosstalk 19 54 58 52 46
    ICF-Exact 18 48 53 50 42
    ICF (approximate) 19 51 55 56 45
    ACF-Exact 17 43 50 50 40
    ACF (approximate) 17 45 52 51 41

    Approximating 7 of every 8 pyramid scales incurs only a 1%1\% increase in mean MR for ACF (40%40\% to 41%41\%) and a 3%3\% increase for ICF (42%42\% to 45%45\%), while accelerating ACF to over 30 fps on 640×480640 \times 480 images.

  8. Knowl 8 — Performance of Deformable Part Models Using Fast Feature Pyramids on PASCAL VOC 2007

    data/table

    Average Precision (AP, in percent, higher is better) for Deformable Part Models on the 20 object classes of PASCAL VOC 2007 comparing exact HOG pyramids (DPM) to pyramids where 9 out of 10 scales per octave are approximated via the power law (∼\simDPM):

    Method plane bike bird boat bottle bus car cat chair cow
    DPM 26.3 59.4 2.3 10.2 21.2 46.2 52.2 7.9 15.9 17.4
    ∼\simDPM 24.1 54.7 1.6 9.8 20.0 42.1 50.1 8.0 13.8 16.7
    Method table dog horse moto person plant sheep sofa train tv
    DPM 10.9 2.9 53.4 37.6 38.2 4.9 16.6 29.7 38.2 40.8
    ∼\simDPM 8.9 2.5 49.4 38.3 36.0 4.2 14.9 24.4 35.8 35.0

    Mean AP across all 20 classes is 26.6%26.6\% with exact pyramids and 24.5%24.5\% with fast feature pyramids, demonstrating that the power-law approximation generalizes to multi-component deformable models with a minor (2.1%2.1\%) drop in detection performance.

  9. Knowl 9 — Detector Rescaling via Feature Power Law in Integral Channel Features

    model/method

    In Integral Channel Features (ICF), feature values are defined as rectangular sums over channel maps: f(I)=AfΩ(I)f(I) = A f_\Omega(I), where AA is the rectangle area and fΩ(I)f_\Omega(I) is the average channel value. Rather than rescaling channel images, detector responses at scale offset α=s/s′\alpha = s/s' can be approximated directly from integral channels computed at reference scale s′s' by rescaling the classifier itself:

    1. Each rectangular feature bounding box of width ww and height hh is rescaled to dimensions αw×αh\alpha w \times \alpha h.
    2. The rectangle sum evaluated on the reference integral image is scaled by α−λΩ\alpha^{-\lambda_\Omega}.

    A hybrid implementation computes exact integral feature channels once per octave and scales the detector model to evaluate detection windows within half an octave of each reference level.

  10. Knowl 10 — Scope and Failure Modes of Fast Feature Pyramid Approximation

    limitation

    The power-law feature approximation relies on the broad fractal spatial spectra of natural scenes and fails or degrades in three specific regimes:

    1. Narrow band-pass spectra: Highly structured, periodic textures (such as certain Brodatz textures) with high-frequency energy do not follow the scale-invariance power law, causing significant underestimation of feature energy upon downsampling.
    2. Highly sparse channels: Channels where the feature value is frequently near zero (fΩ(I)≈0f_\Omega(I) \approx 0, e.g., the output of a sparse classifier) exhibit high variance in the scaling ratio and cannot be reliably extrapolated.
    3. Extreme scale regimes: Scales far outside natural optical viewing ranges (e.g., severe microscopic magnification) break the scale stationarity of image statistics.

Coverage note — None was omitted; all contributed theory, equations, algorithms, detector models (ACF, ICF, DPM), empirical scaling parameter fits, and full benchmark evaluation tables from the paper are represented.

References

  1. 1.D. Hubel and T. Wiesel, “Receptive fields and functional architecture of monkey striate cortex,” Journal of Physiology, 1968.
  2. 2.C. Malsburg, “Self-organization of orientation sensitive cells in the striate cortex,” Biological Cybernetics, vol. 14, no. 2, 1973.
  3. 3.L. Maffei and A. Fiorentini, “The visual cortex as a spatial frequency analyser,” Vision Research, vol. 13, no. 7, 1973.
  4. 4.P. Burt and E. Adelson, “The laplacian pyramid as a compact image code,” IEEE Transactions on Communications, 1983.
  5. 5.J. Daugman, “Uncertainty relation for resolution in space, spatial frequency, and orientation optimized by two-dimensional visual cortical filters,” Journal of the Optical Society of America A, 1985.
  6. 6.J. Koenderink and A. Van Doorn, “Representation of local geometry in the visual system,” Biological cybernetics, vol. 55, no. 6, 1987.
  7. 7.D. J. Field, “Relations between the statistics of natural images and the response properties of cortical cells,” Journal of the Optical Society of America A, vol. 4, pp. 2379–2394, 1987.
  8. 8.S. Mallat, “‘A theory for multiresolution signal decomposition: the wavelet representation,” PAMI, vol. 11, no. 7, 1989.
  9. 9.P. Vaidyanathan, “‘Multirate digital filters, filter banks, polyphase networks, and applications: A tutorial,” Proceedings of the IEEE, vol. 78, no. 1, 1990.
  10. 10.M. Vetterli, “‘A theory of multirate filter banks,” IEEE Conference on Acoustics, Speech and Signal Processing, vol. 35, no. 3, 1987.
  11. 11.E. Simoncelli and E. Adelson, “Noise removal via bayesian wavelet coring,” in ICIP, vol. 1, 1996.
  12. 12.W. T. Freeman and E. H. Adelson, “The design and use of steerable filters,” PAMI, vol. 13, pp. 891–906, 1991.
  13. 13.J. Malik and P. Perona, “‘Preattentive texture discrimination with early vision mechanisms,” Journal of the Optical Society of America A, vol. 7, pp. 923–932, May 1990.
  14. 14.D. Jones and J. Malik, “Computational framework for determining stereo correspondence from a set of linear spatial filters,” Image and Vision Computing, vol. 10, no. 10, pp. 699–708, 1992.
  15. 15.E. Adelson and J. Bergen, “Spatiotemporal energy models for the perception of motion,” Journal of the Optical Society of America A, vol. 2, no. 2, pp. 284–299, 1985.
  16. 16.Y. Weiss and E. Adelson, “‘A unified mixture framework for motion segmentation: Incorporating spatial coherence and estimating the number of models,” in CVPR, 1996.
  17. 17.L. Itti, C. Koch, and E. Niebur, “‘A model of saliency-based visual attention for rapid scene analysis,” PAMI, vol. 20, no. 11, 1998.
  18. 18.P. Perona and J. Malik, “Detecting and localizing edges composed of steps, peaks and roofs,” in ICCV, 1990.
  19. 19.M. Lades, J. Vorbruggen, J. Buhmann, J. Lange, C. von der Malsburg, R. Wurtz, and W. Konen, “‘Distortion invariant object recognition in the dynamic link architecture,” IEEE Transactions on Computers, vol. 42, no. 3, pp. 300–311, 1993.
  20. 20.D. G. Lowe, “‘Object recognition from local scale-invariant features,” in ICCV, 1999.
  21. 21.N. Dalal and B. Triggs, “‘Histograms of oriented gradients for human detection,” in CVPR, 2005.
  22. 22.R. De Valois, D. Albrecht, and L. Thorell, “‘Spatial frequency selectivity of cells in macaque visual cortex,” Vision Research, vol. 22, no. 5, pp. 545–559, 1982.
  23. 23.Y. LeCun, P. Haffner, L. Bottou, and Y. Bengio, “‘Gradient-based learning applied to document recognition,” in Proc. of IEEE, 1998.
  24. 24.M. Riesenhuber and T. Poggio, “‘Hierarchical models of object recognition in cortex,” Nature Neuroscience, vol. 2, 1999.
  25. 25.D. G. Lowe, “‘Distinctive image features from scale-invariant keypoints,” IJCV, vol. 60, no. 2, pp. 91–110, 2004.
  26. 26.A. Krizhevsky, I. Sutskever, and G. Hinton, “Imagenet classification with deep convolutional neural networks,” in NIPS, 2012.
  27. 27.P. Viola and M. Jones, “‘Rapid object detection using a boosted cascade of simple features,” in CVPR, 2001.
  28. 28.P. Viola, M. Jones, and D. Snow, “Detecting pedestrians using patterns of motion and appearance,” IJCV, vol. 63(2), 2005.
  29. 29.P. Doll'ar, Z. Tu, P. Perona, and S. Belongie, “Integral channel features,” in BMVC, 2009.
  30. 30.R. Benenson, M. Mathias, R. Timofte, and L. Van Gool, “‘Pedestrian detection at 100 frames per second,” in CVPR, 2012.
  31. 31.P. Doll'ar, C. Wojek, B. Schiele, and P. Perona, “‘Pedestrian detection: An evaluation of the state of the art,” PAMI, vol. 99, 2011.
  32. 32.www.vision.caltech.edu/Image_Datasets/CaltechPedestrians/.
  33. 33.D. L. Ruderman and W. Bialek, “Statistics of natural images: Scaling in the woods,” Physical Review Letters, vol. 73, no. 6, pp. 814–817, Aug 1994.
  34. 34.E. Switkes, M. Mayer, and J. Sloan, “‘Spatial frequency analysis of the visual environment: anisotropy and the carpentered environment hypothesis,” Vision Research, vol. 18, no. 10, 1978.
  35. 35.P. Felzenszwalb, R. Girshick, D. McAllester, and D. Ramanan, “‘Object detection with discriminatively trained part based models,” PAMI, vol. 32, no. 9, pp. 1627–1645, 2010.
  36. 36.C. Wojek, S. Walk, and B. Schiele, “‘Multi-cue onboard pedestrian detection,” in CVPR, 2009.
  37. 37.A. Ess, B. Leibe, and L. Van Gool, “‘Depth and appearance for mobile scene analysis,” in ICCV, 2007.
  38. 38.M. Everingham, L. Van Gool, C. K. I. Williams, J. Winn, and A. Zisserman, “‘The PASCAL visual object classes (VOC) challenge,” IJCV, vol. 88, no. 2, pp. 303–338, Jun. 2010.
  39. 39.P. Doll'ar, S. Belongie, and P. Perona, “‘The fastest pedestrian detector in the west,” in BMVC, 2010.
  40. 40.P. Doll'ar, R. Appel, and W. Kienzle, “‘Crosstalk cascades for frame-rate pedestrian detection,” in ECCV, 2012.
  41. 41.T. Lindeberg, “Scale-space for discrete signals,” PAMI, vol. 12, no. 3, pp. 234–254, 1990.
  42. 42.J. L. Crowley, O. Riff, and J. H. Piater, “‘Fast computation of characteristic scale using a half-octave pyramid,” in International Conference on Scale-Space Theories in Computer Vision, 2002.
  43. 43.R. S. Eaton, M. R. Stevens, J. C. McBride, G. T. Foil, and M. S. Snorrason, “‘A systems view of scale space,” in ICVS, 2006.
  44. 44.P. Felzenszwalb, R. Girshick, and D. McAllester, “‘Cascade object detection with deformable part models,” in CVPR, 2010.
  45. 45.M. Pedersoli, A. Vedaldi, and J. Gonzalez, “‘A coarse-to-fine approach for fast deformable object detection,” in CVPR, 2011.
  46. 46.C. H. Lampert, M. B. Blaschko, and T. Hofmann, “‘Efficient subwindow search: A branch and bound framework for object localization,” PAMI, vol. 31, pp. 2129–2142, Dec 2009.
  47. 47.L. Bourdev and J. Brandt, “‘Robust object detection via soft cascade,” in CVPR, 2005.
  48. 48.C. Zhang and P. Viola, “‘Multiple-instance pruning for learning efficient cascade detectors,” in NIPS, 2007.
  49. 49.J. Šochman and J. Matas, “‘Waldboost - learning for time constrained sequential detection,” in CVPR, 2005.
  50. 50.H. Masnadi-Shirazi and N. Vasconcelos, “‘High detection-rate cascades for real-time object detection,” in ICCV, 2007.
  51. 51.F. Fleuret and D. Geman, “‘Coarse-to-fine face detection,” IJCV, vol. 41, no. 1-2, pp. 85–107, 2001.
  52. 52.P. Felzenszwalb and D. Huttenlocher, “‘Efficient matching of pictorial structures,” in CVPR, 2000.
  53. 53.C. Papageorgiou and T. Poggio, “‘A trainable system for object detection,” IJCV, vol. 38, no. 1, pp. 15–33, 2000.
  54. 54.M. Weber, M. Welling, and P. Perona, “‘Unsupervised learning of models for recognition,” in ECCV, 2000.
  55. 55.S. Agarwal and D. Roth, “‘Learning a sparse representation for object detection,” in ECCV, 2002.
  56. 56.R. Fergus, P. Perona, and A. Zisserman, “‘Object class recognition by unsupervised scale-invariant learning,” in CVPR, 2003.
  57. 57.B. Leibe, A. Leonardis, and B. Schiele, “‘Robust object detection with interleaved categorization and segmentation,” IJCV, vol. 77, no. 1-3, pp. 259–289, May 2008.
  58. 58.C. Gu, J. J. Lim, P. Arbelaez, and J. Malik, “‘Recognition using regions,” in CVPR, 2009.
  59. 59.B. Alexe, T. Deselaers, and V. Ferrari, “What is an object?” in CVPR, 2010.
  60. 60.C. Wojek, G. Dorko, A. Schulz, and B. Schiele, “‘Sliding-windows for rapid object class localization: A parallel technique,” in DAGM, 2008.
  61. 61.L. Zhang and R. Nevatia, “‘Efficient scan-window based object detection using gpgpu,” in Visual Computer Vision on GPU’s (CVGPU), 2008.
  62. 62.B. Bilgic, “‘Fast human detection with cascaded ensembles,” Master’s thesis, MIT, February 2010.
  63. 63.Q. Zhu, S. Avidan, M. Yeh, and K. Cheng, “‘Fast human detection using a cascade of histograms of oriented gradients,” in CVPR, 2006.
  64. 64.F. M. Porikli, “Integral histogram: A fast way to extract histograms in cartesian spaces,” in CVPR, 2005.
  65. 65.A. Ess, B. Leibe, K. Schindler, and L. Van Gool, “‘Robust multi-person tracking from a mobile platform,” PAMI, vol. 31, pp. 1831–1846, 2009.
  66. 66.M. Bajracharya, B. Moghaddam, A. Howard, S. Brennan, and L. H. Matthies, “‘A fast stereo-based system for detecting and tracking pedestrians from a moving vehicle,” The International Journal of Robotics Research, vol. 28, 2009.
  67. 67.D. L. Ruderman, “The statistics of natural images,” Network: Computation in Neural Systems, vol. 5, no. 4, pp. 517–548, 1994.
  68. 68.S. G. Ghurye, “‘A characterization of the exponential function,” The American Mathematical Monthly, vol. 64, no. 4, 1957.
  69. 69.J. Friedman, T. Hastie, and R. Tibshirani, “‘Additive logistic regression: a statistical view of boosting,” The Annals of Statistics, vol. 38, no. 2, pp. 337–374, 2000.
  70. 70.P. Sabzmeydani and G. Mori, “‘Detecting pedestrians by learning shapelet features,” in CVPR, 2007.
  71. 71.Z. Lin and L. S. Davis, “‘A pose-invariant descriptor for human detection and segmentation,” in ECCV, 2008.
  72. 72.S. Maji, A. Berg, and J. Malik, “‘Classification using intersection kernel SVMs is efficient,” in CVPR, 2008.
  73. 73.X. Wang, T. X. Han, and S. Yan, “‘An hog-lbp human detector with partial occlusion handling,” in ICCV, 2009.
  74. 74.C. Wojek and B. Schiele, “‘A performance evaluation of single and multi-feature people detection,” in DAGM, 2008.
  75. 75.W. Schwartz, A. Kembhavi, D. Harwood, and L. Davis, “‘Human detection using partial least squares analysis,” in ICCV, 2009.
  76. 76.S. Walk, N. Majer, K. Schindler, and B. Schiele, “‘New features and insights for pedestrian detection,” in CVPR, 2010.
  77. 77.D. Park, D. Ramanan, and C. Fowlkes, “‘Multiresolution models for object detection,” in ECCV, 2010.

Citation

MLA
Dollar, P., et al. “Fast Feature Pyramids for Object Detection”. IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 36, no. 8, 2014, pp. 1532–45, https://doi.org/10.1109/TPAMI.2014.2300479.
APA
Dollar, P., Appel, R., Belongie, S., & Perona, P. (2014). Fast Feature Pyramids for Object Detection. IEEE Transactions on Pattern Analysis and Machine Intelligence, 36(8), 1532–1545. https://doi.org/10.1109/TPAMI.2014.2300479
Chicago
Dollar, P., R. Appel, S. Belongie, and P. Perona. 2014. “Fast Feature Pyramids for Object Detection”. IEEE Transactions on Pattern Analysis and Machine Intelligence 36 (8): 1532–45. https://doi.org/10.1109/TPAMI.2014.2300479.
Harvard
Dollar, P. et al. (2014) “Fast Feature Pyramids for Object Detection”, IEEE Transactions on Pattern Analysis and Machine Intelligence, 36(8), pp. 1532–1545. Available at: https://doi.org/10.1109/TPAMI.2014.2300479.
Vancouver
1. Dollar P, Appel R, Belongie S, Perona P (2014) Fast Feature Pyramids for Object Detection. IEEE Transactions on Pattern Analysis and Machine Intelligence 36:1532–1545

BibTeX

@article{Dollar_2014, title={Fast Feature Pyramids for Object Detection}, volume={36}, ISSN={2160-9292}, url={http://dx.doi.org/10.1109/TPAMI.2014.2300479}, DOI={10.1109/tpami.2014.2300479}, number={8}, journal={IEEE Transactions on Pattern Analysis and Machine Intelligence}, publisher={Institute of Electrical and Electronics Engineers (IEEE)}, author={Dollar, Piotr and Appel, Ron and Belongie, Serge and Perona, Pietro}, year={2014}, month=Aug, pages={1532–1545} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF