Pedestrian Detection: An Evaluation of the State of the Art

Piotr DollárChristian WojekBernt SchielePietro Perona

article2012TPAMI3,360 citations

Establishes a standardized evaluation protocol and the large-scale Caltech Pedestrian Dataset while benchmarking sixteen detectors across six datasets to identify critical performance bottlenecks for small and occluded pedestrians.

Listen

This paper evaluates the current state of pedestrian detection algorithms in computer vision through a unified benchmark. Pedestrian detection supports applications such as automotive safety, robotics, surveillance, and elderly care, yet nearly 5,000 pedestrian fatalities occur annually in the United States alone. Prior work suffered from inconsistent datasets and evaluation protocols that prevented reliable comparisons of methods or identification of failure modes.

The authors set out to quantify detector performance, rank leading approaches, expose main weaknesses, and outline productive research directions. They assembled the Caltech Pedestrian Dataset, which contains roughly 350,000 annotated bounding boxes across 250,000 video frames recorded from a vehicle in urban traffic. They also developed a refined per-frame evaluation protocol that measures performance by scale and occlusion level while standardizing bounding-box aspect ratios and using expanded filtering to avoid over- or under-counting errors. Sixteen pre-trained detectors were then tested on this dataset plus five others (INRIA, ETH, TUD-Brussels, Daimler-DB, and Caltech-Japan) under identical conditions.

The study reveals that detection accuracy has advanced steadily yet remains far from adequate for practical use. Log-average miss rates exceed 80 percent across all annotated pedestrians and stay above 50 percent even for clearly visible pedestrians at least 50 pixels tall. Performance drops sharply below 80-pixel height and with any occlusion, and it collapses entirely at far scales or under heavy occlusion. Detectors that combine gradient histograms with additional cues such as texture, color self-similarity, or motion perform best, while part-based and motion-aware methods gain the most from higher resolution. Detector rankings prove reasonably stable across datasets, and statistical tests confirm that the top few methods are not significantly separable given current sample sizes.

These results indicate that current algorithms cannot yet meet the requirements of safety-critical systems that must operate at medium scales and tolerate partial occlusion. The mismatch between typical research focus on high-resolution, unoccluded cases and real-world operating conditions is pronounced. The authors recommend directing effort toward improved handling of medium-scale pedestrians, occlusion patterns, motion features at low resolution, temporal integration, contextual cues, and training on larger, more varied data. They also note that faster implementations and region-of-interest selection will be essential for deployment.

The evaluation relies on pre-trained detectors rather than retraining on the new dataset, and computational limits restricted testing to every thirtieth frame. Nevertheless, the breadth of detectors, datasets, and conditions, together with public release of all code and annotations, supports high confidence in the reported rankings and the identified gaps.

Cover for Pedestrian Detection: An Evaluation of the State of the Art

Abstract

Pedestrian detection is a key problem in computer vision, with several applications that have the potential to positively impact quality of life. In recent years, the number of approaches to detecting pedestrians in monocular images has grown steadily. However, multiple datasets and widely varying evaluation protocols are used, making direct comparisons difficult. To address these shortcomings, we perform an extensive evaluation of the state of the art in a unified framework. We make three primary contributions: (1) we put together a large, well-annotated and realistic monocular pedestrian detection dataset and study the statistics of the size, position and occlusion patterns of pedestrians in urban scenes, (2) we propose a refined per-frame evaluation methodology that allows us to carry out probing and informative comparisons, including measuring performance in relation to scale and occlusion, and (3) we evaluate the performance of sixteen pre-trained state-of-the-art detectors across six datasets. Our study allows us to assess the state of the art and provides a framework for gauging future efforts. Our experiments show that despite significant progress, performance still has much room for improvement. In particular, detection is disappointing at low resolutions and for partially occluded pedestrians.

Table of Contents

  • 1 INTRODUCTION
  • 1.1 Contributions
  • 2 THE CALTECH PEDESTRIAN DATASET
  • 2.1 Data Collection and Ground Truthing
  • 2.2 Dataset Statistics
  • 2.2.1 Scale Statistics
  • 2.2.2 Occlusion Statistics
  • 2.2.3 Position Statistics
  • 2.3 Training and Testing Data
  • 2.4 Comparison of Pedestrian Datasets
  • 3 EVALUATION METHODOLOGY
  • 3.1 Full Image Evaluation
  • 3.2 Filtering Ground Truth
  • 3.3 Filtering Detections
  • 3.4 Standardizing Aspect Ratios
  • 3.5 Per-Window Versus Full Image Evaluation
  • 4 DETECTION ALGORITHMS
  • 4.1 Survey of the State of the Art
  • 4.2 Evaluated Detectors
  • 5 PERFORMANCE EVALUATION
  • 5.1 Performance on the Caltech Dataset
  • 5.2 Evaluation on Multiple Datasets
  • 5.3 Statistical Significance
  • 5.4 Runtime Analysis
  • 6 DISCUSSION
  • 6.1 Statistics of The Caltech Pedestrian Dataset
  • 6.2 Overall Performance
  • 6.3 State of the Art Detectors
  • 6.4 Research directions
  • Acknowledgments
  • REFERENCES

Knowls

  1. Knowl 1 — Statistics and Properties of the Caltech Pedestrian Dataset

    experimental setup

    The Caltech Pedestrian Dataset comprises approximately 10 hours of 640×480640 \times 480 30 Hz color video (106\,\sim 10^6 frames) recorded from a vehicle driving through urban traffic across five Los Angeles metropolitan areas, stabilized using an inverse compositional image alignment algorithm. It includes 250,000 annotated frames containing 350,000 pedestrian bounding boxes spanning 2,300\,\sim 2,300 unique individuals.

    Key geometric, scale, and occlusion properties include:

    • Bounding Box Height Distribution: Pedestrian heights follow a log-normal distribution with a geometric mean (log-average) of 50 pixels and a median of 48 pixels. Heights are partitioned into three scales: Near (ge80\\ge 80 pixels, 16% of instances), Medium (308030\text{--}80 pixels, 69% of instances), and Far (le30\\le 30 pixels, 15% of instances).

    • Pinhole Distance Relation: For a camera with pixel focal length f1000f \approx 1000 pixels and a pedestrian of height H1.8 mH \approx 1.8\text{ m}, observed pixel height hh relates to distance dd via d1800/h md \approx 1800/h\text{ m}. At an urban vehicle speed of 55 km/h55\text{ km/h}, an 80-pixel pedestrian is 1.5 s away, whereas a 30-pixel pedestrian is 4.0 s away, making the medium scale the most critical range for automotive collision avoidance.

    • Aspect Ratio: Bounding box aspect ratios (w/hw/h) follow a log-normal distribution with a log-average value of μ=0.41\mu = 0.41.

    • Occlusion Characteristics: Two bounding boxes are labeled per occluded pedestrian: the full extent estimate (BBfullBB_{\text{full}}) and the visible region (BBvisBB_{\text{vis}}). Over 70% of pedestrians are occluded in at least one frame (29% never occluded, 53% occluded in some frames, 19% occluded in all frames). Occlusion spatial probability is heavily biased toward the lower body (feet occluded first, head remaining visible). Quantizing bounding box occlusions reveals that only 7 discrete spatial configurations account for 97%\,\sim 97\% of all partial occlusions in urban driving.

  2. Knowl 2 — Log-Average Miss Rate Evaluation Metric

    model/method

    Single-frame pedestrian detection performance is evaluated using a full-image protocol. A detected bounding box BBdtBB_{\text{dt}} matches a ground truth bounding box BBgtBB_{\text{gt}} if their intersection-over-union overlap aoa_o exceeds 0.5:

    ao=area(BBdtBBgt)area(BBdtBBgt)>0.5a_o = \frac{\text{area}(BB_{\text{dt}} \cap BB_{\text{gt}})}{\text{area}(BB_{\text{dt}} \cup BB_{\text{gt}})} > 0.5

    Matching is performed greedily in descending order of detection confidence, with ties broken by maximum overlap. Each detection and ground truth box may be matched at most once. Unmatched detections count as false positives; unmatched ground truth boxes count as false negatives.

    Detector performance curves plot miss rate against false positives per image (FPPI) on log-log axes by sweeping the confidence threshold. To summarize the curve into a single scalar value, the log-average miss rate (LAMR) is computed by averaging the miss rate at 9 FPPI values evenly spaced in log-space between 10210^{-2} and 10010^0:

    LAMR=19i=08MissRate(102+0.25i)\text{LAMR} = \frac{1}{9} \sum_{i=0}^{8} \text{MissRate}\left(10^{-2 + 0.25 i}\right)

    If a detection curve terminates before reaching a given FPPI threshold, the minimum miss rate achieved by that detector across all tested thresholds is used for all remaining evaluation points.

  3. Knowl 3 — Ground Truth Ignore Regions and Subregion Matching

    model/method

    To evaluate detectors on specific pedestrian subsets (e.g., unoccluded individuals) or prevent ambiguous image regions from skewing results, annotations can be marked as ignore regions (BBigBB_{\text{ig}}). Simply removing ground truth annotations from evaluation causes valid detections in those areas to be penalized as false positives. Instead, ignore regions are evaluated with an asymmetric, subregion-matching criterion:

    1. Ambiguous cases (crowds labeled People, doubtful pedestrians labeled Person?, bounding boxes under 20 pixels in height, or bounding boxes truncated by image boundaries) are set to BBigBB_{\text{ig}}.

    2. Detections BBdtBB_{\text{dt}} are matched against valid ground truth boxes BBgtBB_{\text{gt}} first.

    3. An unmatched detection BBdtBB_{\text{dt}} matches an ignore region BBigBB_{\text{ig}} if its area of overlap relative to the detection area exceeds 0.5:

    ao=area(BBdtBBig)area(BBdt)>0.5a_o = \frac{\text{area}(BB_{\text{dt}} \cap BB_{\text{ig}})}{\text{area}(BB_{\text{dt}})} > 0.5

    1. Detections matching BBigBB_{\text{ig}} do not count as true positives and do not count as false positives.

    2. Multiple detections are allowed to match a single BBigBB_{\text{ig}}, and unmatched BBigBB_{\text{ig}} do not count as false negatives.

  4. Knowl 4 — Expanded Filtering Protocol for Sub-Scale Evaluation

    model/method

    When evaluating pedestrian detectors within a restricted ground truth height range [S0,S1][S_0, S_1] (in pixels), standard detection filtering methods introduce systematic evaluation biases:

    • Strict filtering removes all detections outside [S0,S1][S_0, S_1] prior to matching. If a ground truth object of height within [S0,S1][S_0, S_1] is matched by a detection of height slightly outside that interval (e.g., S0ϵS_0 - \epsilon), strict filtering discards the detection and creates an artificial false negative, under-reporting detector performance.

    • Post filtering matches all detections across all scales to ground truth within [S0,S1][S_0, S_1] and then removes all unmatched detections falling outside [S0,S1][S_0, S_1]. This undercounts false positives and over-reports performance because generating dense spurious hypotheses outside [S0,S1][S_0, S_1] increases true positive matches without incurring false positive penalties.

    • Expanded filtering resolves both biases by discarding detections only if their height falls outside an expanded range [S0/r,S1r][S_0 / r, \, S_1 \cdot r] with expansion factor r=1.25r = 1.25 prior to matching against ground truth restricted to [S0,S1][S_0, S_1] (with ground truth outside [S0,S1][S_0, S_1] set to ignore). Matches to ground truth outside [S0,S1][S_0, S_1] count as neither true nor false positives. Detector rankings are robust to the choice of rr in this formulation.

  5. Knowl 5 — Standardization of Bounding Box Aspect Ratios

    model/method

    Pedestrian bounding box widths exhibit high variance due to pose changes (such as limb movement during walking strides) and differing dataset or detector conventions (detector aspect ratios w/hw/h vary from 0.34 to 0.50). In contrast, height is a stable indicator of scale and distance.

    To eliminate the effect of arbitrary bounding box width variations during benchmarking, all ground truth and detected bounding boxes are standardized to a constant aspect ratio of w/h=0.41w/h = 0.41 (matching the log-mean aspect ratio of pedestrians in the Caltech dataset).

    For any bounding box with original height hh, center coordinates (xc,yc)(x_c, y_c), and width ww, standardization fixes the height and center while redefining width and horizontal boundaries:

    wstd=0.41h,xleft=xcwstd2w_{\text{std}} = 0.41 \cdot h, \quad x_{\text{left}} = x_c - \frac{w_{\text{std}}}{2}

    Standardizing aspect ratios ensures consistent overlap area computation without altering detector localization accuracy.

  6. Knowl 6 — Discrepancy Between Per-Window and Full-Image Evaluation

    empirical result

    Per-window (PW) evaluation measures classifier error on pre-cropped positive and negative image patches, isolating the binary classifier from the full detection pipeline. Full-image evaluation tests the complete multiscale detector across full scenes.

    PW performance correlates only weakly with full-image detection performance and leads to substantially different rankings among detectors for several reasons:

    • Sampling and Boundary Artifacts: Classifiers evaluated on pre-cropped image windows can exploit subtle boundary artifacts created by differing positive and negative crop generation procedures, causing severe classifier overfitting that does not generalize to full images.

    • System Parameters Omission: PW ignores spatial stride, scale stride, scale interpolation, and non-maximum suppression (NMS), all of which heavily affect end-to-end performance.

    • Untested Full-Image Error Cases: PW fails to evaluate multiple common detection failure modes, including false positives on body subparts, false positives at incorrect scales or positions, and false negatives caused by spatial sliding window misalignments or aggressive NMS suppression.

  7. Knowl 7 — Performance and Runtime Benchmark of Sixteen Pedestrian Detectors

    data/table

    Sixteen state-of-the-art pedestrian detectors were evaluated under the Reasonable setting (pedestrians 50\ge 50 pixels tall under no or partial [135%][1\text{--}35\%] occlusion) on the Caltech Pedestrian Dataset test set. Runtimes on 640×480640 \times 480 images were normalized to a single modern processor for detection scales 100\ge 100 pixels and 50\ge 50 pixels.

    Detector Feature Categories Classifier NMS Speed 100\ge 100px (fps) Speed 50\ge 50px (fps) Reasonable LAMR
    MultiFtr+Motion GradHist+Grad+Motion Linear SVM Mean Shift 0.020 0.004 51%
    ChnFtrs GradHist+Grad+LUV AdaBoost Pairwise Max* 1.183 0.278 56%
    FPDW GradHist+Grad+LUV AdaBoost Pairwise Max* 6.492 2.670 57%
    FeatSynth Part-Synthesized Linear SVM 60%
    MultiFtr+CSS GradHist+Grad+CSS Linear SVM Mean Shift 0.027 0.005 61%
    Pls Grad+Texture+Color PLS+QDA Pairwise Max* 0.018 0.005 62%
    LatSvm-V2 GradHist (Deformable Parts) Latent SVM Pairwise Max 0.629 0.164 63%
    HogLbp GradHist+LBP Linear SVM Mean Shift 0.062 0.014 68%
    MultiFtr GradHist+Grad AdaBoost Mean Shift 0.072 0.017 68%
    HOG GradHist Linear SVM Mean Shift 0.239 0.054 68%
    HikSvm GradHist HIK SVM Mean Shift 0.185 0.036 73%
    FtrMine GradHist+Grad+Haar+Color AdaBoost Pairwise Max 0.080 0.020 74%
    LatSvm-V1 GradHist (Deformable Parts) Latent SVM Pairwise Max 0.392 0.098 80%
    PoseInv GradHist+Shapelet AdaBoost Mean Shift 0.474 0.101 86%
    Shapelet Shapelet Gradients AdaBoost Mean Shift 0.051 0.010 91%
    VJ Haar Grayscale AdaBoost Mean Shift 0.447 0.089 95%

    The data shows that multi-cue feature integration (such as combining gradient histograms with color, texture, or optical flow) consistently outperforms single-feature models (e.g., pure HOG at 68% LAMR vs MultiFtr+Motion at 51%). Fast multiscale approximation in FPDW provides a >5×>5\times speedup over ChnFtrs with virtually identical accuracy (57% vs 56% LAMR).

  8. Knowl 8 — Performance Degradation Across Scale Ranges and Occlusion Levels

    empirical result

    Empirical evaluation on the Caltech Pedestrian Dataset demonstrates severe performance degradation of state-of-the-art detectors when evaluated across scale and occlusion subsets:

    • Scale Degradation: On unoccluded near-scale pedestrians (height 80\ge 80 pixels), the best detector (MultiFtr+Motion) achieves a 22% log-average miss rate, and multiple detectors achieve 3040%30\text{--}40\%. On medium-scale pedestrians (308030\text{--}80 pixels, which constitute 69%\sim 69\% of all instances), log-average miss rates degrade sharply to 77%ext80%77\% ext{--}80\%, meaning that at an operating point of 1 false alarm per 10 images, roughly 80% of pedestrians are missed. At far scales (<30< 30 pixels), all detectors fail, exhibiting 95%\ge 95\% log-average miss rates. Detectors relying on motion, texture, and part models suffer the largest relative performance drops at small scales.

    • Occlusion Degradation: On pedestrians 50\ge 50 pixels high, partial occlusion (135%1\text{--}35\% area occluded) increases the log-average miss rate of top detectors from 48%ext54%48\% ext{--}54\% (unoccluded) to 73%73\%. Part-based models (such as LatSvm) degrade as severely under partial occlusion as holistic sliding-window models. Under heavy occlusion (3580%35\text{--}80\% occluded), all evaluated detectors fail catastrophically (ge93%\\ge 93\% log-average miss rate).

  9. Knowl 9 — Cross-Dataset Detector Ranking and Statistical Significance Analysis

    empirical result

    Evaluation of 14 pedestrian detectors across 28 data folds (11 Caltech folds, 13 Caltech-Japan folds, 3 ETH sequences, and 1 TUD-Brussels sequence under the Reasonable setting) using the non-parametric Friedman test with Shaffer post-hoc analysis at significance level α=0.05\alpha = 0.05 demonstrates:

    • Mean Detector Ranks: MultiFtr+Motion achieved the top mean rank of 2.4 (ranking 1st in 17 of 28 folds), followed by ChnFtrs (mean rank 3.3), FPDW (3.8), Pls (4.5), MultiFtr+CSS (5.2), LatSvm-V2 (5.4), HogLbp (6.6), MultiFtr (6.6), HOG (8.2), HikSvm (9.8), LatSvm-V1 (10.8), PoseInv (12.2), Shapelet (12.9), and VJ (13.5).

    • Statistical Significance: MultiFtr+Motion, ChnFtrs, and FPDW are statistically significantly superior to baseline HOG (p<0.05p < 0.05). However, pairwise differences among the top six detectors (MultiFtr+Motion, ChnFtrs, FPDW, Pls, MultiFtr+CSS, LatSvm-V2) are not statistically significant.

    • Cross-Dataset Consistency: Relative detector rankings remain largely consistent across diverse datasets, although absolute dataset difficulty varies: INRIA contains high-resolution pedestrians and yields the best results (20--22% LAMR for top models), followed by Daimler-DB (29% LAMR), ETH/TUD-Brussels/Caltech (51--55% LAMR), and Caltech-Japan (64% LAMR).

Coverage note — Detailed literature summaries of previous individual pedestrian detection architectures from the survey in Section 4.1 (which represent prior work) and the interactive cubic-interpolation user interface implementation details of the video labeling tool were omitted as they are not primary algorithmic/empirical contributions of this benchmark paper.

References

  1. 1.U. Shankar, ‘‘Pedestrian roadway fatalities,’’ Department of Transportation, Tech. Rep., 2003.
  2. 2.D. Geronimo, A. M. Lopez, A. D. Sappa, and T. Graf, ‘‘Survey on pedestrian detection for advanced driver assistance systems,’’ IEEE Trans. Pattern Analysis and Machine Intelligence, vol. 32, no. 7, pp. 1239–1258, 2010.
  3. 3.P. Doll'ar, C. Wojek, B. Schiele, and P. Perona, ‘‘Pedestrian Detection: A Benchmark,’’ in IEEE Conf. Computer Vision and Pattern Recognition, 2009.
  4. 4.A. Ess, B. Leibe, and L. Van Gool, ‘‘Depth and appearance for mobile scene analysis,’’ in IEEE Intl. Conf. Computer Vision, 2007.
  5. 5.C. Wojek, S. Walk, and B. Schiele, ‘‘Multi-cue onboard pedestrian detection,’’ in IEEE Conf. Computer Vision and Pattern Recognition, 2009.
  6. 6.M. Enzweiler and D. M. Gavrila, ‘‘Monocular pedestrian detection: Survey and experiments,’’ IEEE Trans. Pattern Analysis and Machine Intelligence, 2009.
  7. 7.N. Dalal and B. Triggs, ‘‘Histograms of oriented gradients for human detection,’’ in IEEE Conf. Computer Vision and Pattern Recognition, 2005.
  8. 8.J. L. Barron, D. J. Fleet, S. S. Beauchemin, and T. A. Burkitt, ‘‘Performance of optical flow techniques,’’ Intl. Journal of Computer Vision, vol. 12, no. 1, pp. 43–77, 1994.
  9. 9.S. Baker, D. Scharstein, J. Lewis, S. Roth, M. Black, and R. Szeliski, ‘‘A database and eval. methodology for optical flow,’’ in IEEE Intl. Conf. Computer Vision, 2007.
  10. 10.D. Martin, C. Fowlkes, and J. Malik, ‘‘Learning to detect natural image boundaries using local brightness, color, and texture cues,’’ IEEE Trans. Pattern Analysis and Machine Intelligence, vol. 26, no. 5, pp. 530–549, 2004.
  11. 11.D. Scharstein and R. Szeliski, ‘‘A taxonomy and evaluation of dense two-frame stereo correspondence algorithms,’’ Intl. Journal of Computer Vision, vol. 47, pp. 7–42, 2002.
  12. 12.L. Fei-Fei, R. Fergus, and P. Perona, ‘‘One-shot learning of object categories,’’ IEEE Trans. Pattern Analysis and Machine Intelligence, vol. 28, no. 4, pp. 594–611, 2006.
  13. 13.G. Griffin, A. Holub, and P. Perona, ‘‘Caltech-256 object category dataset,’’ California Inst. of Technology, Tech. Rep. 7694, 2007.
  14. 14.M. Everingham, L. Van Gool, C. K. I. Williams, J. Winn, and A. Zisserman, ‘‘The PASCAL visual object classes (VOC) challenge,’’ Intl. Journal of Computer Vision, vol. 88, no. 2, pp. 303–338, Jun. 2010.
  15. 15.S. Baker and I. Matthews, ‘‘Lucas-kanade 20 years on: A unifying framework,’’ Intl. Journal of Computer Vision, vol. 56, no. 3, pp. 221–255, 2004.
  16. 16.C. Papageorgiou and T. Poggio, ‘‘A trainable system for object detection,’’ Intl. Journal of Computer Vision, vol. 38, no. 1, pp. 15–33, 2000.
  17. 17.B. Wu and R. Nevatia, ‘‘Detection of multiple, partially occluded humans in a single image by bayesian combination of edgelet part det.’’ in IEEE Intl. Conf. Computer Vision, 2005.
  18. 18.——, ‘‘Cluster boosted tree classifier for multi-view, multi-pose object det.’’ in IEEE Intl. Conf. Computer Vision, 2007.
  19. 19.D. Geronimo, A. Sappa, A. L'opez, and D. Ponsa, ‘‘Adaptive image sampling and windows classification for on-board ped. det.’’ in Intl. Conf. Computer Vision Systems, 2005.
  20. 20.M. Andriluka, S. Roth, and B. Schiele, ‘‘People-tracking-bydetection and people-detection-by-tracking,’’ in IEEE Conf. Computer Vision and Pattern Recognition, 2008.
  21. 21.S. Munder and D. M. Gavrila, ‘‘An experimental study on pedestrian classification,’’ IEEE Trans. Pattern Analysis and Machine Intelligence, vol. 28, no. 11, pp. 1863–1868, 2006.
  22. 22.G. Overett, L. Petersson, N. Brewer, L. Andersson, and N. Pettersson, ‘‘A new pedestrian dataset for supervised learning,’’ in IEEE Intelligent Vehicles Symposium, 2008.
  23. 23.B. Russell, A. Torralba, K. P. Murphy, and W. T. Freeman, ‘‘LabelMe: A database and web-based tool for image ann.’’ Intl. Journal of Computer Vision, vol. 77, no. 1-3, pp. 157–173, 2008.
  24. 24.A. T. Nghiem, F. Bremond, M. Thonnat, and V. Valentin, ‘‘ETISEO, performance evaluation for video surveillance systems,’’ in IEEE Intl. Conf. on Advanced Video and Signal Based Surveillance, 2007.
  25. 25.E. Seemann, M. Fritz, and B. Schiele, ‘‘Towards robust pedestrian detection in crowded image sequences,’’ in IEEE Conf. Computer Vision and Pattern Recognition, 2007.
  26. 26.G. Salton and M. J. McGill, Introduction to Modern Information Retrieval. New York, NY, USA: McGraw-Hill, Inc., 1986.
  27. 27.M. Hussein, F. Porikli, and L. Davis, ‘‘A comp. eval. framework and a comparative study for human det.’’ IEEE Trans. Intelligent Transportation Systems, vol. 10, pp. 417–427, Sept. 2009.
  28. 28.S. Walk, N. Majer, K. Schindler, and B. Schiele, ‘‘New features and insights for pedestrian detection,’’ in IEEE Conf. Computer Vision and Pattern Recognition, 2010.
  29. 29.P. Doll'ar, Z. Tu, P. Perona, and S. Belongie, ‘‘Integral channel features,’’ in British Machine Vision Conf., 2009.
  30. 30.D. M. Gavrila and S. Munder, ‘‘Multi-cue pedestrian detection and tracking from a moving vehicle,’’ Intl. Journal of Computer Vision, vol. 73, pp. 41–59, 2007.
  31. 31.B. Leibe, A. Leonardis, and B. Schiele, ‘‘Robust object detection with interleaved categorization and segmentation,’’ Intl. Journal of Computer Vision, vol. 77, no. 1-3, pp. 259–289, May 2008.
  32. 32.C. H. Lampert, M. B. Blaschko, and T. Hofmann, ‘‘Beyond sliding windows: Object localization by eff. subwindow search,’’ in IEEE Conf. Computer Vision and Pattern Recognition, 2008.
  33. 33.P. Sabzmeydani and G. Mori, ‘‘Detecting pedestrians by learning shapelet features,’’ in IEEE Conf. Computer Vision and Pattern Recognition, 2007.
  34. 34.S. Maji, A. Berg, and J. Malik, ‘‘Classification using intersection kernel SVMs is efficient,’’ in IEEE Conf. Computer Vision and Pattern Recognition, 2008.
  35. 35.C. Gu, J. J. Lim, P. Arbelaez, and J. Malik, ‘‘Recog. using regions,’’ in IEEE Conf. Computer Vision and Pattern Recognition, 2009.
  36. 36.B. Leibe, E. Seemann, and B. Schiele, ‘‘Pedestrian detection in crowded scenes,’’ in IEEE Conf. Computer Vision and Pattern Recognition, 2005.
  37. 37.E. Seemann, B. Leibe, K. Mikolajczyk, and B. Schiele, ‘‘An evaluation of local shape-based features for pedestrian detection,’’ in British Machine Vision Conf., 2005.
  38. 38.I. Alonso, D. Llorca, M. Sotelo, L. Bergasa, P. R. de Toro, J. Nuevo, M. Ocana, and M. Garrido, ‘‘Combination of feature extraction methods for SVM pedestrian detection,’’ IEEE Trans. Intelligent Transportation Systems, vol. 8, no. 2, pp. 292–307, June 2007.
  39. 39.M. Bajracharya, B. Moghaddam, A. Howard, S. Brennan, and L. H. Matthies, ‘‘A fast stereo-based system for detecting and tracking pedestrians from a moving vehicle,’’ The Intl. Journal of Robotics Research, vol. 28, 2009.
  40. 40.A. Ess, B. Leibe, K. Schindler, and L. Van Gool, ‘‘Robust multiperson tracking from a mobile platform,’’ IEEE Trans. Pattern Analysis and Machine Intelligence, vol. 31, pp. 1831–1846, 2009.
  41. 41.C. Wojek, S. Roth, K. Schindler, and B. Schiele, ‘‘Monocular 3d scene modeling and inference: Understanding multi-object traffic scenes,’’ in European Conf. Computer Vision, 2010.
  42. 42.E. Dickmanns, Dynamic Vision for Perception and Control of Motion. Springer, 2007.
  43. 43.T. Gandhi and M. Trivedi, ‘‘Pedestrian protection systems: Issues, survey, and challenges,’’ IEEE Trans. Intelligent Transportation Systems, vol. 8, no. 3, pp. 413–430, Sept. 2007.
  44. 44.P. A. Viola and M. J. Jones, ‘‘Robust real-time face det.’’ Intl. Journal of Computer Vision, vol. 57, no. 2, pp. 137–154, 2004.
  45. 45.D. G. Lowe, ‘‘Distinctive image features from scale-invariant keypoints,’’ Intl. Journal of Computer Vision, vol. 60, no. 2, pp. 91–110, 2004.
  46. 46.Q. Zhu, S. Avidan, M. Yeh, and K. Cheng, ‘‘Fast human detection using a cascade of histograms of oriented gradients,’’ in IEEE Conf. Computer Vision and Pattern Recognition, 2006.
  47. 47.F. M. Porikli, ‘‘Integral histogram: A fast way to extract histograms in cartesian spaces,’’ in IEEE Conf. Computer Vision and Pattern Recognition, 2005.
  48. 48.A. Shashua, Y. Gdalyahu, and G. Hayun, ‘‘Ped. det. for driving assistance systems: Single-frame classification and system level performance,’’ in IEEE Intl. Conf. Intelligent Vehicles, 2004.
  49. 49.D. M. Gavrila and V. Philomin, ‘‘Real-time object det. for smart vehicles,’’ in IEEE Intl. Conf. Computer Vision, 1999, pp. 87–93.
  50. 50.D. M. Gavrila, ‘‘A bayesian, exemplar-based approach to hierarchical shape matching,’’ IEEE Trans. Pattern Analysis and Machine Intelligence, 2007.
  51. 51.Y. Liu, S. Shan, W. Zhang, X. Chen, and W. Gao, ‘‘Granularitytunable gradients partition descriptors for human det.’’ in IEEE Conf. Computer Vision and Pattern Recognition, 2009.
  52. 52.Y. Liu, S. Shan, X. Chen, J. Heikkila, W. Gao, and M. Pietikainen, ‘‘Spatial-temporal granularity-tunable gradients partition descriptors for human det.’’ in European Conf. Computer Vision, 2010.
  53. 53.P. A. Viola, M. J. Jones, and D. Snow, ‘‘Detecting pedestrians using patterns of motion and appearance,’’ Intl. Journal of Computer Vision, vol. 63(2), pp. 153–161, 2005.
  54. 54.N. Dalal, B. Triggs, and C. Schmid, ‘‘Human detection using oriented histograms of flow and appearance,’’ in European Conf. Computer Vision, 2006.
  55. 55.N. Dalal, ‘‘Finding people in images and videos,’’ Ph.D. dissertation, Institut Nat. Polytechnique de Grenoble, July 2006.
  56. 56.C. Wojek and B. Schiele, ‘‘A performance evaluation of single and multi-feature people detection,’’ in DAGM Symposium Pattern Recognition, 2008.
  57. 57.G. Mori, S. Belongie, and J. Malik, ‘‘Efficient shape matching using shape contexts,’’ IEEE Trans. Pattern Analysis and Machine Intelligence, pp. 1832–1837, 2005.
  58. 58.B. Wu and R. Nevatia, ‘‘Optimizing discrimination-efficiency tradeoff in integrating heterogeneous local features for object det.’’ in IEEE Conf. Computer Vision and Pattern Recognition, 2008.
  59. 59.X. Wang, T. X. Han, and S. Yan, ‘‘An hog-lbp human detector with partial occlusion handling,’’ in IEEE Intl. Conf. Computer Vision, 2009.
  60. 60.T. Ojala, M. Pietikainen, and T. Maenpaa, ‘‘Multiresolution grayscale and rotation invariant texture classification with local binary patterns,’’ IEEE Trans. Pattern Analysis and Machine Intelligence, vol. 24, no. 7, pp. 971–987, Jul. 2002.
  61. 61.S. Hussain and B. Triggs, ‘‘Feature sets and dimensionality reduction for visual object det.’’ in British Machine Vision Conf., 2010.
  62. 62.P. Ott and M. Everingham, ‘‘Implicit color segmentation features for pedestrian and object detection,’’ in IEEE Intl. Conf. Computer Vision, 2009.
  63. 63.P. Doll'ar, S. Belongie, and P. Perona, ‘‘The fastest pedestrian detector in the west,’’ in British Machine Vision Conf., 2010.
  64. 64.O. Tuzel, F. Porikli, and P. Meer, ‘‘Ped. det. via classification on riemannian manifolds,’’ IEEE Trans. Pattern Analysis and Machine Intelligence, vol. 30, no. 10, pp. 1713–1727, Oct. 2008.
  65. 65.B. Babenko, P. Doll'ar, Z. Tu, and S. Belongie, ‘‘Simultaneous learning and alignment: Multi-instance and multi-pose learning,’’ in ECCV Faces in Real-Life Images, 2008.
  66. 66.S. Walk, K. Schindler, and B. Schiele, ‘‘Disparity statistics for pedestrian detection: Combining appearance, motion and stereo,’’ in European Conf. Computer Vision, 2010.
  67. 67.P. Doll'ar, Z. Tu, H. Tao, and S. Belongie, ‘‘Feature mining for image classification,’’ in IEEE Conf. Computer Vision and Pattern Recognition, 2007.
  68. 68.A. Bar-Hillel, D. Levi, E. Krupka, and C. Goldberg, ‘‘Part-based feature synthesis for human detection,’’ in European Conf. Computer Vision, 2010.
  69. 69.W. Schwartz, A. Kembhavi, D. Harwood, and L. Davis, ‘‘Human detection using partial least squares analysis,’’ in IEEE Intl. Conf. Computer Vision, 2009.
  70. 70.Z. Lin and L. S. Davis, ‘‘A pose-invariant descriptor for human det. and seg.’’ in European Conf. Computer Vision, 2008.
  71. 71.P. Felzenszwalb, D. McAllester, and D. Ramanan, ‘‘A discriminatively trained, multiscale, deformable part model,’’ in IEEE Conf. Computer Vision and Pattern Recognition, 2008.
  72. 72.P. F. Felzenszwalb, R. B. Girshick, D. McAllester, and D. Ramanan, ‘‘Object detection with discriminatively trained part based models,’’ IEEE Trans. Pattern Analysis and Machine Intelligence, vol. 99, no. PrePrints, 2009.
  73. 73.A. Mohan, C. Papageorgiou, and T. Poggio, ‘‘Example-based object det. in images by components,’’ IEEE Trans. Pattern Analysis and Machine Intelligence, vol. 23, no. 4, pp. 349–361, Apr. 2001.
  74. 74.K. Mikolajczyk, C. Schmid, and A. Zisserman, ‘‘Human detection based on a probabilistic assembly of robust part detectors,’’ in European Conf. Computer Vision, 2004.
  75. 75.M. Enzweiler, A. Eigenstetter, B. Schiele, and D. M. Gavrila, ‘‘Multi-cue ped. classification with partial occlusion handling,’’ in IEEE Conf. Computer Vision and Pattern Recognition, 2010.
  76. 76.L. Bourdev and J. Malik, ‘‘Poselets: Body part detectors trained using 3d human pose annotations,’’ in IEEE Intl. Conf. Computer Vision, 2009.
  77. 77.M. Enzweiler and D. M. Gavrila, ‘‘Integrated pedestrian classification and orientation estimation,’’ in IEEE Conf. Computer Vision and Pattern Recognition, 2010.
  78. 78.D. Tran and D. Forsyth, ‘‘Configuration estimates improve pedestrian finding,’’ in Advances in Neural Information Processing Systems, 2008.
  79. 79.M. Weber, M. Welling, and P. Perona, ‘‘Unsupervised learning of models for recog.’’ in European Conf. Computer Vision, 2000.
  80. 80.R. Fergus, P. Perona, and A. Zisserman, ‘‘Object class recognition by unsupervised scale-invariant learning,’’ in IEEE Conf. Computer Vision and Pattern Recognition, 2003.
  81. 81.S. Agarwal and D. Roth, ‘‘Learning a sparse representation for object det.’’ in European Conf. Computer Vision, 2002.
  82. 82.P. Doll'ar, B. Babenko, S. Belongie, P. Perona, and Z. Tu, ‘‘Multiple component learning for object detection,’’ in European Conf. Computer Vision, 2008.
  83. 83.Z. Lin, G. Hua, and L. S. Davis, ‘‘Multiple instance feature for robust part-based object detection,’’ in IEEE Conf. Computer Vision and Pattern Recognition, 2009.
  84. 84.D. Park, D. Ramanan, and C. Fowlkes, ‘‘Multiresolution models for obj. det.’’ in European Conf. Computer Vision, 2010.
  85. 85.R. M. Haralick, K. Shanmugam, and I. Dinstein, ‘‘Textural features for image classification,’’ IEEE Trans. on Systems, Man, and Cybernetics, vol. 3, no. 6, pp. 610–621, 1973.
  86. 86.E. Shechtman and M. Irani, ‘‘Matching local self-similarities across images and videos,’’ in IEEE Conf. Computer Vision and Pattern Recognition, 2007.
  87. 87.J. Demsar, ‘‘Statistical Comparisons of Classifiers over Multiple Data Sets,’’ Journal of Machine Learning Research, vol. 7, pp. 1–30, 2006.
  88. 88.S. Garc'ıa and F. Herrera, ‘‘An extension on ’’Statistical Comparisons of Classifiers over Multiple Data Sets’’ for all pairwise comparisons,’’ Journal of Machine Learning Research, vol. 9, pp. 2677–2694, 2008.
  89. 89.W. Zhang, G. J. Zelinsky, and D. Samaras, ‘‘Real-time accurate object detection using multiple resolutions,’’ in IEEE Intl. Conf. Computer Vision, 2007.
  90. 90.C. Wojek, G. Dork'o, A. Schulz, and B. Schiele, ‘‘Sliding-windows for rapid object class localization: A parallel technique,’’ in DAGM Symposium Pattern Recognition, 2008.

Citation

MLA
Dollar, P., et al. “Pedestrian Detection: An Evaluation of the State of the Art”. IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 34, no. 4, 2012, pp. 743–61, https://doi.org/10.1109/TPAMI.2011.155.
APA
Dollar, P., Wojek, C., Schiele, B., & Perona, P. (2012). Pedestrian Detection: An Evaluation of the State of the Art. IEEE Transactions on Pattern Analysis and Machine Intelligence, 34(4), 743–761. https://doi.org/10.1109/TPAMI.2011.155
Chicago
Dollar, P., C. Wojek, B. Schiele, and P. Perona. 2012. “Pedestrian Detection: An Evaluation of the State of the Art”. IEEE Transactions on Pattern Analysis and Machine Intelligence 34 (4): 743–61. https://doi.org/10.1109/TPAMI.2011.155.
Harvard
Dollar, P. et al. (2012) “Pedestrian Detection: An Evaluation of the State of the Art”, IEEE Transactions on Pattern Analysis and Machine Intelligence, 34(4), pp. 743–761. Available at: https://doi.org/10.1109/TPAMI.2011.155.
Vancouver
1. Dollar P, Wojek C, Schiele B, Perona P (2012) Pedestrian Detection: An Evaluation of the State of the Art. IEEE Transactions on Pattern Analysis and Machine Intelligence 34:743–761

BibTeX

@article{Dollar_2012, title={Pedestrian Detection: An Evaluation of the State of the Art}, volume={34}, ISSN={2160-9292}, url={http://dx.doi.org/10.1109/TPAMI.2011.155}, DOI={10.1109/tpami.2011.155}, number={4}, journal={IEEE Transactions on Pattern Analysis and Machine Intelligence}, publisher={Institute of Electrical and Electronics Engineers (IEEE)}, author={Dollar, P. and Wojek, C. and Schiele, B. and Perona, P.}, year={2012}, month=Apr, pages={743–761} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF