A Trainable System for Object Detection

CONSTANTINE PAPAGEORGIOUTOMASO POGGIO

article2000IJCV1,483 citations

Presents a general framework for object detection in cluttered scenes by combining an overcomplete dictionary of multiscale Haar wavelets with support vector machines to achieve real-time detection across multiple visual domains.

Listen

As image and video data expand across digital databases, automotive safety systems, and surveillance operations, automated visual object detection has become a vital operational capability. Detecting distinct object classes in cluttered, unconstrained real-world scenes remains difficult due to variations in lighting, color, texture, pose, and background clutter. Prior methods often rely on restrictive operational assumptions, such as static backgrounds, moving targets, manual body-part models, or tracking over time. The article demonstrates a trainable object detection framework that learns implicitly from examples without domain-specific handcrafting, evaluating its effectiveness across static image benchmarks for faces, people, and vehicles.

The framework combines a dense Haar wavelet representation—which measures local, multiscale intensity differences across adjacent regions—with a Support Vector Machine classifier. Support Vector Machines are statistical learning models designed to classify complex data accurately while controlling classifier complexity to prevent overfitting. The system evaluates static grayscale and color image sets, applying a sliding examination window across multiple resized versions of an image to detect targets at multiple scales. Training datasets ranged from over 1,000 aligned positive samples for cars and 1,800 for people to over 2,400 for faces, paired with thousands of negative non-target patterns.

The findings show that the overcomplete Haar wavelet representation consistently outperforms traditional pixel and principal component analysis methods across all tested domains. In face detection benchmarks, the system achieved a 90% detection rate with a false positive rate of roughly 1 in 100,000 patterns, which translates to about one false detection per image. For people detection, color wavelet features achieved a 90% detection rate at 1 false positive per 10,000 patterns, significantly outperforming raw pixel and complete wavelet baselines. Unsigned wavelet magnitudes yielded higher accuracy than signed gradients across both faces and people, demonstrating that modeling boundary strength while discarding gradient polarity produces a cleaner, more generalizable class model.

These results demonstrate that a single visual representation can detect diverse object categories without needing custom-engineered architectures for each domain. For automotive safety, where full unoptimized scans took up to 20 minutes per frame, the authors developed a streamlined grayscale implementation using 29 salient wavelet features and a reduced decision boundary. Integrated with DaimlerChrysler's stereo obstacle tracker in a demonstration vehicle, this deployment narrowed the search space, enabling pedestrian detection to operate at over 10 frames per second with under 15 milliseconds of processing time per identified obstacle.

Organizations deploying automated detection should adopt modular frameworks paired with front-end focus-of-attention filters to balance speed and accuracy in high-throughput environments. Future work should expand training sets to eliminate false alarms caused by unrepresented poses, integrate temporal motion data to suppress remaining false positives in video feeds, and explore component-based detection models for complex, multi-part objects like vehicles.

Confidence in the system's core capabilities is high, backed by extensive testing across out-of-sample static images. However, operational limitations exist, including susceptibility to missed detections when objects are heavily rotated, partially occluded, or clipped at image borders. Decision-makers should account for these boundary conditions when applying the system in unconstrained visual environments.

  • Paper: Example-Based Learning for View-Based Human Face Detection, Kah Kay Sung et al. (1998). Sung and Poggio established the foundational example-based sliding-window framework for face detection using statistical models and negative pattern curation that directly precedes and motivates this trainable wavelet-based detection architecture.
  • Paper: Probabilistic Visual Learning for Object Representation, B. Moghaddam et al. (1997). Moghaddam and Pentland introduced appearance-based visual learning and eigenspace density estimation for object representation that this paper builds upon and contrasts against dense wavelet representations.
  • Paper: Histograms of Oriented Gradients for Human Detection, Navneet Dalal et al. (2005). Dalal and Triggs advance the dense gradient and SVM paradigm introduced here by establishing Histograms of Oriented Gradients (HOG) as a substantially more robust representation for pedestrian detection.
  • Paper: Detecting Faces in Images: A Survey, Ming-Hsuan Yang et al. (2002). This comprehensive survey provides an exhaustive comparative overview of appearance-based and SVM detection methods, situating this paper's wavelet architecture within the broader evolution of face detection.
  • Paper: Human Detection Using Oriented Histograms of Flow and Appearance, Navneet Dalal et al. (2006). This work directly addresses this paper's proposed future direction by integrating optical flow and temporal motion descriptors with static gradient features to drastically cut false positives in pedestrian video streams.
  • Paper: Object Detection with Discriminatively Trained Part-Based Models, Pedro F. Felzenszwalb et al. (2010). Felzenszwalb et al. directly realize this paper's proposed component-based future direction by formulating deformable part-based models trained discriminatively with latent SVMs.
  • Paper: Sharing visual features for multiclass and multiview object detection, Antonio Torralba et al. (2007). Torralba et al. expand upon generic multi-class object detection by developing shared visual features and boosted classifiers to scale learning across dozens of categories efficiently.
  • Paper: Pedestrian Detection: An Evaluation of the State of the Art, Piotr Dollár et al. (2012). Dollár et al. systematically benchmark modern pedestrian detection algorithms across diverse datasets, tracing the subsequent performance lineage originating from early trainable systems like this one.
  • Paper: Fast Feature Pyramids for Object Detection, Piotr Dollar et al. (2014). Dollar et al. solve the multi-scale sliding-window computational bottleneck inherent to dense feature architectures by demonstrating fast feature pyramid extrapolation across scales.
  • Paper: Object Detection in 20 Years: A Survey, Zhengxia Zou et al. (2019). This survey provides a definitive retrospective charting the historical progression of object detection from early hand-crafted feature and SVM sliding-window systems to contemporary deep networks.
Cover for A Trainable System for Object Detection

Abstract

This paper presents a general, trainable system for object detection in unconstrained, cluttered scenes. The system derives much of its power from a representation that describes an object class in terms of an overcomplete dictionary of local, oriented, multiscale intensity differences between adjacent regions, efficiently computable as a Haar wavelet transform. This example-based learning approach implicitly derives a model of an object class by training a support vector machine classifier using a large set of positive and negative examples. We present results on face, people, and car detection tasks using the same architecture. In addition, we quantify how the representation affects detection performance by considering several alternate representations including pixels and principal components. We also describe a real-time application of our person detection system as part of a driver assistance system.

Table of Contents

  • 1. Introduction
  • 2. Architecture and Representation
  • 2.1. Wavelets
  • 2.2. The Wavelet Representation
  • 2.3. Support Vector Machine Classification
  • 3. Experiments
  • 3.1. Pixels, Wavelets, PCA
  • 3.2. Signed vs. Unsigned Wavelets
  • 3.3. Complete vs. Overcomplete
  • 3.4. Color vs. Gray Level
  • 3.5. Faces, People, and Cars
  • 4. A Real-Time Application
  • 4.1. Speed Optimizations
  • 4.2. Integration with the DaimlerChrysler Urban Traffic Assistant
  • 5. Conclusion
  • Acknowledgments
  • References

Knowls

  1. Knowl 1 — Trainable Multiscale Object Detection Framework

    model/method

    The object detection framework detects visual object classes (such as human faces, pedestrians, and automobiles) in static, unconstrained, cluttered scenes without hand-crafted geometric models or background assumptions.

    The framework operates in two phases:

    1. Training Phase: Labeled positive image patches (aligned and scaled instances of the target object class) and negative image patches (cluttered background patterns) are transformed into feature vectors using an overcomplete dictionary of 2D Haar wavelets. A support vector machine (SVM) classifier is trained on these feature vectors to determine an optimal separating decision surface that minimizes structural risk.

    2. Testing Phase: To perform multiscale detection in an arbitrary test image, the system constructs a resolution pyramid by iteratively resizing the image. A fixed-size detection window is raster-scanned across each scale level. For every window position, local overcomplete Haar wavelet features are extracted directly in the wavelet domain and classified by the trained SVM as containing the target object or background.

  2. Knowl 2 — Quadruple Density Haar Wavelet Transform Algorithm

    algorithm

    To obtain dense spatial sampling for object pattern modeling while maintaining an O(n)O(n) time complexity in the number of pixels nn, the quadruple-density discrete Haar wavelet transform eliminates standard decimation steps and interleaves transformed sub-bands.

    For a 1D signal with scaling filter h={…,0,12,12,0,0,… }h = \{\dots, 0, \frac{1}{2}, \frac{1}{2}, 0, 0, \dots\} and wavelet filter g={…,0,−12,12,0,0,… }g = \{\dots, 0, -\frac{1}{2}, \frac{1}{2}, 0, 0, \dots\}, the transform proceeds across levels n=1,…,Ln = 1, \dots, L as follows:

    Input: 1D signal ff, decomposition levels LL
    Output: Quadruple-density wavelet coefficients {γn}\{\gamma_n\} for n=1,…,Ln = 1, \dots, L
    λ0←f\lambda_0 \leftarrow f
    for n=1n = 1 to LL:
        λdense←convolve(λn−1,h) without downsampling\lambda_{\text{dense}} \leftarrow \text{convolve}(\lambda_{n-1}, h) \text{ without downsampling}
        λeven←elements of λdense at even indices\lambda_{\text{even}} \leftarrow \text{elements of } \lambda_{\text{dense}} \text{ at even indices}
        λodd←elements of λdense at odd indices\lambda_{\text{odd}} \leftarrow \text{elements of } \lambda_{\text{dense}} \text{ at odd indices}
        γeven←convolve(λeven,g) without downsampling\gamma_{\text{even}} \leftarrow \text{convolve}(\lambda_{\text{even}}, g) \text{ without downsampling}
        γodd←convolve(λodd,g) without downsampling\gamma_{\text{odd}} \leftarrow \text{convolve}(\lambda_{\text{odd}}, g) \text{ without downsampling}
        γn←interleave(γeven,γodd)\gamma_n \leftarrow \text{interleave}(\gamma_{\text{even}}, \gamma_{\text{odd}})
        λn←λeven\lambda_n \leftarrow \lambda_{\text{even}}
    return {γ1,γ2,…,γL}\{\gamma_1, \gamma_2, \dots, \gamma_L\}

    The 2D quadruple-density transform is computed by applying the 1D transform sequentially along rows and columns at each scale, producing square support wavelets spaced at intervals of 142n\frac{1}{4}2^n pixels at scale level nn.

  3. Knowl 3 — Overcomplete Haar Wavelet Feature Extraction and Normalization

    model/method

    Image patches are mapped to an overcomplete feature space designed to capture local, oriented, multiscale intensity differences while ensuring robustness to contrast, illumination, and color variations:

    • 2D Non-Standard Basis: At each scale, three oriented wavelet functions are defined from the 1D scaling function ϕ\phi and mother wavelet ψ\psi:

      • Vertical: ψ(x,y)=ψ(x)⊗ϕ(y)\psi(x, y) = \psi(x) \otimes \phi(y) (responds to vertical edges)
      • Horizontal: ψ(x,y)=ϕ(x)⊗ψ(y)\psi(x, y) = \phi(x) \otimes \psi(y) (responds to horizontal edges)
      • Diagonal: ψ(x,y)=ψ(x)⊗ψ(y)\psi(x, y) = \psi(x) \otimes \psi(y) (responds to corner and diagonal features)
    • Unsigned Magnitude Representation: Feature values are taken as the absolute magnitude of the wavelet response, ∣γ∣|\gamma|, discarding the sign of the intensity gradient. This provides invariance to contrast reversals (e.g., light silhouettes on dark backgrounds vs. dark silhouettes on light backgrounds).

    • Multichannel Color Pooling: For RGB images, the 2D wavelet transform is computed independently per color channel, and each feature takes the maximum absolute value across channels: γcolor=max⁡(∣γR∣,∣γG∣,∣γB∣)\gamma_{\text{color}} = \max\left(|\gamma_R|, |\gamma_G|, |\gamma_B|\right)

    • Local Area Normalization: To compensate for global lighting differences, each coefficient γc(x,y)\gamma_c(x, y) of class c∈{vertical,horizontal,diagonal}×{scale}c \in \{\text{vertical}, \text{horizontal}, \text{diagonal}\} \times \{\text{scale}\} is divided by the mean absolute coefficient value within that class across the entire detection window: γ~c(x,y)=∣γc(x,y)∣1Nc∑(u,v)∣γc(u,v)∣\tilde{\gamma}_c(x, y) = \frac{|\gamma_c(x, y)|}{\frac{1}{N_c} \sum_{(u, v)} |\gamma_c(u, v)|}

    • Intermediate Scale Selection: The finest wavelet scales (dominated by high-frequency texture noise) and coarsest scales (exceeding object sub-structures) are discarded, retaining only two intermediate scales.

  4. Knowl 4 — Quadratic Support Vector Machine Decision Rule

    equation

    The classification decision for a candidate feature vector x∈Rdx \in \mathbb{R}^d extracted from an image window is evaluated using an SVM with a second-order polynomial kernel:

    f(x)=θ(∑i=1Nsαiyi(x⋅xi+1)2+b)f(x) = \theta\left( \sum_{i=1}^{N_s} \alpha_i y_i (x \cdot x_i + 1)^2 + b \right)

    where:

    • Ns∈NN_s \in \mathbb{N} is the number of support vectors selected during optimization,
    • xi∈Rdx_i \in \mathbb{R}^d is the ii-th support vector,
    • yi∈{−1,+1}y_i \in \{-1, +1\} is the binary training label corresponding to xix_i,
    • αi>0\alpha_i > 0 is the Lagrange multiplier for the ii-th support vector,
    • b∈Rb \in \mathbb{R} is the bias threshold,
    • (x⋅xi+1)2(x \cdot x_i + 1)^2 is the quadratic polynomial kernel K(x,xi)K(x, x_i),
    • θ(z)\theta(z) is the step threshold function defined as θ(z)=1\theta(z) = 1 if z≥0z \ge 0 and −1-1 if z<0z < 0.
  5. Knowl 5 — Wavelet Feature Configurations for Faces, Pedestrians, and Cars

    model/method

    The object detection architecture is instantiated for three distinct visual categories using specific scale and window configurations:

    • Face Detection:

      • Window size: 19×1919 \times 19 pixels, grayscale.
      • Wavelet scales: 2×22 \times 2 (double density, 17×1717 \times 17 positions per orientation) and 4×44 \times 4 (quadruple density, 17×1717 \times 17 positions per orientation).
      • Total feature dimension: 3×(289+289)=1,7343 \times (289 + 289) = 1,734 features.
    • Pedestrian Detection:

      • Window size: 128×64128 \times 64 pixels, color (RGB).
      • Target alignment: Centered body with shoulder-to-foot height ≈80\approx 80 pixels.
      • Wavelet scales: 16×1616 \times 16 (29×13=37729 \times 13 = 377 positions per orientation) and 32×3232 \times 32 (15×5=7515 \times 5 = 75 positions per orientation).
      • Total feature dimension: 3×(377+75)=1,3263 \times (377 + 75) = 1,326 features (after max-RGB pooling).
    • Car Detection:

      • Window size: 128×128128 \times 128 pixels, color (RGB frontal and rear views).
      • Target alignment: Front or rear bumper width normalized to 6464 pixels.
      • Wavelet scales: 16×1616 \times 16 and 32×3232 \times 32.
      • Total feature dimension: 3,0303,030 features (after max-RGB pooling).
  6. Knowl 6 — Empirical Superiority of Overcomplete Unsigned Wavelet Representations

    empirical result

    ROC curve evaluations across face and pedestrian detection benchmarks demonstrate key properties of feature representations:

    • Overcomplete Wavelets vs. Pixels and PCA: In face detection (tested on 105 images with 3,909,200 non-face patterns), gray unsigned wavelets outperformed histogram-equalized pixels (361 features), raw pixels (361 features), and PCA projections (361 features), achieving a 90%90\% detection rate at 1 false positive per 100,000100,000 inspected windows (≈1\approx 1 false positive per image). In pedestrian detection, overlapping 8×88 \times 8 pixel averages and their PCA projections performed substantially worse than wavelets.

    • Overcomplete vs. Complete Dictionary: In pedestrian detection, the overcomplete quadruple-density representation (1,3261,326 features) drastically outperformed the complete standard Haar basis (120120 features) across all false positive rates.

    • Unsigned vs. Signed Wavelets: For both face and pedestrian classes, unsigned (absolute magnitude) wavelets consistently outperformed signed wavelets, because unsigned features reduce intra-class variance arising from lighting direction and background contrast polarity.

    • Color vs. Grayscale: Color unsigned wavelets outperformed grayscale unsigned wavelets on pedestrian detection across all operating points.

  7. Knowl 7 — Computational Speed Optimizations for Pedestrian Classification

    model/method

    To transition the pedestrian detection engine from an unoptimized research speed of 1 frame per 20 minutes to real-time performance, four optimization stages are applied:

    1. Grayscale Processing: Transforming images to single-channel grayscale eliminates per-channel color transform computations.

    2. Salient Feature Selection: The feature vector is pruned from 1,326 coefficients down to 29 manually selected salient wavelets encoding the outer silhouette and head/extremity corners:

      • Scale 32×3232 \times 32: 6 vertical coefficients, 1 horizontal coefficient.
      • Scale 16×1616 \times 16: 14 vertical coefficients, 8 horizontal coefficients.
    3. Reduced Set Vectors: The reduced set method approximates the trained SVM decision surface (which originally contains ≈1,000\approx 1,000 support vectors) using exactly 29 synthetic support vectors, transforming the quadratic polynomial evaluation into a 29-dimensional dot product over 29 vectors.

    4. Focus of Attention via Stereo Vision: Candidate evaluation is restricted to regions of interest and bounded depth ranges provided by a front-end stereo vision module, reducing the search space to at most three scale levels.

  8. Knowl 8 — Real-Time Embedded Pedestrian Detection in the Urban Traffic Assistant

    empirical result

    The optimized pedestrian detection module (using 29 grayscale unsigned wavelet features and 29 reduced set synthetic support vectors) was integrated into the DaimlerChrysler Urban Traffic Assistant (UTA) system and deployed in an experimental demonstration vehicle.

    UTA provided real-time 3D obstacle bounding boxes and depth estimates from a binocular stereo vision system running at 25 Hz on a 200 MHz PowerPC 604 processor. When evaluated on real-world driving sequences through urban traffic in Esslingen/Stuttgart, Germany, the integrated pedestrian detection module achieved an operational processing rate exceeding 10 Hz, spending less than 15 ms of classification time per detected obstacle.

  9. Knowl 9 — Pose Sensitivity and Viewpoint Limitations of Holistic Full-Pattern Detectors

    limitation

    The holistic full-pattern wavelet framework relies on a consistent 2D silhouette and internal structure across the training ensemble. While this assumption holds effectively for upright pedestrian silhouettes and frontal faces across moderate variations, it limits performance on object categories with severe 3D appearance variations under changing viewpoints (such as multi-view car detection).

    For classes with large out-of-plane rotations and viewpoint changes, full-pattern template matching is insufficient, requiring component-based architectures that independently detect parts (e.g., wheels, headlights, windshields) and evaluate their spatial geometric configuration.

Coverage note — None: all contributed methods, algorithmic details, experimental comparisons, speed optimizations, real-time vehicle deployment results, and architectural limitations have been covered.

References

  1. 1.Betke, M., Haritaoglu, E., and Davis, L. 1997. Highway scene analysis in hard real-time. In Proceedings of Intelligent Transportation Systems.
  2. 2.Betke, M. and Nguyen, H. 1998. Highway scene analysis form a moving vehicle under reduced visibility conditions. In Proceedings of Intelligent Vehicles, pp. 131–136.
  3. 3.Beymer, D., McLauchlan, P., Coifman, B., and Malik, J. 1997. A real-time computer vision system for measuring traffic parameters. In Proceedings of Computer Vision and Pattern Recognition, pp. 495–501.
  4. 4.Bregler, C. and Malik, J. 1996. Learning appearance based models: Mixtures of second moment experts. In Advances in Neural Information Processing Systems.
  5. 5.Burges, C. 1996. Simplified support vector decision rules. In Proceedings of 13th International Conference on Machine Learning.
  6. 6.Burges, C. 1998. A tutorial on support vector machines for pattern recognition. In Proceedings of Data Mining and Knowledge Discovery, U. Fayyad (Ed.), pp. 1–43.
  7. 7.Forsyth, D. and Fleck, M. 1997. Body plans. In Proceedings of Computer Vision and Pattern Recognition, pp. 678–683.
  8. 8.Forsyth, D. and Fleck, M. 1999. Automatic detection of human nudes, International Journal of Computer Vision, 32(1):63–77.
  9. 9.Franke, U., Gavrila, D., Goerzig, S., Lindner, F., Paetzold, F., and Woehler, C. 1998. Autonomous driving goes downtown. IEEE Intelligent Systems, pp. 32–40.
  10. 10.Haritaoglu, I., Harwood, D., and Davis, L. 1998. W4: Who? When? Where? What? A real time system for detecting and tracking people. In Face and Gesture Recognition, pp. 222–227.
  11. 11.Heisele, B. and Wohler, C. 1998. Motion-based recognition of pedestrians. In Proceedings of International Conference on Pattern Recognition, pp. 1325–1330.
  12. 12.Hogg, D. 1983. Model-based vision: A program to see a walking person. Image and Vision Computing, 1(1):5–20.
  13. 13.Itti, L. and Koch, C. 1999. A comparison of feature combination strategies for saliency-based visual attention systems. In Human Vision and Electronic Imaging, vol. 3644, pp. 473–482.
  14. 14.Itti, L., Koch, C., and Niebur, E. 1998. A model of saliency-based visual attention for rapid scene analysis. IEEE Transactions on Pattern Analysis and Machine Intelligence, 20(11):1254–1259.
  15. 15.Joachims, T. 1997. Text categorization with support vector machines. Technical Report LS-8 Report 23, University of Dortmund.
  16. 16.Lipson, P. 1996. Context and configuration based scene classification. Ph.D. thesis, Massachusetts Institute of Technology.
  17. 17.Lipson, P., Grimson, W., and Sinha, P. 1997. Configuration based scene classification and image indexing. In Proceedings of Computer Vision and Pattern Recognition, pp. 1007–1013.
  18. 18.Mallat, S. 1989. A theory for multiresolution signal decomposition: The wavelet representation. IEEE Transactions on Pattern Analysis and Machine Intelligence, 11(7):674–693.
  19. 19.McKenna, S. and Gong, S. 1997. Non-intrusive person authentication for access control by visual tracking and face recognition. In Audio- and Video-based Biometric Person Authentication, J. Bigun, G. Chollet, and G. Borgefors (Eds.), pp. 177–183.
  20. 20.Moghaddam, B. and Pentland, A. 1995. Probabilistic visual learning for object detection. In Proceedings of 6th International Conference on Computer Vision.
  21. 21.Mohan, A. 1999. Robust object detection in images by components. Master’s Thesis, Massachusetts Institute of Technology.
  22. 22.Osuna, E., Freund, R., and Girosi, F. 1997a. Support vector machines: Training and applications. A.I. Memo 1602, MIT Artificial Intelligence Laboratory.
  23. 23.Osuna, E., Freund, R., and Girosi, F. 1997b. Training support vector machines: An application to face detection. In Proceedings of Computer Vision and Pattern Recognition, pp. 130–136.
  24. 24.Rohr, K. 1993. Incremental recognition of pedestrians from image sequences. In Proceedings of Computer Vision and Pattern Recognition, pp. 8–13.
  25. 25.Rowley, H., Baluja, S., and Kanade, T. 1998. Neural network-based face detection. IEEE Transactions on Pattern Analysis and Machine Intelligence, 20(1):23–38.
  26. 26.Shio, A. and Sklansky, J. 1991. Segmentation of people in motion. In IEEE Workshop on Visual Motion, pp. 325–332.
  27. 27.Sinha, P. 1994. Qualitative image-based representations for object recognition. A.I. Memo 1505, MIT Artificial Intelligence Laboratory.
  28. 28.Stollnitz, E., DeRose, T., and Salesin, D. 1994. Wavelets for computer graphics: A primer. Technical Report 94-09-11, Department of Computer Science and Engineering, University of Washington.
  29. 29.Sung, K.-K. 1995. Learning and example selection for object and pattern detection. Ph.D. Thesis, MIT Artificial Intelligence Laboratory.
  30. 30.Sung, K.-K. and Poggio, T. 1994. Example-based learning for view-based human face detection. A.I. Memo 1521, MIT Artificial Intelligence Laboratory.
  31. 31.Vaillant, R., Monrocq, C., and Cun, Y.L. 1994. Original approach for the localisation of objects in images. IEE Proceedings Vision Image Signal Processing, 141(4):245–250.
  32. 32.Vapnik, V. 1995. The Nature of Statistical Learning Theory. Springer Verlag.
  33. 33.Vapnik, V. 1998. Statistical Learning Theory. John Wiley and Sons: New York.
  34. 34.Wren, C., Azarbayejani, A., Darrell, T., and Pentland, A. 1995. Pfinder: Real-time tracking of the human body. Technical Report 353, MIT Media Laboratory.

Citation

MLA
Papageorgiou, C., and T. Poggio. “A Trainable System for Object Detection”. International Journal of Computer Vision, vol. 38, no. 1, 2000, pp. 15–33, https://doi.org/10.1023/A:1008162616689.
APA
Papageorgiou, C., & Poggio, T. (2000). A Trainable System for Object Detection. International Journal of Computer Vision, 38(1), 15–33. https://doi.org/10.1023/A:1008162616689
Chicago
Papageorgiou, C., and T. Poggio. 2000. “A Trainable System for Object Detection”. International Journal of Computer Vision 38 (1): 15–33. https://doi.org/10.1023/A:1008162616689.
Harvard
Papageorgiou, C. and Poggio, T. (2000) “A Trainable System for Object Detection”, International Journal of Computer Vision, 38(1), pp. 15–33. Available at: https://doi.org/10.1023/A:1008162616689.
Vancouver
1. Papageorgiou C, Poggio T (2000) A Trainable System for Object Detection. International Journal of Computer Vision 38:15–33

BibTeX

@article{Papageorgiou_2000, title={A Trainable System for Object Detection}, volume={38}, ISSN={1573-1405}, url={http://dx.doi.org/10.1023/A:1008162616689}, DOI={10.1023/a:1008162616689}, number={1}, journal={International Journal of Computer Vision}, publisher={Springer Science and Business Media LLC}, author={Papageorgiou, Constantine and Poggio, Tomaso}, year={2000}, month=June, pages={15–33} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF