Ensemble Tracking

S. Avidan

article2005TPAMI1,590 citations

Proposes a visual tracking framework that treats tracking as an online binary classification problem by combining AdaBoost and mean-shift optimization to adaptively distinguish target objects from complex backgrounds across video frames.

Listen

Visual tracking is essential for applications such as automated surveillance, driver assistance systems, and human-computer interfaces. However, conventional tracking algorithms often fail when target objects encounter complex visual conditions, changing backgrounds, or appearance variations. Many traditional systems focus exclusively on modeling the tracked object rather than separating it from its immediate background, leading to tracking loss when colors or textures overlap.

The article demonstrates an online tracking framework called ensemble tracking, which frames the tracking task as an active binary classification problem to continually distinguish foreground objects from changing background environments. Rather than relying on a static object representation, the approach evaluates how combining multiple simple models in real time improves tracking robustness under dynamic real-world conditions.

The framework trains a collection of simple classifiers on individual video frames to label pixels as either object or background. An adaptive boosting procedure combines these simple models into a single weighted classifier that generates a spatial confidence map for subsequent frames. A mode-seeking search algorithm identifies the confidence peak to locate the object's new position. The tracker continuously adapts by discarding the oldest weak classifier, updating the weights of remaining models, and learning a new classifier from the most recent frame. Testing evaluated this framework across several real-world video sequences—including moving cameras, pedestrians crossing in front of similarly colored backgrounds, out-of-plane facial rotations, and a 225-frame grayscale vehicle sequence—using multi-scale image processing and pixel-level color and edge orientation features.

The evaluation produced four key operational findings. First, continuous online model updating is necessary to maintain target lock; a static baseline model lost track of its target by frame 30, whereas the adaptive tracker followed the target successfully through the entire sequence. Second, the system dynamically shifts feature emphasis as conditions change, automatically relying more heavily on edge orientations when the target and background share identical colors. Third, the framework operates effectively on high-dimensional feature spaces and low-information grayscale footage where conventional color-histogram trackers struggle. Fourth, incorporating an outlier rejection mechanism successfully cleans noisy training labels caused by rectangular bounding boxes, producing clearer confidence maps and preventing tracking drift.

These findings indicate that treating tracking as an ongoing classification problem delivers stable, high-performance tracking without requiring complex offline training or static camera setups. By breaking computation into sequential, lightweight learning tasks, the method achieves adaptability at low computational cost, significantly mitigating risks related to false detections, lighting shifts, and partial occlusions in automated vision systems.

Organizations developing computer vision systems should adopt this ensemble-based framework to enhance tracking reliability in environments with mobile cameras or visually confusing backgrounds. Further work should explore porting the prototype implementation from development environments to optimized production code to achieve higher processing frame rates, as well as integrating predefined domain-specific detectors as permanent classifiers within the ensemble.

Confidence in these findings is high for the evaluated video sequences and feature sets. However, readers should note that the current implementation requires manual target initialization in the first frame and currently operates at a few frames per second in a development environment, meaning additional optimization is necessary before deployment in hard real-time systems.

Cover for Ensemble Tracking

Abstract

We consider tracking as a binary classification problem, where an ensemble of weak classifiers is trained on-line to distinguish between the object and the background. The ensemble of weak classifiers is combined into a strong classifier using AdaBoost. The strong classifier is then used to label pixels in the next frame as either belonging to the object or the background, giving a confidence map. The peak of the map, and hence the new position of the object, is found using mean shift. Temporal coherence is maintained by updating the ensemble with new weak classifiers that are trained on-line during tracking. We show a realization of this method and demonstrate it on several video sequences.

Table of Contents

  • 1 Introduction
  • 2 Background
  • 3 Ensemble Tracking
  • 3.1 The weak classifier
  • 3.2 Ensemble update
  • 3.3 Outlier rejection
  • 3.4 Multi-resolution tracking
  • 4 Experiments
  • 5 Conclusions
  • References

Knowls

  1. Knowl 1 — Online Ensemble Tracking via AdaBoost and Mean Shift

    algorithm

    Ensemble Tracking treats visual tracking as an online binary classification problem that discriminates foreground object pixels (y=+1y = +1) from surrounding background pixels (y=−1y = -1). The tracker maintains an ensemble of TT weak linear classifiers combined into a strong classifier H(x)=∑t=1Tαtht(x)H(x) = \sum_{t=1}^T \alpha_t h_t(x). At each video frame, pixel classifications form a confidence map whose spatial peak is located using mean shift. Temporal adaptation is achieved by discarding the KK oldest weak classifiers, updating the weights αt\alpha_t of the remaining T−KT-K classifiers, and training KK new weak classifiers using the AdaBoost sample weight distribution.

    Input: Video frames I1,…,InI_1, \dots, I_n; initial bounding rectangle r1r_1
    Output: Object bounding rectangles r2,…,rnr_2, \dots, r_n
    Initialization (for frame I1I_1):
      Extract pixel feature vectors and labels {xi,yi}i=1N\{x_i, y_i\}_{i=1}^N from I1I_1 using r1r_1
      Initialize sample weights wi←1Nw_i \leftarrow \frac{1}{N} for all i∈{1,…,N}i \in \{1, \dots, N\}
      for t=1t = 1 to TT do
        Normalize sample weights: wi←wi∑j=1Nwjw_i \leftarrow \frac{w_i}{\sum_{j=1}^N w_j}
        Train weak classifier ht(x)∈{−1,+1}h_t(x) \in \{-1, +1\}
        Compute classification error errt=∑i=1Nwi⋅I(ht(xi)≠yi)\text{err}_t = \sum_{i=1}^N w_i \cdot \mathbb{I}(h_t(x_i) \ne y_i)
        Compute classifier weight αt=12ln⁡(1−errterrt)\alpha_t = \frac{1}{2} \ln\left(\frac{1 - \text{err}_t}{\text{err}_t}\right)
        Update sample weights: wi←wiexp⁡(αt⋅I(ht(xi)≠yi))w_i \leftarrow w_i \exp(\alpha_t \cdot \mathbb{I}(h_t(x_i) \ne y_i))
      end for
      Form strong classifier H(x)=∑t=1Tαtht(x)H(x) = \sum_{t=1}^T \alpha_t h_t(x)
    Tracking loop (for each frame IjI_j, j=2,…,nj = 2, \dots, n):
      Extract pixel feature vectors {xi}i=1N\{x_i\}_{i=1}^N over the search window in IjI_j
      Evaluate H(xi)H(x_i) to compute confidence map LjL_j
      Run mean shift on LjL_j initialized at rj−1r_{j-1} to find mode rectangle rjr_j
      Assign labels yi=+1y_i = +1 for pixels inside rjr_j and yi=−1y_i = -1 outside
      Remove the KK oldest weak classifiers from the ensemble
      Reset sample weights wi←1Nw_i \leftarrow \frac{1}{N}
      for l=K+1l = K+1 to TT do
        Normalize sample weights: wi←wi∑k=1Nwkw_i \leftarrow \frac{w_i}{\sum_{k=1}^N w_k}
        Select ht(x)h_t(x) from remaining classifiers that minimizes error err\text{err}
        Update αt\alpha_t and sample weights {wi}i=1N\{w_i\}_{i=1}^N
        Remove selected ht(x)h_t(x) from pool of unweighted existing classifiers
      end for
      for t=1t = 1 to KK do
        Normalize sample weights: wi←wi∑k=1Nwkw_i \leftarrow \frac{w_i}{\sum_{k=1}^N w_k}
        Train new weak classifier ht(x)h_t(x) on current frame data
        Compute error errt\text{err}_t and weight αt\alpha_t
        if errt≥0.4\text{err}_t \ge 0.4 then
          Abort training new weak classifier for this step
        end if
        Update sample weights {wi}i=1N\{w_i\}_{i=1}^N
      end for
      Update strong classifier H(x)=∑t=1Tαtht(x)H(x) = \sum_{t=1}^T \alpha_t h_t(x)

    In typical implementations, K=1K = 1 weak classifier is added and removed per frame, and T=5T = 5 weak classifiers are retained in the temporal ensemble.

  2. Knowl 2 — Pixel Feature Representation and Least-Squares Weak Classifier

    model/method

    Each pixel in an image is represented as a dd-dimensional feature vector x∈Rdx \in \mathbb{R}^d capturing local appearance and geometry:

    1. Color Sequences (d=11d=11): An 11-dimensional feature vector consisting of an 8-bin local orientation histogram computed over a 5×55 \times 5 pixel neighborhood around the pixel, concatenated with the 3 color channels (R,G,BR, G, B). To ensure robustness against noise, only gradient edges with intensity differences greater than or equal to a threshold of 10 intensity levels are counted in the histogram.
    2. Grayscale Sequences (d=9d=9): A 9-dimensional feature vector consisting of the 8-bin local orientation histogram (with gradient threshold 10 over a 5×55 \times 5 window) concatenated with the single grayscale intensity value.

    The weak learner trains a linear binary classifier h(x):Rd→{−1,+1}h(x): \mathbb{R}^d \to \{-1, +1\} using a least-squares formulation on weighted positive (object) and negative (background) pixel samples. While least-squares classifiers require slightly more computation than color histograms, they scale efficiently to high dimensions (d=9d=9 or d=11d=11). Combining these classifiers with spatial mean shift enables indirect mean shift optimization over high-dimensional feature spaces.

  3. Knowl 3 — Confidence-Weight-Based Outlier Rejection for Online Labeling

    model/method

    Because object bounding rectangles generally enclose non-rectangular objects, some background pixels are located inside the positive bounding box. AdaBoost is sensitive to such label noise and places disproportionately large weights on difficult or mislabeled pixels. To prevent these outliers from corrupting subsequent weak learners, an outlier rejection rule is applied during label assignment:

    yi={+1if inside(rj,pi)∧(wi<3N)−1otherwisey_i = \begin{cases} +1 & \text{if } \text{inside}(r_j, p_i) \land \left(w_i < \frac{3}{N}\right) \\ -1 & \text{otherwise} \end{cases}

    where:

    • pip_i is the 2D spatial coordinate of pixel sample ii,
    • rjr_j is the current estimated object bounding rectangle at frame jj,
    • inside(rj,pi)\text{inside}(r_j, p_i) evaluates to true if pixel pip_i lies within rectangle rjr_j,
    • wiw_i is the normalized sample weight of pixel ii produced by testing the sample against the current strong classifier, with ∑k=1Nwk=1\sum_{k=1}^N w_k = 1,
    • NN is the total number of pixel samples within the search region, setting the outlier threshold to Θ=3N\Theta = \frac{3}{N}.

    Pixels located inside the bounding rectangle whose classification difficulty exceeds Θ\Theta have their label flipped from +1+1 to −1-1, filtering out background clutter within the object box and yielding cleaner confidence maps.

  4. Knowl 4 — Margin-to-Confidence Transformation and Mean Shift Object Localization

    model/method

    Given a strong classifier H(x)=∑t=1Tαtht(x)H(x) = \sum_{t=1}^T \alpha_t h_t(x), where ht(x)∈{−1,+1}h_t(x) \in \{-1, +1\} and αt>0\alpha_t > 0, the classification margin for a pixel at spatial position pp with feature vector x(p)x(p) is converted into a normalized confidence value c(x(p))∈[0,1]c(x(p)) \in [0, 1] via the mapping:

    c(x(p))={0if H(x(p))<0H(x(p))max⁡qH(x(q))if H(x(p))≥0c(x(p)) = \begin{cases} 0 & \text{if } H(x(p)) < 0 \\ \frac{H(x(p))}{\max_{q} H(x(q))} & \text{if } H(x(p)) \ge 0 \end{cases}

    where negative margins are clipped to zero and non-negative margins are linearly scaled into the interval [0,1][0, 1].

    The 2D array of values c(x(p))c(x(p)) across the search region constitutes the spatial confidence map LjL_j for frame jj. Mean shift is executed on LjL_j, initialized at the previous bounding rectangle center rj−1r_{j-1}, iteratively climbing the density gradient to find the local peak (mode) representing the updated object bounding rectangle rjr_j.

  5. Knowl 5 — Multi-Resolution Feature Integration via Pyramid Confidence Maps

    model/method

    To capture features and spatial structures at multiple spatial scales, ensemble tracking runs in parallel across multiple levels of an image pyramid (e.g., three levels: original resolution 1×1\times, half-size 12×\frac{1}{2}\times, and quarter-size 14×\frac{1}{4}\times):

    1. At each pyramid level ss, an independent strong classifier H(s)(x)H^{(s)}(x) is maintained and updated with its own weak classifier per frame.
    2. In each frame, each pyramid level computes a multi-scale confidence map Lj(s)L_j^{(s)} over its downscaled input image.
    3. All individual confidence maps {Lj(s)}\{L_j^{(s)}\} are spatially rescaled (interpolated) back to the dimensions of the original image.
    4. The rescaled confidence maps are combined via a weighted average (or score-based combination) into a single composite confidence map LjL_j.
    5. Mean shift optimization is executed on the combined map LjL_j to determine the updated object position rjr_j.
  6. Knowl 6 — Dynamic Cue Re-Weighting Under Background Camouflage

    empirical result

    In a video sequence tracking a pedestrian crossing a street who passes in front of a vehicle with the exact same color, the ensemble tracker automatically modulates classifier feature weights across time. When the pedestrian is in front of the road, RGB color features carry the largest positive weights among the weak classifiers. When the pedestrian stands directly in front of the identically colored vehicle, color loses its discriminative power; the online ensemble adapts by automatically increasing the weights of the 8-bin local orientation histogram features, successfully maintaining continuous track of the pedestrian through complete background color camouflage.

  7. Knowl 7 — Tracking Robustness of Online Ensemble Updating Compared to Static Classifiers

    empirical result

    In a comparative evaluation on a 90-frame video sequence exhibiting background and appearance variation:

    • An adaptive ensemble tracker (updating K=1K=1 weak classifier per frame in an ensemble of T=5T=5 classifiers) maintains accurate tracking across the entire 90 frames.
    • A static tracker (trained with 5 weak classifiers on frame 1 without subsequent updates) loses the target and permanently locks onto background clutter at frame 30.

    This demonstrates that continuous online replacement and re-weighting of weak classifiers is necessary to accommodate changing target appearances and dynamic backgrounds.

  8. Knowl 8 — Tracking in Grayscale Video Sequences using 9D Orientation-Intensity Features

    empirical result

    Conventional mean shift tracking based on color histograms often fails on single-channel grayscale video due to insufficient discriminative information in a 1D color histogram. Evaluated on a 225-frame grayscale sequence tracking a car moving through an intersection recorded from a freely moving camera, the ensemble tracker using a 9-dimensional feature vector (8-bin orientation histogram from a 5×55 \times 5 window with edge threshold 10 plus 1 intensity value) successfully tracks the car across all 225 frames, with the learned confidence maps progressively conforming to the vehicle's geometric shape over time.

Coverage note — None was omitted; all significant contributed methods, algorithms, and experimental evaluations from the paper are captured.

References

  1. 1.Avidan S., Support Vector Tracking. IEEE Trans. on Pattern Analysis and Machine Intelligence, 2004.
  2. 2.Black, M. J. and Jepson, A. EigenTracking: Robust matching and tracking of articulated objects using a view-based representation. International Journal of Computer Vision, 26(1), pp. 63-84, 1998.
  3. 3.Bobick, A., S.Intille, J.Davis, F.Baird, C.Pinhanez, L.Campbell, Y.Ivanov, A.Schutte, and A.Wilson. The KidsRoom. In Communications of the ACM, 43(3). 2000
  4. 4.Collins T. R., Liu, Y. On-Line Selection of Discriminative Tracking Features. Proceedings of the International Conference on Computer Vision (ICCV '03), France, 2003.
  5. 5.Comanciu, D., Visvanathan R., Meer, P. Kernel-Based Object Tracking. IEEE Trans. on Pattern Analysis and Machine Intelligence (PAMI), 25:5, pp 564-575, 2003.
  6. 6.Crowley, J., Berard, F. Multi-Modal Tracking of Faces for Video Communications. Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR '97), Puerto Rico 1997.
  7. 7.Darrell, T., Gordon, G., Harville, M., and Woodfill, J. Integrated person tracking using stereo, color, and pattern detection. Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR '98), pp. 601-609, Santa Barbara, June 1998.
  8. 8.Freund, Y. An adaptive version of the boost by majority algorithm. In Machine Learning, 43(3):293-318, June 2001.
  9. 9.Freund, Y. Schapire, R. E. A decision-theoretic generalization of on-line learning and an application to boosting. In Computational Learning Theory: Eurocolt 95, pp 23-37, 1995.
  10. 10.Georescu, B., Shimshoni, I., Meer, P. Mean Shift Based Clustering in High Dimensions: A Texture Classification Example. Proceedings of the International Conference on Computer Vision (ICCV '03), France, 2003.
  11. 11.Ho, J., Lee K., Yang, M., Kriegman, D. Visual Tracking Using Learned Linear Subspaces. In IEEE Conf. on Computer Vision and Pattern Recognition, 2004.
  12. 12.Isard, M., Blake, A. CONDENSATION - Conditional Density Propagation for Visual Tracking, International Journal of Computer Vision, Vol 29(1), pp:5-28, 1998.
  13. 13.Jepson, A.D., Fleet, D.J. and El-Maraghi, T. Robust. on-line appearance models for vision tracking. In IEEE Transactions on Pattern Analysis and Machine Intelligence 25(10):1296-1311.
  14. 14.Levi, K., Weiss, Y. Learning Object Detection from a Small Number of Examples: The Importance of Good Features. In IEEE Conf. on Computer Vision and Pattern Recognition, 2004.
  15. 15.Stauffer, C. and E. Grimson, Learning Patterns of Activity Using Real-Time Tracking, PAMI, 22(8):747-757, 2000.

Citation

MLA
Avidan, S. “Ensemble Tracking”. 2005 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR'05), vol. 2, 2005, pp. 494–501, https://doi.org/10.1109/CVPR.2005.144.
APA
Avidan, S. (2005). Ensemble Tracking. 2005 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR'05), 2, 494–501. https://doi.org/10.1109/CVPR.2005.144
Chicago
Avidan, S. 2005. “Ensemble Tracking”. 2005 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR'05) 2: 494–501. https://doi.org/10.1109/CVPR.2005.144.
Harvard
Avidan, S. (2005) “Ensemble Tracking”, 2005 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR'05). IEEE, pp. 494–501. Available at: https://doi.org/10.1109/CVPR.2005.144.
Vancouver
1. Avidan S (2005) Ensemble Tracking. In: 2005 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR'05). IEEE, pp 494–501

BibTeX

@inproceedings{Avidan, title={Ensemble Tracking}, volume={2}, url={http://dx.doi.org/10.1109/CVPR.2005.144}, DOI={10.1109/cvpr.2005.144}, booktitle={2005 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR′05)}, publisher={IEEE}, author={Avidan, S.}, pages={494–501} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF