Robust Object Tracking with Online Multiple Instance Learning

Boris BabenkoMing-Hsuan YangSerge Belongie

article2011TPAMI2,097 citations

Presents an online Multiple Instance Learning framework for visual object tracking that handles label ambiguity in self-training classifiers to prevent tracking drift during real-time video processing.

Listen

Visual object tracking is a foundational computer vision capability required in surveillance, robotics, and automated video analysis. A prominent real-time approach, known as tracking-by-detection, updates an internal classifier in every video frame to separate the target object from its surrounding background. However, these systems face a critical self-training vulnerability: when an object experiences sudden motion, illumination shifts, or partial obstruction, the tracker can select an imprecise bounding box. Updating standard supervised classifiers with slightly misaligned examples introduces labeling noise, which quickly degrades the model and leads to tracking failure or persistent drift.

To overcome this limitation, the article presents and evaluates MILTrack, a tracking framework powered by a novel online Multiple Instance Learning algorithm. Instead of assuming that a single cropped image patch represents the exact target, the system groups a set of candidate patches near the predicted location into a positive set, termed a bag. The algorithm requires only that at least one instance within the bag represents the true object, allowing the learner to resolve visual ambiguities autonomously during model updates.

Across multiple challenging video sequences featuring out-of-plane rotations, fast motion, and occlusions, the proposed method substantially outperformed existing state-of-the-art baselines. Evaluated on both average center location error and tracking precision at a 20-pixel threshold, MILTrack maintained superior stability without needing sequence-specific parameter adjustments. In heavily occluded and fast-moving scenarios, such as the challenging tiger toy sequences, MILTrack reduced tracking errors by more than half compared to traditional online boosting. Furthermore, the experiments showed that simply feeding multiple positive patches into standard supervised learners degraded their accuracy, proving that the multi-instance formulation is what drives the performance gains. The framework also successfully incorporated scale adaptation and achieved real-time execution speeds of approximately 25 frames per second.

These findings indicate that handling data ambiguity directly within the learning algorithm creates a significantly more resilient visual tracking system. In operational environments, this translates to reduced operational drift, lower risk of target loss, and minimized overhead from manual parameter tuning. Practitioners seeking robust object tracking should consider adopting online multiple-instance formulations for adaptive appearance modeling.

Despite these strengths, the article notes that adaptive appearance models still face inherent limitations when a target is completely occluded for extended durations or entirely leaves the camera view. Future development should focus on integrating online multi-instance appearance tracking with pre-trained object detectors to re-acquire targets after prolonged disappearances, as well as extending the formulation to articulated or non-rigid targets.

  • Paper: High-Speed Tracking with Kernelized Correlation Filters, João F. Henriques et al. (2014). This paper directly builds on the tracking-by-detection paradigm by formulating Kernelized Correlation Filters to eliminate redundancies in training samples while vastly improving speed.
  • Paper: Tracking-Learning-Detection, Zdenek Kalal et al. (2012). This work extends online tracking-by-detection principles by combining frame-by-frame tracking with an explicit online learning framework and detector error correction.
Cover for Robust Object Tracking with Online Multiple Instance Learning

Abstract

In this paper, we address the problem of tracking an object in a video given its location in the first frame and no other information. Recently, a class of tracking techniques called "tracking by detection" has been shown to give promising results at real-time speeds. These methods train a discriminative classifier in an online manner to separate the object from the background. This classifier bootstraps itself by using the current tracker state to extract positive and negative examples from the current frame. Slight inaccuracies in the tracker can therefore lead to incorrectly labeled training examples, which degrade the classifier and can cause drift. In this paper, we show that using Multiple Instance Learning (MIL) instead of traditional supervised learning avoids these problems and can therefore lead to a more robust tracker with fewer parameter tweaks. We propose a novel online MIL algorithm for object tracking that achieves superior results with real-time performance. We present thorough experimental results (both qualitative and quantitative) on a number of challenging video clips.

Table of Contents

  • 1 INTRODUCTION
  • 2 ADAPTIVE APPEARANCE MODELS
  • 3 TRACKING WITH ONLINE MIL
  • 3.1 System Overview and Motion Model
  • 3.2 Multiple Instance Learning
  • 3.3 Online Boosting
  • 3.4 Online Multiple Instance Boosting
  • 3.5 Discussion
  • 3.6 Implementation Details
  • 3.6.1 Weak Classifiers
  • 3.6.2 Image Features
  • 4 EXPERIMENTS
  • 4.1 Evaluation Methodology
  • 4.2 Tracking Object Location
  • 4.2.1 Sylvester and David Indoors
  • 4.2.2 Occluded Face, Occluded Face 2
  • 4.2.3 Cola Can, Surfer
  • 4.2.4 Tiger 1, Tiger 2
  • 4.2.5 Coupon Book
  • 4.3 Tracking Object Location and Scale
  • 4.3.1 David Indoor
  • 4.3.2 Snack Bar
  • 4.3.3 Tea Box
  • 5 DISCUSSION/CONCLUSIONS
  • ACKNOWLEDGMENTS
  • REFERENCES

Knowls

  1. Knowl 1 — Online Multiple Instance Boosting Algorithm

    algorithm

    Online Multiple Instance Boosting (Online MILBoost) trains an additive strong classifier H(x)=∑k=1Khk(x)H(x) = \sum_{k=1}^K h_k(x) in an online setting where training examples are grouped into labeled bags rather than individually labeled instances. In Multiple Instance Learning (MIL), a bag is positive (yi=1y_i = 1) if at least one instance inside it is positive, and negative (yi=0y_i = 0) if all instances inside it are negative.

    Input: Training dataset of bags and bag labels {(Xi,yi)}i=1N\{(X_i, y_i)\}_{i=1}^N where Xi={xi1,xi2,…,ximi}X_i = \{x_{i1}, x_{i2}, \dots, x_{im_i}\} and yi∈{0,1}y_i \in \{0, 1\}; candidate pool of MM weak stump classifiers {h1,…,hM}\{h_1, \dots, h_M\}; number of selected weak classifiers KK.
    Output: Additive strong classifier H(x)=∑k=1Khk(x)H(x) = \sum_{k=1}^K h_k(x), yielding instance conditional probability p(y=1∣x)=σ(H(x))p(y=1|x) = \sigma(H(x)).
    1: for m=1m = 1 to MM do
    2: Update candidate weak classifier hmh_m in parallel using all instances xijx_{ij} with assigned label yiy_i.
    3: end for
    4: Initialize Hij=0H_{ij} = 0 for all i∈{1,…,N}i \in \{1, \dots, N\} and j∈{1,…,mi}j \in \{1, \dots, m_i\}.
    5: for k=1k = 1 to KK do
    6: for m=1m = 1 to MM do
    7: for all i,ji, j do
    8: pijm=σ(Hij+hm(xij))p_{ij}^m = \sigma(H_{ij} + h_m(x_{ij}))
    9: end for
    10: for all ii do
    11: pim=1−∏j=1mi(1−pijm)p_i^m = 1 - \prod_{j=1}^{m_i} (1 - p_{ij}^m)
    12: end for
    13: Lm=∑i=1N(yilog⁡(pim)+(1−yi)log⁡(1−pim))\mathcal{L}^m = \sum_{i=1}^N \left( y_i \log(p_i^m) + (1 - y_i) \log(1 - p_i^m) \right)
    14: end for
    15: m∗=arg⁡max⁡m∈{1,…,M}Lmm^* = \arg\max_{m \in \{1, \dots, M\}} \mathcal{L}^m
    16: hk(x)←hm∗(x)h_k(x) \leftarrow h_{m^*}(x)
    17: for all i,ji, j do
    18: Hij←Hij+hk(xij)H_{ij} \leftarrow H_{ij} + h_k(x_{ij})
    19: end for
    20: end for
    21: return H(x)=∑k=1Khk(x)H(x) = \sum_{k=1}^K h_k(x)

    In typical tracking implementations, M=250M = 250 candidate weak classifiers are maintained, from which K=50K = 50 classifiers are greedily selected at each frame update. The sigmoid link is σ(v)=11+e−v\sigma(v) = \frac{1}{1 + e^{-v}}.

  2. Knowl 2 — MILTrack Object Tracking Framework

    algorithm

    MILTrack performs visual tracking by combining a greedy motion model with an adaptive appearance model updated via Multiple Instance Learning. At each video frame, candidate image patches are evaluated in a local search neighborhood, the target state is updated to the maximum response location, and positive and negative bags are extracted to update the online classifier.

    Input: Video frame at time step tt, previous object location lt−1l_{t-1}, search radius ss, positive bag radius rr (r<sr < s), outer negative radius β\beta, number of negative samples NnegN_{\text{neg}}, strong MIL classifier H(x)H(x).
    Output: Estimated target location ltl_t, updated classifier H(x)H(x).
    1: Crop candidate image patches within search window: Xs={x:∥l(x)−lt−1∥<s}X^s = \{x : \|l(x) - l_{t-1}\| < s\}.
    2: Compute Haar-like feature vector for every patch x∈Xsx \in X^s.
    3: Estimate object presence probability for all x∈Xsx \in X^s: p(y=1∣x)=σ(H(x))p(y=1|x) = \sigma(H(x)).
    4: Update object location greedily: lt=l(arg⁡max⁡x∈Xsp(y=1∣x))l_t = l(\arg\max_{x \in X^s} p(y=1|x)).
    5: Crop positive bag patches: Xr={x:∥l(x)−lt∥<r}X^r = \{x : \|l(x) - l_t\| < r\}, and assign bag label y=1y = 1.
    6: Crop negative candidate patches: Xr,β={x:r<∥l(x)−lt∥<β}X^{r,\beta} = \{x : r < \|l(x) - l_t\| < \beta\}.
    7: Randomly sample NnegN_{\text{neg}} patches from Xr,βX^{r,\beta} and place each into its own negative bag with label y=0y = 0.
    8: Update H(x)H(x) with Online Multiple Instance Boosting using positive bag XrX^r and the NnegN_{\text{neg}} negative bags.
    9: return lt,H(x)l_t, H(x)

    Standard default tracker parameters are search radius s=35s = 35 pixels, positive radius r=4r = 4 pixels (producing ∣Xr∣=45|X^r| = 45 patches), outer negative radius β=50\beta = 50 pixels, and Nneg=65N_{\text{neg}} = 65 randomly sampled negative patches (with 1,000 negative patches sampled on initial frame initialization).

  3. Knowl 3 — Online Weak Classifier Parameterization and Update Rules

    model/method

    Each weak classifier h(x)h(x) in the online boosting pool is associated with a single randomly generated Haar-like feature f(x)f(x) (computed via integral images from 2 to 4 weighted rectangles). The weak learner models class-conditional feature distributions as Gaussians:

    p(f(x)∣y=1)∼N(μ1,σ1),p(f(x)∣y=0)∼N(μ0,σ0)p(f(x)|y = 1) \sim \mathcal{N}(\mu_1, \sigma_1), \quad p(f(x)|y = 0) \sim \mathcal{N}(\mu_0, \sigma_0)

    Assuming equal class priors p(y=1)=p(y=0)p(y = 1) = p(y = 0), the weak classifier returns the log odds ratio:

    h(x)=log⁡(p(y=1∣f(x))p(y=0∣f(x)))=log⁡(p(f(x)∣y=1)p(f(x)∣y=0))h(x) = \log\left( \frac{p(y = 1 | f(x))}{p(y = 0 | f(x))} \right) = \log\left( \frac{p(f(x) | y = 1)}{p(f(x) | y = 0)} \right)

    When a batch of nn instances {(x1,y1),…,(xn,yn)}\{(x_1, y_1), \dots, (x_n, y_n)\} is presented to the weak learner, the positive parameters (μ1,σ1)(\mu_1, \sigma_1) are updated online using a learning rate γ∈(0,1)\gamma \in (0, 1):

    μ1←γμ1+(1−γ)1n∑i:yi=1f(xi)\mu_1 \leftarrow \gamma \mu_1 + (1 - \gamma) \frac{1}{n} \sum_{i: y_i = 1} f(x_i)

    σ1←γσ1+(1−γ)1n∑i:yi=1(f(xi)−μ1)2\sigma_1 \leftarrow \gamma \sigma_1 + (1 - \gamma) \sqrt{\frac{1}{n} \sum_{i: y_i = 1} (f(x_i) - \mu_1)^2}

    Parameters μ0\mu_0 and σ0\sigma_0 are updated using identical formulas over instances where yi=0y_i = 0. The default learning rate parameter is γ=0.85\gamma = 0.85.

  4. Knowl 4 — Noisy-OR Bag Probability and Objective Function in Multiple Instance Tracking

    equation

    In Multiple Instance Learning for tracking, the instance-level classifier output H(x)H(x) is mapped to an instance probability via the sigmoid function:

    p(y=1∣x)=σ(H(x))=11+e−H(x)p(y = 1 | x) = \sigma(H(x)) = \frac{1}{1 + e^{-H(x)}}

    For a bag Xi={xi1,xi2,…,ximi}X_i = \{x_{i1}, x_{i2}, \dots, x_{im_i}\}, the bag probability p(yi=1∣Xi)p(y_i = 1 | X_i) is formulated using the Noisy-OR (NOR) model:

    p(yi=1∣Xi)=1−∏j=1mi(1−p(yi=1∣xij))p(y_i = 1 | X_i) = 1 - \prod_{j=1}^{m_i} (1 - p(y_i = 1 | x_{ij}))

    The learning objective maximized during online boosting is the bag log-likelihood:

    L=∑i(yilog⁡p(yi=1∣Xi)+(1−yi)log⁡(1−p(yi=1∣Xi)))\mathcal{L} = \sum_{i} \left( y_i \log p(y_i = 1 | X_i) + (1 - y_i) \log(1 - p(y_i = 1 | X_i)) \right)

    Under this formulation, a positive bag achieves a high probability if at least one instance has a high probability. For negative bags, where all instances must be negative, placing all negative instances into a single bag yields an identical loss to placing each negative instance into its own separate bag.

  5. Knowl 5 — Scale-Space Extension in MILTrack

    model/method

    To track object scale alongside 2D translation, a discrete scale parameter SS is defined as the scale step size. At time step tt, candidate image patches are extracted across three scale hypotheses: the current scale ltsl_t^s, one scale step larger lts+Sl_t^s + S, and one scale step smaller lts−Sl_t^s - S.

    The tracker evaluates the appearance classifier p(y=1∣x)p(y=1|x) across all candidate patches from these three scales within the search radius ss, selecting the patch with the highest probability response to jointly update the tracker's position and scale state. Appearance model training patches for positive and negative bags are then cropped exclusively from the newly selected optimal scale.

  6. Knowl 6 — Object Location Tracking Performance Comparison

    data/table

    The quantitative evaluation measures object location tracking across nine video sequences comparing MILTrack(45) (r=4r=4, 45 positive patches per bag) against Online AdaBoost with 1 positive sample (OAB(1)), Online AdaBoost with 45 positive samples (OAB(45)), SemiBoost, and FragTrack. Trackers are evaluated on Average Center Location Error (pixels) and Precision at a fixed Euclidean distance threshold of 20 pixels (corresponding approximately to ≥50%\ge 50\% bounding box overlap).

    Average Center Location Error (px) Precision at Threshold 20 (Score ∈[0,1]\in [0, 1])
    Video Clip OAB(1) OAB(45) SemiBoost Frag MILTrack(45) OAB(1) OAB(45) SemiBoost Frag MILTrack(45)
    Sylvester 25 79 16 11 11 0.64 0.04 0.69 0.86 0.90
    David Indoor 49 72 39 46 23 0.16 0.08 0.46 0.45 0.52
    Cola Can 25 57 13 63 20 0.45 0.16 0.78 0.14 0.55
    Occluded Face 43 105 7 6 27 0.22 0.02 0.97 0.95 0.43
    Occluded Face 2 21 93 23 45 20 0.61 0.03 0.60 0.44 0.60
    Surfer 23 43 9 139 11 0.51 0.33 0.96 0.28 0.93
    Tiger 1 35 57 42 39 16 0.48 0.22 0.44 0.28 0.81
    Tiger 2 33 33 61 37 18 0.51 0.40 0.30 0.22 0.83
    Coupon Book 25 58 67 56 15 0.67 0.15 0.37 0.41 0.69

    MILTrack(45) achieves the lowest center error on 6 out of 9 sequences and the highest precision on 6 out of 9 sequences, maintaining stability during large appearance changes, out-of-plane rotations (Tiger 1 and Tiger 2), and distractor objects (Coupon Book).

  7. Knowl 7 — Performance Degradation of Supervised Online Boosting with Multiple Positives vs MIL Robustness

    empirical result

    In traditional supervised tracking-by-detection, extracting multiple positive examples in a small neighborhood (r=4r = 4, yielding 45 positive instances) causes severe performance degradation compared to extracting a single positive example (r=1r = 1). When all 45 patches are labeled positive in Online AdaBoost (OAB(45)), the classifier is forced to find a decision boundary treating slightly misaligned or partially occluded patches as true positives, confusing the appearance model and causing tracker drift.

    Experimentally, OAB(45) location error increases from 25 to 79 pixels on Sylvester, from 49 to 72 pixels on David Indoor, and from 43 to 105 pixels on Occluded Face compared to OAB(1). In contrast, placing all 45 patches into a single positive bag in Multiple Instance Learning (MILTrack(45)) allows the learning algorithm the flexibility to select features that maximize bag likelihood based on the single best instance, reducing location error to 11 pixels on Sylvester, 23 pixels on David Indoor, and 27 pixels on Occluded Face.

  8. Knowl 8 — Object Location and Scale Tracking Evaluation

    data/table

    Location and scale tracking evaluation compares Online AdaBoost with scale (extOABs(1) ext{OAB}_s(1)), Incremental Visual Tracking with scale (extIVTs ext{IVT}_s), fixed-scale MILTrack (extMILTrack(45) ext{MILTrack}(45)), and multi-scale MILTrack (extMILTracks(45) ext{MILTrack}_s(45)) across three challenging video sequences containing scale variations and background clutter.

    Location Mean Error (pixels) Precision at Threshold 20 (Score ∈[0,1]\in [0, 1])
    Video Clip OABs(1)\text{OAB}_s(1) IVTs\text{IVT}_s MILTrack(45)\text{MILTrack}(45) MILTracks(45)\text{MILTrack}_s(45) OABs(1)\text{OAB}_s(1) IVTs\text{IVT}_s MILTrack(45)\text{MILTrack}(45) MILTracks(45)\text{MILTrack}_s(45)
    David Indoor 28 5 23 20 0.13 0.98 0.52 0.75
    Snack Bar 18 30 12 9 0.76 0.57 0.90 0.98
    Tea Box 17 14 10 10 0.74 0.70 0.91 0.87

    While generative subspace tracking (extIVTs ext{IVT}_s) excels on faces with moderate lighting changes (David Indoor, achieving 5 px error), it fails on textured backgrounds and out-of-plane rotations (Snack Bar error of 30 px, Tea Box error of 14 px) because it does not incorporate negative background models. MILTracks(45)\text{MILTrack}_s(45) achieves the lowest error on Snack Bar (9 px) and Tea Box (10 px) and improves precision over fixed-scale MILTrack.

  9. Knowl 9 — Limitations of MILTrack Adaptive Appearance Tracking

    limitation

    The online Multiple Instance Learning tracking framework exhibits specific failure modes:

    1. Prolonged Full Occlusion and Target Disappearance: Because the appearance model continuously updates in an online fashion, prolonged total occlusion or target departure from the field of view forces the positive bag to be sampled entirely from background or occluder pixels, leading to unavoidable model corruption and permanent drift.
    2. Bounding Box Representation of Articulated Objects: The use of single rectangular bounding boxes cannot accurately isolate non-rigid, articulated, or highly deformable targets from background clutter, which requires part-based or deformable contour models.
    3. Weak Classifier Approximations: Passing bag labels directly to update instance-level weak classifiers and greedily selecting weak classifiers sequentially based on current-frame log-likelihood approximates the joint batch MIL optimization, relying on the weak classifiers' internal history parameters to prevent catastrophic forgetting.

Coverage note — None was omitted; all contributed algorithms, equations, empirical tracking comparisons, scale tracking models, and system limitations are fully represented.

References

  1. 1.S. Birchfield, “Elliptical Head Tracking Using Intensity Gradients and Color Histograms,” Proc. IEEE Conf. Computer Vision and Pattern Recognition, pp. 232-237, 1998.
  2. 2.M. Isard and J. Maccormick, “Bramble: A Bayesian Multiple-Blob Tracker,” Proc. IEEE Int’l Conf. Computer Vision, vol. 2, pp. 34-41, 2001.
  3. 3.K. Branson and S. Belongie, “Tracking Multiple Mouse Contours (without Too Many Samples),” Proc. IEEE Conf. Computer Vision and Pattern Recognition, vol. 1, 2005.
  4. 4.V. Lepetit and P. Fua, “Keypoint Recognition Using Randomized Trees,” IEEE Trans. Pattern Analysis and Machine Intelligence, vol. 28, no. 9, pp. 1465-1479, Sept. 2006.
  5. 5.A. Yilmaz, O. Javed, and M. Shah, “Object Tracking: A Survey,” ACM Computing Surveys, vol. 38, no. 4, 2006.
  6. 6.G. Hager and P. Belhumeur, “Efficient Region Tracking with Parametric Models of Geometry and Illumination,” IEEE Trans. Pattern Analysis and Machine Intelligence, vol. 20, no. 10, pp. 1025-1039, Oct. 1998.
  7. 7.M. Black and A. Jepson, “Eigentracking: Robust Matching and Tracking of Articulated Objects Using a View-Based Representation,” Int’l J. Computer Vision , vol. 26, no. 1, pp. 63-84, 1998.
  8. 8.D. Comaniciu, V. Ramesh, and P. Meer, “Real-Time Tracking of Non-Rigid Objects Using Mean Shift,” Proc. IEEE Conf. Computer Vision and Pattern Recognition, vol. 2, pp. 142-149, 2000.
  9. 9.A. Adam, E. Rivlin, and I. Shimshoni, “Robust Fragments-Based Tracking Using the Integral Histogram,” Proc. IEEE Conf. Computer Vision and Pattern Recognition, vol. 1, pp. 798-805, 2006.
  10. 10.A.D. Jepson, D.J. Fleet, and T.F. El-Maraghi, “Robust Online Appearance Models for Visual Tracking,” IEEE Trans. Pattern Analysis and Machine Intelligence, vol. 25, no. 10, pp. 1296-1311, Oct. 2003.
  11. 11.I. Matthews, T. Ishikawa, and S. Baker, “The Template Update Problem,” IEEE Trans. Pattern Analysis and Machine Intelligence, vol. 26, no. 6, pp. 810-815, June 2004.
  12. 12.D. Ross, J. Lim, R.-S. Lin, and M.-H. Yang, “Incremental Learning for Robust Visual Tracking,” Int’l J. Computer Vision, vol. 77, no. 1, pp. 125-141, 2008.
  13. 13.T.G. Dietterich, R.H. Lathrop, and L.T. Perez, “Solving the Multiple-Instance Problem with Axis Parallel Rectangles,” Artificial Intelligence, vol. 89, pp. 31-71, 1997.
  14. 14.P. Viola, J.C. Platt, and C. Zhang, “Multiple Instance Boosting for Object Detection,” Proc. Neural Information Processing Systems, pp. 1417-1426, 2005.
  15. 15.P. Dolla´r, B. Babenko, S. Belongie, P. Perona, and Z. Tu, “Multiple Component Learning for Object Detection,” Proc. European Conf. Computer Vision, 2008.
  16. 16.S. Andrews, I. Tsochantaridis, and T. Hofmann, “Support Vector Machines for Multiple-Instance Learning,” Proc. Neural Information Processing Systems, pp. 577-584, 2003.
  17. 17.C. Galleguillos, B. Babenko, A. Rabinovich, and S. Belongie, “Weakly Supervised Object Recognition and Localization with Stable Segmentations,” Proc. European Conf. Computer Vision, 2008.
  18. 18.S. Vijayanarasimhan and K. Grauman, “Keywords to Visual Categories: Multiple-Instance Learning for Weakly Supervised Object Categorization,” Proc. IEEE Conf. Computer Vision and Pattern Recognition, 2008.
  19. 19.K. Okuma, A. Taleghani, N. De Freitas, J. Little, and D. Lowe, “A Boosted Particle Filter: Multitarget Detection and Tracking,” Proc. European Conf. Computer Vision, pp. 28-39, 2004.
  20. 20.M. Isard and A. Blake, “Contour Tracking by Stochastic Propagation of Conditional Density,” Proc. European Conf. Computer Vision, vol. 1064, pp. 343-356, 1996.
  21. 21.L. Vese and T. Chan, “A Multiphase Level Set Framework for Image Segmentation Using the Mumford and Shah Model,” Int’l J. Computer Vision , vol. 50, no. 3, pp. 271-293, 2002.
  22. 22.M. Salzmann, V. Lepetit, and P. Fua, “Deformable Surface Tracking Ambiguities,” Proc. IEEE Conf. Computer Vision and Pattern Recognition, 2007.
  23. 23.A.O. Balan and M.J. Black, “An Adaptive Appearance Model Approach for Model-Based Articulated Object Tracking,” Proc. IEEE Conf. Computer Vision and Pattern Recognition, vol. 1, pp. 758-765, 2006.
  24. 24.R. Lin, D. Ross, J. Lim, and M.-H. Yang, “Adaptive Discriminative Generative Model and Its Applications,” Proc. Neural Information Processing Systems, pp. 801-808, 2004.
  25. 25.H. Grabner, M. Grabner, and H. Bischof, “Real-Time Tracking via Online Boosting,” Proc. Conf. British Machine Vision, pp. 47-56, 2006.
  26. 26.X. Liu and T. Yu, “Gradient Feature Selection for Online Boosting,” Proc. IEEE Int’l J. Computer Vision, pp. 1-8, 2007.
  27. 27.S. Avidan, “Ensemble Tracking,” Proc. IEEE Conf. Computer Vision and Pattern Recognition, vol. 2, pp. 494-501, 2005.
  28. 28.S. Avidan, “Support Vector Tracking,” IEEE Trans. Pattern Analysis and Machine Intelligence, vol. 26, no. 8, pp. 1064-1072, Aug. 2004.
  29. 29.J. Wang, X. Chen, and W. Gao, “Online Selecting Discriminative Tracking Features Using Particle Filter,” Proc. IEEE Conf. Computer Vision and Pattern Recognition, vol. 2, pp. 1037-1042, 2005.
  30. 30.R.T. Collins, Y. Liu, and M. Leordeanu, “Online Selection of Discriminative Tracking Features,” IEEE Trans. Pattern Analysis and Machine Intelligence, vol. 27, no. 10, pp. 1631-1643, Oct. 2005.
  31. 31.G. Mori and J. Malik, “Recovering 3D Human Body Configurations Using Shape Contexts,” IEEE Trans. Pattern Analysis and Machine Intelligence, vol. 28, no. 7, pp. 1052-1062, July 2006.
  32. 32.P. Viola and M. Jones, “Rapid Object Detection Using a Boosted Cascade of Simple Features,” Proc. IEEE Conf. Computer Vision and Pattern Recognition, vol. 1, pp. 511-518, 2001.
  33. 33.H. Grabner, C. Leistner, and H. Bischof, “Semi-Supervised Online Boosting for Robust Tracking,” Proc. European Conf. Computer Vision, 2008.
  34. 34.N.C. Oza, “Online Ensemble Learning,” PhD Thesis, Univ. of California, 2001.
  35. 35.P. Dolla´r, Z. Tu, H. Tao, and S. Belongie, “Feature Mining for Image Classification,” Proc. IEEE Conf. Computer Vision and Pattern Recognition, 2007.
  36. 36.Z. Khan, T. Balch, and F. Dellaert, “A Rao-Blackwellized Particle Filter for Eigentracking,” Proc. IEEE Conf. Computer Vision and Pattern Recognition, vol. 2, 2004.
  37. 37.J.H. Friedman, “Greedy Function Approximation: A Gradient Boosting Machine,” The Annals of Statistics, vol. 29, no. 5, pp. 1189-1232, 2001.
  38. 38.B. Babenko, P. Dolla´r, Z. Tu, and S. Belongie, “Simultaneous Learning and Alignment: Multi-Instance and Multi-Pose Learning,” Proc. Faces in Real-Life Images, 2008.
  39. 39.Y. Freund and R.E. Schapire, “A Decision-Theoretic Generalization of Online Learning and an Application to Boosting,” J. Computer and System Sciences, vol. 55, pp. 119-139, 1997.
  40. 40.J. Friedman, T. Hastie, and R. Tibshirani, “Additive Logistic Regression: A Statistical View of Boosting,” The Annals of Statistics, vol. 28, no. 2, pp. 337-407, 2000.
  41. 41.C. Leistner, A. Saffari, P. Roth, and H. Bischof, “On Robustness of Online Boosting—A Competitive Study,” Proc. Third IEEE Workshop Online Computer Vision, 2009.
  42. 42.H. Grabner and H. Bischof, “Online Boosting and Vision,” Proc. IEEE Conf. Computer Vision and Pattern Recognition, pp. 260-267, 2006.
  43. 43.A. Chan and N. Vasconcelos, “Modeling, Clustering, and Segmenting Video with Mixtures of Dynamic Textures,” IEEE Trans. Pattern Analysis and Machine Intelligence, vol. 30, no. 5, pp. 909-926, May 2008.
  44. 44.M. Everingham, L. Van Gool, C.K.I. Williams, J. Winn, and A. Zisserman, “The PASCAL Visual Object Classes Challenge 2010 (VOC2010) Results,” http://www.pascal-network.org/challenges/VOC/voc2010/workshop/index.html, 2011.
  45. 45.S. Stalder, H. Grabner, and L. van Gool, “Beyond Semi-Supervised Tracking: Tracking Should Be as Simple as Detection, But Not Simpler than Recognition,” Proc. Workshop Online Learning in Computer Vision , 2009.
  46. 46.P. Felzenszwalb, D. McAllester, and D. Ramanan, “A Discriminatively Trained, Multiscale, Deformable Part Model,” Proc. IEEE Conf. Computer Vision and Pattern Recognition, 2008.
  47. 47.L. Mu, J. Kwok, and L. Bao-liang, “Online Multiple Instance Learning with No Regret,” Proc. IEEE Conf. Computer Vision and Pattern Recognitio, 2010.

Citation

MLA
Babenko, B., et al. “Robust Object Tracking with Online Multiple Instance Learning”. IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 33, no. 8, 2011, pp. 1619–32, https://doi.org/10.1109/TPAMI.2010.226.
APA
Babenko, B., Ming-Hsuan Yang, & Belongie, S. (2011). Robust Object Tracking with Online Multiple Instance Learning. IEEE Transactions on Pattern Analysis and Machine Intelligence, 33(8), 1619–1632. https://doi.org/10.1109/TPAMI.2010.226
Chicago
Babenko, B., Ming-Hsuan Yang, and S. Belongie. 2011. “Robust Object Tracking with Online Multiple Instance Learning”. IEEE Transactions on Pattern Analysis and Machine Intelligence 33 (8): 1619–32. https://doi.org/10.1109/TPAMI.2010.226.
Harvard
Babenko, B., Ming-Hsuan Yang and Belongie, S. (2011) “Robust Object Tracking with Online Multiple Instance Learning”, IEEE Transactions on Pattern Analysis and Machine Intelligence, 33(8), pp. 1619–1632. Available at: https://doi.org/10.1109/TPAMI.2010.226.
Vancouver
1. Babenko B, Ming-Hsuan Yang, Belongie S (2011) Robust Object Tracking with Online Multiple Instance Learning. IEEE Transactions on Pattern Analysis and Machine Intelligence 33:1619–1632

BibTeX

@article{Babenko_2011, title={Robust Object Tracking with Online Multiple Instance Learning}, volume={33}, ISSN={0162-8828}, url={http://dx.doi.org/10.1109/TPAMI.2010.226}, DOI={10.1109/tpami.2010.226}, number={8}, journal={IEEE Transactions on Pattern Analysis and Machine Intelligence}, publisher={Institute of Electrical and Electronics Engineers (IEEE)}, author={Babenko, B. and Ming-Hsuan Yang and Belongie, S.}, year={2011}, month=Aug, pages={1619–1632} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF