Tracking-Learning-Detection

Zdenek KalalKrystian MikolajczykJiri Matas

article2012TPAMI3,272 citations

Proposes a real-time framework that decomposes long-term visual tracking into tracking, detection, and self-correcting P-N learning to sustain tracking of unknown objects through full occlusions, camera disappearances, and severe appearance changes.

Listen

The paper addresses the challenge of long-term tracking of unknown objects in video streams captured by moving cameras. Objects frequently change appearance, undergo scale and illumination shifts, suffer partial or full occlusions, and move in or out of view, while processing must occur in real time and continue indefinitely. Existing trackers accumulate drift and fail on disappearance, while detectors require offline training and cannot handle novel objects.

The work develops and evaluates the TLD framework, which decomposes the task into three simultaneously operating components: a frame-to-frame tracker, a detector that localizes all previously seen appearances, and a learning process that identifies and corrects detector errors from the video itself. The learning component, called P-N learning, employs two independentexpertsthat estimate missed detections and false alarms, then augments the detector’s training set accordingly. The process is modeled as a discrete dynamical system whose stability conditions are derived analytically.

Experiments on six established benchmark sequences and four new, more demanding sequences demonstrate that the initial detector improves substantially after one pass through each video. On the new dataset the final detector reaches f-measures between 0.25 and 0.95, while the complete TLD system attains an overall f-measure of 0.81, more than three times higher than the strongest competing tracker. The system runs at 20 frames per second after a single-frame initialization and maintains real-time performance on QVGA imagery.

These results show that online, error-canceling learning can produce a detector sufficiently accurate and general to re-initialize tracking after prolonged absences or drastic appearance changes, eliminating the need for offline training or manual re-initialization. The approach therefore enables reliable, indefinite tracking of arbitrary objects under realistic conditions.

The main limitations are reduced reliability under full out-of-plane rotation, difficulty with highly articulated objects, and restriction to a single target. Extensions that also adapt the tracker, incorporate background subtraction for static-camera cases, and scale to multiple targets are identified as the next steps needed to broaden applicability.

  • Paper: Incremental Learning for Robust Visual Tracking, David A. Ross et al. (2008). This paper establishes foundational techniques for incremental appearance model updates during visual tracking, providing the adaptive representation strategies built upon by the TLD framework.
  • Paper: Kernel-Based Object Tracking, Dorin Comaniciu et al. (2003). This kernel-based tracking approach provides baseline methodology for real-time localization and scale adaptation that informs the tracker component of the TLD architecture.
  • Paper: High-Speed Tracking with Kernelized Correlation Filters, João F. Henriques et al. (2014). This paper directly builds on and outperforms the TLD framework by introducing circulant-matrix correlation filters that achieve higher accuracy at significantly faster frame rates.
Cover for Tracking-Learning-Detection

Abstract

This paper investigates long-term tracking of unknown objects in a video stream. The object is defined by its location and extent in a single frame. In every frame that follows, the task is to determine the object's location and extent or indicate that the object is not present. We propose a novel tracking framework (TLD) that explicitly decomposes the long-term tracking task into tracking, learning and detection. The tracker follows the object from frame to frame. The detector localizes all appearances that have been observed so far and corrects the tracker if necessary. The learning estimates detector's errors and updates it to avoid these errors in the future. We study how to identify detector's errors and learn from them. We develop a novel learning method (P-N learning) which estimates the errors by a pair ofexperts”: (i) P-expert estimates missed detections, and (ii) N-expert estimates false alarms. The learning process is modeled as a discrete dynamical system and the conditions under which the learning guarantees improvement are found. We describe our real-time implementation of the TLD framework and the P-N learning. We carry out an extensive quantitative evaluation which shows a significant improvement over state-of-the-art approaches.

Table of Contents

  • 1 INTRODUCTION
  • 2 RELATED WORK
  • 2.1 Object tracking
  • 2.2 Object detection
  • 2.3 Machine learning
  • 2.4 Most related approaches
  • 3 TRACKING-LEARNING-DETECTION
  • 4 P-N LEARNING
  • 4.1 Formalization
  • 4.2 Stability
  • 4.3 Experiments with simulated experts
  • 4.4 Design of real experts
  • 5 IMPLEMENTATION OF TLD
  • 5.1 Prerequisites
  • 5.2 Object model
  • 5.3 Object detector
  • 5.3.1 Patch variance
  • 5.3.2 Ensemble classifier
  • 5.3.3 Nearest neighbor classifier
  • 5.4 Tracker
  • 5.5 Integrator
  • 5.6 Learning component
  • 5.6.1 Initialization
  • 5.6.2 P-expert
  • 5.6.3 N-expert
  • 6 QUANTITATIVE EVALUATION
  • 6.1 Comparison 1: CoGD
  • 6.2 Comparison 2: Prost
  • 6.3 TLD dataset
  • 6.4 Improvement of the object detector
  • 6.5 Comparison 3: TLD dataset
  • CONCLUSIONS
  • LIMITATIONS AND FUTURE WORK
  • ACKNOWLEDGMENTS
  • REFERENCES

Knowls

  1. Knowl 1 — Tracking-Learning-Detection (TLD) Framework

    model/method

    The Tracking-Learning-Detection (TLD) framework addresses long-term visual tracking of an unknown target object defined by a single bounding box in an initial frame. TLD decomposes the long-term tracking task into three distinct, simultaneously operating components coordinated by an Integrator:

    1. Tracker: Follows the target from frame to frame under the assumption of limited inter-frame displacement and continuous visibility. It produces smooth trajectories and requires no prior training, but it accumulates drift over time and cannot recover if the object becomes occluded or leaves the field of view.

    2. Detector: Performs an independent full-image scan in every frame to localize all known appearances of the target learned so far. It does not suffer from drift and operates regardless of whether the object was visible in previous frames, but it is prone to false positives and false negatives.

    3. Learner (P-N Learning): Observes the joint operation of both the tracker and detector, identifies their classification errors in real time using structural constraints, and updates the detector's model online. Learning expands the detector's generality to newly discovered target appearances while sharpening its discriminability against background clutter.

    4. Integrator: Fuses the tracker's bounding box and the detector's candidate bounding boxes to produce the final object state estimate for each frame, and uses high-confidence detections to re-initialize the tracker upon tracking failure.

  2. Knowl 2 — Stability Analysis of P-N Learning via Discrete Dynamical System

    theoretical result

    P-N learning is modeled as a discrete dynamical system tracking how classifier errors evolve across bootstrap iterations on an unlabeled set XuX_u. Let α(k)\alpha(k) denote the number of false positives and β(k)\beta(k) denote the number of false negatives produced by the classifier at iteration kk.

    The quality of the P-expert (which labels negative classifications as positive) and N-expert (which labels positive classifications as negative) is parameterized by four quality metrics:

    • P-precision: P+=nc+nc++nf+P^+ = \frac{n_c^+}{n_c^+ + n_f^+} (fraction of true positives among all samples relabeled positive by the P-expert)
    • P-recall: R+=nc+βR^+ = \frac{n_c^+}{\beta} (fraction of classifier false negatives identified by the P-expert)
    • N-precision: P=ncnc+nfP^- = \frac{n_c^-}{n_c^- + n_f^-} (fraction of true negatives among all samples relabeled negative by the N-expert)
    • N-recall: R=ncαR^- = \frac{n_c^-}{\alpha} (fraction of classifier false positives identified by the N-expert)

    where nc+,nf+n_c^+, n_f^+ are the number of correct and false positive relabelings by the P-expert, and nc,nfn_c^-, n_f^- are the correct and false negative relabelings by the N-expert.

    Defining the system error state vector x(k)=[α(k)β(k)]\vec{x}(k) = \begin{bmatrix} \alpha(k) \\ \beta(k) \end{bmatrix}, error propagation is governed by the linear dynamical recurrence:

    x(k+1)=Mx(k)\vec{x}(k+1) = \mathbf{M} \vec{x}(k)

    where the transition matrix M\mathbf{M} is defined as:

    M=[1R1P+P+R+1PPR1R+]\mathbf{M} = \begin{bmatrix} 1 - R^- & \frac{1 - P^+}{P^+} R^+ \\ \frac{1 - P^-}{P^-} R^- & 1 - R^+ \end{bmatrix}

    Convergence Theorem: The error state vector x(k)\vec{x}(k) converges to zero as kk \to \infty if and only if both eigenvalues λ1,λ2\lambda_1, \lambda_2 of the matrix M\mathbf{M} have absolute values strictly less than one (λ1<1|\lambda_1| < 1 and λ2<1|\lambda_2| < 1). Experts satisfying this condition are termed error-canceling because their mutual error compensation stabilizes the online learning process.

    In the symmetric case where P+=R+=P=R=1ϵP^+ = R^+ = P^- = R^- = 1 - \epsilon for error rate ϵ\epsilon, the eigenvalues simplify to λ1=0\lambda_1 = 0 and λ2=2ϵ\lambda_2 = 2\epsilon, which guarantees error decay and stable improvement whenever ϵ<0.5\epsilon < 0.5.

  3. Knowl 3 — Spatio-Temporal P-Expert and N-Expert for Video Streams

    model/method

    P-N learning exploits the spatio-temporal structure inherent in video sequences to identify detector errors without ground-truth supervision:

    • P-Expert (Temporal Continuity): Assumes that the target object moves along a smooth, continuous spatio-temporal trajectory. It tracks the target's location from the previous frame to the current frame using a frame-to-frame tracker. A trajectory is classified as reliable if it enters the core of the object model (where Conservative similarity ScS_c exceeds a set threshold) and remains reliable until tracking failure is flagged or re-initialization occurs. If the detector predicts a negative label at the current reliable location, the P-expert identifies a false negative and outputs positive training patches (the current bounding box plus 10 closest bounding boxes perturbed by affine warps and Gaussian noise), thereby expanding the detector's generalization to new appearances.

    • N-Expert (Spatial Exclusivity): Assumes that the target object can occupy at most one location in any given video frame. When the trajectory is reliable, the N-expert takes the maximally confident location output by the integrator/tracker and labels all distant, non-overlapping candidate bounding boxes (bounding box overlap <0.2< 0.2) as negative training examples. This identifies false alarms and increases the detector's discriminative power against background clutter.

  4. Knowl 4 — Three-Stage Cascaded Object Detector

    model/method

    To evaluate approximately 50,000 scanning windows per frame in real time on standard resolution images (e.g., 320×240320 \times 240), the TLD detector organizes classification into a three-stage cascade where each stage can rapidly reject negative candidates:

    1. Patch Variance Filter: Rejects any candidate patch pp whose gray-level variance E[p2](E[p])2\mathbb{E}[p^2] - (\mathbb{E}[p])^2 is less than 50%50\% of the variance of the initial target patch defined in the first frame. This value is computed in O(1)O(1) time per patch using integral images and typically filters out over 50%50\% of background patches (e.g., uniform regions such as sky and road).

    2. Ensemble Classifier: Evaluates surviving patches using an ensemble of nn randomized base classifiers. Each base classifier ii evaluates d=13d = 13 predefined, pairwise pixel intensity comparisons on a blurred patch to generate a 13-bit binary code x{0,1}13x \in \{0, 1\}^{13}. The code indexes a posterior probability lookup table Pi(y=1x)=#p#p+#nP_i(y=1|x) = \frac{\#p}{\#p + \#n}, where #p\#p and #n\#n are counts of positive and negative patches assigned that code during learning. The ensemble calculates the mean posterior Pˉ(y=1x)=1ni=1nPi(y=1x)\bar{P}(y=1|x) = \frac{1}{n} \sum_{i=1}^n P_i(y=1|x) and passes patches only if Pˉ(y=1x)>0.5\bar{P}(y=1|x) > 0.5.

    3. Nearest Neighbor (NN) Classifier: Evaluates the remaining candidate patches (typically 50\approx 50) using the Relative similarity metric Sr(p,M)S_r(p, M). The patch is classified as the target object if Sr(p,M)>θNNS_r(p, M) > \theta_{\text{NN}} with threshold θNN=0.6\theta_{\text{NN}} = 0.6.

  5. Knowl 5 — TLD Object Model and Similarity Metrics

    definition

    The TLD object model M={p1+,p2+,,pm+,p1,p2,,pn}M = \{p_1^+, p_2^+, \dots, p_m^+, p_1^-, p_2^-, \dots, p_n^-\} is an online collection of historical positive (p+p^+) and negative (pp^-) image patches, each normalized to a canonical resolution of 15×1515 \times 15 pixels. The positive patches are ordered chronologically, where p1+p_1^+ is the initial template and pm+p_m^+ is the most recently added positive patch.

    The similarity between two patches pi,pjp_i, p_j is defined using the Normalized Correlation Coefficient (NCC):

    S(pi,pj)=0.5(NCC(pi,pj)+1)[0,1]S(p_i, p_j) = 0.5 \cdot (\text{NCC}(p_i, p_j) + 1) \in [0, 1]

    For an arbitrary patch pp and model MM, the following similarity measures are defined:

    • Positive Nearest Neighbor Similarity: S+(p,M)=maxpi+MS(p,pi+)S^+(p, M) = \max_{p_i^+ \in M} S(p, p_i^+)
    • Negative Nearest Neighbor Similarity: S(p,M)=maxpiMS(p,pi)S^-(p, M) = \max_{p_i^- \in M} S(p, p_i^-)
    • Conservative Positive Similarity: S50%+(p,M)=maxpi+M,im/2S(p,pi+)S^+_{50\%}(p, M) = \max_{p_i^+ \in M, \, i \le m/2} S(p, p_i^+) (similarity restricted to the oldest 50%50\% of positive patches)
    • Relative Similarity: Sr(p,M)=S+(p,M)S+(p,M)+S(p,M)[0,1]S_r(p, M) = \frac{S^+(p, M)}{S^+(p, M) + S^-(p, M)} \in [0, 1]
    • Conservative Similarity: Sc(p,M)=S50%+(p,M)S50%+(p,M)+S(p,M)[0,1]S_c(p, M) = \frac{S^+_{50\%}(p, M)}{S^+_{50\%}(p, M) + S^-(p, M)} \in [0, 1]

    Model Update Rule: A new labeled patch pp from P-N experts is inserted into MM if the 1-NN prediction differs from the expert label, or if the classification margin satisfies Sr(p,M)θNN<λ|S_r(p, M) - \theta_{\text{NN}}| < \lambda (with λ=0.1\lambda = 0.1 and θNN=0.6\theta_{\text{NN}} = 0.6).

  6. Knowl 6 — Median-Flow Tracker with Failure Detection

    model/method

    The tracking module in TLD tracks target bounding boxes between consecutive frames using a Median-Flow point-tracker augmented with an explicit failure detection rule:

    1. Point Tracking: A regular grid of 10×1010 \times 10 points is placed within the bounding box. Displacements did_i for each point between frame tt and frame t+1t+1 are estimated using a two-level pyramidal Lucas-Kanade optical flow tracker on 10×1010 \times 10 pixel patches.

    2. Median Motion Filtering: The overall bounding box displacement is computed by taking the median over the 50%50\% most reliable point displacement vectors.

    3. Failure Detection: Let dmd_m denote the median displacement vector. The motion residual for each tracked point is defined as didm|d_i - d_m|. A tracking failure is declared if:

    medianididm>10 pixels\text{median}_i |d_i - d_m| > 10 \text{ pixels}

    When this threshold is exceeded (typically caused by rapid motion, background occlusion, or target disappearance), the tracker outputs no bounding box, signaling the framework that tracker output is invalid and preventing drift.

  7. Knowl 7 — Integrator Decision Rule

    model/method

    The Integrator merges the current bounding box hypotheses produced by the frame-to-frame tracker and the scanning-window detector into a unified output:

    1. If neither the tracker nor the detector outputs a bounding box, the object is declared not visible.
    2. If bounding boxes are produced by either or both components, the Integrator outputs the single bounding box that achieves the maximal Conservative similarity Sc(p,M)S_c(p, M).
    3. The detector's highest-confidence detection is used to re-initialize the tracker state when the tracker has suffered a detected failure or drift.
  8. Knowl 8 — Detector Improvement and P-N Expert Metrics Across TLD Sequences

    data/table

    The quantitative evaluation of P-N learning demonstrates that retraining an initial detector (trained on frame 1) over a video stream via P-N experts increases detector recall substantially while preserving high precision. On the challenging sequences 7–10, where the initial detector fails completely (F-measure 0.00–0.01), P-N learning increases F-measure to 0.25–0.83. All transition matrix eigenvalues λ1,λ2\lambda_1, \lambda_2 remain strictly below 1.0, confirming the empirical validity of the dynamical system stability criterion.

    Sequence Frames Initial Detector Final Detector P-expert N-expert Eigenvalues
    Precision / Recall / F-m Precision / Recall / F-m P+,R+P^+, R^+ P,RP^-, R^- λ1,λ2\lambda_1, \lambda_2
    1. David 761 1.00 / 0.01 / 0.02 1.00 / 0.32 / 0.49 1.00 / 0.08 0.99 / 0.17 0.92 / 0.83
    2. Jumping 313 1.00 / 0.01 / 0.02 0.99 / 0.88 / 0.93 0.86 / 0.24 0.98 / 0.30 0.70 / 0.77
    3. Pedestrian 1 140 1.00 / 0.06 / 0.12 1.00 / 0.12 / 0.22 0.81 / 0.04 1.00 / 0.04 0.96 / 0.96
    4. Pedestrian 2 338 1.00 / 0.02 / 0.03 1.00 / 0.34 / 0.51 1.00 / 0.25 1.00 / 0.24 0.76 / 0.75
    5. Pedestrian 3 184 1.00 / 0.73 / 0.84 0.97 / 0.93 / 0.95 0.98 / 0.78 0.98 / 0.68 0.32 / 0.22
    6. Car 945 1.00 / 0.04 / 0.08 0.99 / 0.82 / 0.90 1.00 / 0.52 1.00 / 0.46 0.48 / 0.54
    7. Motocross 2665 1.00 / 0.00 / 0.00 0.92 / 0.32 / 0.47 0.96 / 0.19 0.84 / 0.08 0.92 / 0.81
    8. Volkswagen 8576 1.00 / 0.00 / 0.00 0.92 / 0.75 / 0.83 0.70 / 0.23 0.99 / 0.09 0.91 / 0.77
    9. Car Chase 9928 0.36 / 0.00 / 0.00 0.90 / 0.42 / 0.57 0.64 / 0.19 0.95 / 0.22 0.76 / 0.83
    10. Panda 3000 0.79 / 0.01 / 0.01 0.51 / 0.16 / 0.25 0.31 / 0.02 0.96 / 0.19 0.81 / 0.99
  9. Knowl 9 — Tracking Performance Benchmark Comparison on TLD Dataset

    data/table

    TLD was evaluated against five baseline tracking algorithms across 10 challenging video sequences (totalling 26,850 frames) containing full occlusions, pose variations, illumination changes, and camera motion: Online Boosting (OB), Semi-Supervised Online Boosting (SB), Beyond Semi-Supervised Tracking (BS), Multiple Instance Learning (MIL), and Co-trained Generative-Discriminative Tracker (CoGD). Trajectories were evaluated using Precision (P), Recall (R), and F-measure (F) with a bounding box overlap threshold >25%> 25\%.

    TLD achieved the highest overall performance on 9 out of 10 sequences, with a weighted mean F-measure of 0.81 compared to 0.22 for the second-best algorithm (CoGD) and 0.13–0.15 for the others, while running in real time at 20 frames per second.

    Sequence Frames OB SB BS MIL CoGD TLD
    1. David 761 0.41/0.29/0.34 0.35/0.35/0.35 0.32/0.24/0.28 0.15/0.15/0.15 1.00/1.00/1.00 1.00/1.00/1.00
    2. Jumping 313 0.47/0.05/0.09 0.25/0.13/0.17 0.17/0.14/0.15 1.00/1.00/1.00 1.00/0.99/1.00 1.00/1.00/1.00
    3. Pedestrian 1 140 0.61/0.14/0.23 0.48/0.33/0.39 0.29/0.10/0.15 0.69/0.69/0.69 1.00/1.00/1.00 1.00/1.00/1.00
    4. Pedestrian 2 338 0.77/0.12/0.21 0.85/0.71/0.77 1.00/0.02/0.04 0.10/0.12/0.11 0.72/0.92/0.81 0.89/0.92/0.91
    5. Pedestrian 3 184 1.00/0.33/0.49 0.41/0.33/0.36 0.92/0.46/0.62 0.69/0.81/0.75 0.85/1.00/0.92 0.99/1.00/0.99
    6. Car 945 0.94/0.59/0.73 1.00/0.67/0.80 0.99/0.56/0.72 0.23/0.25/0.24 0.95/0.96/0.96 0.92/0.97/0.94
    7. Motocross 2665 0.33/0.00/0.01 0.13/0.03/0.05 0.14/0.00/0.00 0.05/0.02/0.03 0.93/0.30/0.45 0.89/0.77/0.83
    8. Volkswagen 8576 0.39/0.02/0.04 0.04/0.04/0.04 0.02/0.01/0.01 0.42/0.04/0.07 0.79/0.06/0.11 0.80/0.96/0.87
    9. Carchase 9928 0.79/0.03/0.06 0.80/0.04/0.09 0.52/0.12/0.19 0.62/0.04/0.07 0.95/0.04/0.08 0.86/0.70/0.77
    10. Panda 3000 0.95/0.35/0.51 1.00/0.17/0.29 0.99/0.17/0.30 0.36/0.40/0.38 0.12/0.12/0.12 0.58/0.63/0.60
    Mean 26850 0.62/0.09/0.13 0.50/0.10/0.14 0.39/0.10/0.15 0.44/0.11/0.13 0.80/0.18/0.22 0.82/0.81/0.81
  10. Knowl 10 — Limitations of the TLD Framework

    limitation

    The TLD system exhibits several specific limitations documented by the authors:

    1. Out-of-Plane Rotations: Under full out-of-plane rotation, the Median-Flow tracker drifts away and fails, and the detector cannot reacquire the object until it returns to an appearance or pose that was previously observed and trained into the model.

    2. Fixed Tracker Component: Online learning updates only the object detector while the frame-to-frame tracker remains static, causing the tracker to make the same recurring tracking errors over time.

    3. Single Target Constraint: TLD tracks a single object instance and lacks mechanisms for joint model training or feature sharing across multiple simultaneous targets.

    4. Articulated Objects: Performance degrades on highly deformable or articulated non-rigid targets (e.g., pedestrians) relative to rigid and semi-rigid objects.

Coverage note — Preliminary benchmark evaluations on older short datasets (Tables 1, 2, and 3 comparing with IVT, ODF, ET, ORF, FT, and PROST) were omitted because they showed saturated performance on short sequences and are fully superseded by the comprehensive evaluation on the 10-sequence TLD dataset.

References

  1. 1.A. Blum and T. Mitchell, “Combining labeled and unlabeled data with co-training,” Conference on Computational Learning Theory, p. 100, 1998.
  2. 2.B. D. Lucas and T. Kanade, “An iterative image registration technique with an application to stereo vision,” International Joint Conference on Artificial Intelligence, vol. 81, pp. 674–679, 1981.
  3. 3.J. Shi and C. Tomasi, “Good features to track,” Conference on Computer Vision and Pattern Recognition, 1994.
  4. 4.P. Sand and S. Teller, “Particle video: Long-range motion estimation using point trajectories,” International Journal of Computer Vision, vol. 80, no. 1, pp. 72–91, 2008.
  5. 5.L. Wang, W. Hu, and T. Tan, “Recent developments in human motion analysis,” Pattern Recognition, vol. 36, no. 3, pp. 585–601, 2003.
  6. 6.D. Ramanan, D. A. Forsyth, and A. Zisserman, “Tracking people by learning their appearance,” IEEE Transactions on Pattern Analysis and Machine Intelligence, pp. 65–81, 2007.
  7. 7.P. Buehler, M. Everingham, D. P. Huttenlocher, and A. Zisserman, “Long term arm and hand tracking for continuous sign language TV broadcasts,” British Machine Vision Conference, 2008.
  8. 8.S. Birchfield, “Elliptical head tracking using intensity gradients and color histograms,” Conference on Computer Vision and Pattern Recognition, 1998.
  9. 9.M. Isard and A. Blake, “CONDENSATION - Conditional Density Propagation for Visual Tracking,” International Journal of Computer Vision, vol. 29, no. 1, pp. 5–28, 1998.
  10. 10.C. Bibby and I. Reid, “Robust real-time visual tracking using pixel-wise posteriors,” European Conference on Computer Vision, 2008.
  11. 11.C. Bibby and I. Reid, “Real-time Tracking of Multiple Occluding Objects using Level Sets,” Computer Vision and Pattern Recognition, 2010.
  12. 12.B. K. P. Horn and B. G. Schunck, “Determining optical flow,” Artificial intelligence, vol. 17, no. 1-3, pp. 185–203, 1981.
  13. 13.T. Brox, A. Bruhn, N. Papenberg, and J. Weickert, “High accuracy optical flow estimation based on a theory for warping,” European Conference on Computer Vision, pp. 25–36, 2004.
  14. 14.J. L. Barron, D. J. Fleet, and S. S. Beauchemin, “Performance of Optical Flow Techniques,” International Journal of Computer Vision, vol. 12, no. 1, pp. 43–77, 1994.
  15. 15.D. Comaniciu, V. Ramesh, and P. Meer, “Kernel-Based Object Tracking,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 25, no. 5, pp. 564–577, 2003.
  16. 16.I. Matthews, T. Ishikawa, and S. Baker, “The Template Update Problem,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 26, no. 6, pp. 810–815, 2004.
  17. 17.N. Dowson and R. Bowden, “Simultaneous Modeling and Tracking (SMAT) of Feature Sets,” Conference on Computer Vision and Pattern Recognition, 2005.
  18. 18.A. Rahimi, L. P. Morency, and T. Darrell, “Reducing drift in differential tracking,” Computer Vision and Image Understanding, vol. 109, no. 2, pp. 97–111, 2008.
  19. 19.A. D. Jepson, D. J. Fleet, and T. F. El-Maraghi, “Robust Online Appearance Models for Visual Tracking,” IEEE Transactions on Pattern Analysis and Machine Intelligence, pp. 1296–1311, 2003.
  20. 20.A. Adam, E. Rivlin, and I. Shimshoni, “Robust Fragments-based Tracking using the Integral Histogram,” Conference on Computer Vision and Pattern Recognition, pp. 798–805, 2006.
  21. 21.M. J. Black and A. D. Jepson, “Eigentracking: Robust matching and tracking of articulated objects using a view-based representation,” International Journal of Computer Vision, vol. 26, no. 1, pp. 63–84, 1998.
  22. 22.D. Ross, J. Lim, R. Lin, and M. Yang, “Incremental Learning for Robust Visual Tracking,” International Journal of Computer Vision, vol. 77, pp. 125–141, Aug. 2007.
  23. 23.J. Kwon and K. M. Lee, “Visual Tracking Decomposition,” Conference on Computer Vision and Pattern Recognition, 2010.
  24. 24.M. Yang, Y. Wu, and G. Hua, “Context-aware visual tracking.,” IEEE transactions on pattern analysis and machine intelligence, vol. 31, pp. 1195–209, July 2009.
  25. 25.H. Grabner, J. Matas, L. Van Gool, and P. Cattin, “Tracking the Invisible: Learning Where the Object Might be,” Conference on Computer Vision and Pattern Recognition, 2010.
  26. 26.S. Avidan, “Support Vector Tracking,” IEEE Transactions on Pattern Analysis and Machine Intelligence, pp. 1064–1072, 2004.
  27. 27.R. Collins, Y. Liu, and M. Leordeanu, “Online Selection of Discriminative Tracking Features,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 27, no. 10, pp. 1631–1643, 2005.
  28. 28.S. Avidan, “Ensemble Tracking,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 29, no. 2, pp. 261–271, 2007.
  29. 29.H. Grabner and H. Bischof, “On-line boosting and vision,” Conference on Computer Vision and Pattern Recognition, 2006.
  30. 30.B. Babenko, M.-H. Yang, and S. Belongie, “Visual Tracking with Online Multiple Instance Learning,” Conference on Computer Vision and Pattern Recognition, 2009.
  31. 31.H. Grabner, C. Leistner, and H. Bischof, “Semi-Supervised On-line Boosting for Robust Tracking,” European Conference on Computer Vision, 2008.
  32. 32.F. Tang, S. Brennan, Q. Zhao, H. Tao, and U. C. Santa Cruz, “Co-tracking using semi-supervised support vector machines,” International Conference on Computer Vision, pp. 1–8, 2007.
  33. 33.Q. Yu, T. B. Dinh, and G. Medioni, “Online tracking and reacquisition using co-trained generative and discriminative trackers,” European Conference on Computer Vision, 2008.
  34. 34.D. G. Lowe, “Distinctive image features from scale-invariant keypoints,” International Journal of Computer Vision, vol. 60, no. 2, pp. 91–110, 2004.
  35. 35.P. Viola and M. Jones, “Rapid object detection using a boosted cascade of simple features,” Conference on Computer Vision and Pattern Recognition, 2001.
  36. 36.V. Lepetit, P. Lagger, and P. Fua, “Randomized trees for real-time keypoint recognition,” Conference on Computer Vision and Pattern Recognition, 2005.
  37. 37.L. Vacchetti, V. Lepetit, and P. Fua, “Stable real-time 3d tracking using online and offline information,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 26, no. 10, p. 1385, 2004.
  38. 38.S. Taylor and T. Drummond, “Multiple target localisation at over 100 fps,” British Machine Vision Conference, 2009.
  39. 39.J. Pilet and H. Saito, “Virtually augmenting hundreds of real pictures: An approach based on learning, retrieval, and tracking,” 2010 IEEE Virtual Reality Conference (VR), pp. 71–78, Mar. 2010.
  40. 40.S. Obdrzalek and J. Matas, “Sub-linear indexing for large scale object recognition,” British Machine Vision Conference, vol. 1, pp. 1–10, 2005.
  41. 41.S. Hinterstoisser, O. Kutter, N. Navab, P. Fua, and V. Lepetit, “Real-time learning of accurate patch rectification,” Conference on Computer Vision and Pattern Recognition, 2009.
  42. 42.O. Chapelle, B. Scholkopf, and A. Zien, Semi-Supervised Learning. Cambridge, MA: MIT Press, 2006.
  43. 43.X. Zhu and A. B. Goldberg, Introduction to semi-supervised learning. Morgan & Claypool Publishers, 2009.
  44. 44.K. Nigam, A. K. McCallum, S. Thrun, and T. Mitchell, “Text classification from labeled and unlabeled documents using EM,” Machine Learning, vol. 39, no. 2, pp. 103–134, 2000.
  45. 45.R. Fergus, P. Perona, and A. Zisserman, “Object class recognition by unsupervised scale-invariant learning,” Conference on Computer Vision and Pattern Recognition, vol. 2, 2003.
  46. 46.C. Rosenberg, M. Hebert, and H. Schneiderman, “Semi-supervised self-training of object detection models,” Workshop on Application of Computer Vision, 2005.
  47. 47.N. Poh, R. Wong, J. Kittler, and F. Roli, “Challenges and Research Directions for Adaptive Biometric Recognition Systems,” Advances in Biometrics, 2009.
  48. 48.A. Levin, P. Viola, and Y. Freund, “Unsupervised improvement of visual detectors using co-training,” International Conference on Computer Vision, 2003.
  49. 49.O. Javed, S. Ali, and M. Shah, “Online detection and classification of moving objects using progressively improving detectors,” Conference on Computer Vision and Pattern Recognition, 2005.
  50. 50.O. Williams, A. Blake, and R. Cipolla, “Sparse bayesian learning for efficient visual tracking,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 27, no. 8, pp. 1292–1304, 2005.
  51. 51.M. Isard and A. Blake, “CONDENSATION Conditional Density Propagation for Visual Tracking,” International Journal of Computer Vision, vol. 29, no. 1, pp. 5–28, 1998.
  52. 52.Y. Li, H. Ai, T. Yamashita, S. Lao, and M. Kawade, “Tracking in Low Frame Rate Video: A Cascade Particle Filter with Discriminative Observers of Different Lifespans,” Conference on Computer Vision and Pattern Recognition, 2007.
  53. 53.K. Okuma, A. Taleghani, N. de Freitas, J. J. Little, and D. G. Lowe, “A boosted particle filter: Multitarget detection and tracking,” European Conference on Computer Vision, 2004.
  54. 54.B. Leibe, K. Schindler, and L. Van Gool, “Coupled Detection and Trajectory Estimation for Multi-Object Tracking,” 2007 IEEE 11th International Conference on Computer Vision, pp. 1–8, Oct. 2007.
  55. 55.M. D. Breitenstein, F. Reichlin, B. Leibe, E. Koller-Meier, and L. V. Gool, “Robust Tracking-by-Detection using a Detector Confidence Particle Filter,” International Conference on Computer Vision, 2009.
  56. 56.K. K. Sung and T. Poggio, “Example-based learning for view-based human face detection,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 20, no. 1, pp. 39–51, 1998.
  57. 57.K. Zhou, J. C. Doyle, and K. Glover, Robust and optimal control. Prentice Hall Englewood Cliffs, NJ, 1996.
  58. 58.K. Ogata, Modern control engineering. Prentice Hall, 2009.
  59. 59.Z. Kalal, J. Matas, and K. Mikolajczyk, “P-N Learning: Bootstrapping Binary Classifiers by Structural Constraints,” Conference on Computer Vision and Pattern Recognition, 2010.
  60. 60.V. Lepetit and P. Fua, “Keypoint recognition using randomized trees.,” IEEE transactions on pattern analysis and machine intelligence, vol. 28, pp. 1465–79, Sept. 2006.
  61. 61.M. Ozuysal, P. Fua, and V. Lepetit, “Fast Keypoint Recognition in Ten Lines of Code,” Conference on Computer Vision and Pattern Recognition, 2007.
  62. 62.M. Calonder, V. Lepetit, and P. Fua, “BRIEF : Binary Robust Independent Elementary Features,” European Conference on Computer Vision, 2010.
  63. 63.L. Breiman, “Random forests,” Machine Learning, vol. 45, no. 1, pp. 5–32, 2001.
  64. 64.Z. Kalal, K. Mikolajczyk, and J. Matas, “Forward-Backward Error: Automatic Detection of Tracking Failures,” International Conference on Pattern Recognition, pp. 23–26, 2010.
  65. 65.J. Y. Bouguet, “Pyramidal Implementation of the Lucas Kanade Feature Tracker Description of the algorithm,” Technical Report, Intel Microprocessor Research Labs, 1999.
  66. 66.Z. Kalal, J. Matas, and K. Mikolajczyk, “Online learning of robust object detectors during unstable tracking,” On-line Learning for Computer Vision Workshop, 2009.
  67. 67.J. Santner, C. Leistner, A. Saffari, T. Pock, and H. Bischof, “PROST: Parallel Robust Online Simple Tracking,” Conference on Computer Vision and Pattern Recognition, 2010.
  68. 68.A. Saffari, C. Leistner, J. Santner, M. Godec, and H. Bischof, “On-line Random Forests,” Online Learning for Computer Vision Workshop, 2009.
  69. 69.S. Stalder, H. Grabner, and L. V. Gool, “Beyond semi-supervised tracking: Tracking should be as simple as detection, but not simpler than recognition,” 2009 IEEE 12th International Conference on Computer Vision Workshops, ICCV Workshops, pp. 1409–1416, Sept. 2009.

Citation

MLA
Kalal, Z., et al. “Tracking-Learning-Detection”. IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 34, no. 7, 2012, pp. 1409–22, https://doi.org/10.1109/TPAMI.2011.239.
APA
Kalal, Z., Mikolajczyk, K., & Matas, J. (2012). Tracking-Learning-Detection. IEEE Transactions on Pattern Analysis and Machine Intelligence, 34(7), 1409–1422. https://doi.org/10.1109/TPAMI.2011.239
Chicago
Kalal, Z., K. Mikolajczyk, and J. Matas. 2012. “Tracking-Learning-Detection”. IEEE Transactions on Pattern Analysis and Machine Intelligence 34 (7): 1409–22. https://doi.org/10.1109/TPAMI.2011.239.
Harvard
Kalal, Z., Mikolajczyk, K. and Matas, J. (2012) “Tracking-Learning-Detection”, IEEE Transactions on Pattern Analysis and Machine Intelligence, 34(7), pp. 1409–1422. Available at: https://doi.org/10.1109/TPAMI.2011.239.
Vancouver
1. Kalal Z, Mikolajczyk K, Matas J (2012) Tracking-Learning-Detection. IEEE Transactions on Pattern Analysis and Machine Intelligence 34:1409–1422

BibTeX

@article{Kalal_2012, title={Tracking-Learning-Detection}, volume={34}, ISSN={2160-9292}, url={http://dx.doi.org/10.1109/TPAMI.2011.239}, DOI={10.1109/tpami.2011.239}, number={7}, journal={IEEE Transactions on Pattern Analysis and Machine Intelligence}, publisher={Institute of Electrical and Electronics Engineers (IEEE)}, author={Kalal, Zdenek and Mikolajczyk, Krystian and Matas, Jiri}, year={2012}, month=July, pages={1409–1422} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF