Fast Online Object Tracking and Segmentation: A Unifying Approach

Qiang WangLi ZhangLuca BertinettoWeiming HuPhilip H. S. Torr

article2018CVPR1,338 citations

Proposes SiamMask, a unified framework that extends fully-convolutional Siamese trackers with a binary segmentation branch to achieve state-of-the-art real-time object tracking and video mask generation at 55 frames per second from a single bounding box initialization.

Listen

Visual reasoning in real-time video streaming is essential for modern applications such as autonomous navigation, automated surveillance, and video editing. However, computer vision has traditionally separated this challenge into two distinct tasks with conflicting trade-offs: visual object tracking, which is fast and operates online but produces coarse rectangular bounding boxes, and video object segmentation, which generates detailed pixel-level masks but is computationally slow and requires complex initialization.

The article demonstrates a unified framework named SiamMask that bridges this divide by performing both online visual tracking and semi-supervised video object segmentation in real-time using only a simple bounding box initialization.

The authors implemented a multi-task learning architecture based on fully-convolutional Siamese neural networks. By extending standard similarity matching and bounding box regression with an additional binary segmentation branch, the network learns to predict spatial masks directly during offline training. The model was trained offline across large-scale video and image datasets (COCO, ImageNet-VID, and YouTube-VOS) and evaluated across major benchmark datasets, including VOT-2016, VOT-2018, DAVIS-2016, DAVIS-2017, and YouTube-VOS, without requiring any online retraining or sequence-specific adaptation.

The article establishes several key findings. First, SiamMask achieves state-of-the-art performance among real-time visual trackers on the VOT-2018 benchmark, securing an Expected Average Overlap score of 0.380 while processing 55 frames per second on a single graphics processing unit. Second, it demonstrates competitive accuracy against dedicated video object segmentation methods while operating four to sixty times faster than existing competitive approaches. Third, deriving rotated minimum bounding rectangles from the predicted pixel masks improves tracking accuracy significantly over traditional axis-aligned boxes, yielding a 10.6% improvement in mean intersection-over-union. Finally, multi-task training benefits overall tracking performance even when mask outputs are not explicitly used at test time.

These results show that high-precision, pixel-level object tracking no longer requires heavy computational overhead, extensive manual annotations, or slow per-video fine-tuning. Integrating mask generation with fast tracking reduces infrastructure costs and latency risks, making fine-grained visual tracking practical for latency-critical and embedded real-world systems.

For practical deployment, organizations should adopt SiamMask with the minimum bounding rectangle strategy when real-time throughput (55 frames per second) is required, reserving the slower optimization-based bounding box strategy for non-real-time workflows prioritizing maximum overlap. Future development should focus on enhancing training data diversity to improve robustness against severe motion blur and indistinct object patterns, which remain the primary causes of tracking failure.

Cover for Fast Online Object Tracking and Segmentation: A Unifying Approach

Abstract

In this paper we illustrate how to perform both visual object tracking and semi-supervised video object segmentation, in real-time, with a single simple approach. Our method, dubbed SiamMask, improves the offline training procedure of popular fully-convolutional Siamese approaches for object tracking by augmenting their loss with a binary segmentation task. Once trained, SiamMask solely relies on a single bounding box initialisation and operates online, producing class-agnostic object segmentation masks and rotated bounding boxes at 55 frames per second. Despite its simplicity, versatility and fast speed, our strategy allows us to establish a new state of the art among real-time trackers on VOT-2018, while at the same time demonstrating competitive performance and the best speed for the semi-supervised video object segmentation task on DAVIS-2016 and DAVIS-2017. The project website is this http URL.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Methodology
  • 3.1 Fully-convolutional Siamese networks
  • 3.2 SiamMask
  • 3.3 Implementation details
  • 4 Experiments
  • 4.1 Evaluation for visual object tracking
  • 4.2 Evaluation for semi-supervised VOS
  • 4.3 Further analysis
  • 5 Conclusion
  • References
  • A Architectural details
  • B Further qualitative results

Knowls

  1. Knowl 1 — SiamMask Multi-Task Architecture

    model/method

    SiamMask extends fully-convolutional Siamese tracking frameworks to simultaneously perform visual object tracking and class-agnostic binary object segmentation from a single bounding box initialization.

    The framework processes two inputs through a shared convolutional backbone fθf_\theta: a target exemplar crop zz of size 127×127×3127 \times 127 \times 3 and a search region crop xx of size 255×255×3255 \times 255 \times 3. A depth-wise cross-correlation (denoted ⋆d\star_d) of their feature representations yields a multi-channel response map of size 17×17×25617 \times 17 \times 256:

    gθ(z,x)=fθ(z)⋆dfθ(x)g_\theta(z, x) = f_\theta(z) \star_d f_\theta(x)

    Each spatial position in this response map is called a Response of a candidate Window (RoW), denoted gθn(z,x)g_\theta^n(z, x) for the nn-th candidate window.

    SiamMask branches into separate heads departing from the shared representation:

    1. Mask branch (hϕh_\phi): A two-layer network consisting of 1×11 \times 1 convolutions (with 256 and 632=396963^2 = 3969 channels) that predicts a flattened w×hw \times h (63×6363 \times 63) binary segmentation mask mn=hϕ(gθn(z,x))m_n = h_\phi(g_\theta^n(z, x)) for each RoW. Information across the RoW is merged with high-resolution features from earlier backbone layers using stacked top-down refinement modules with upsampling and skip connections.

    2. Classification score branch: Predicts positive/negative target alignment scores for each RoW.

    3. Bounding box regression branch (optional): Predicts bounding box proposal offsets via Region Proposal Network (RPN) heads parameterized with anchor boxes.

    Two principal variants are formulated: SiamMask-2B (a two-branch model augmenting SiamFC with mask prediction and similarity classification) and SiamMask (a three-branch model augmenting SiamRPN with mask prediction, classification scores, and bounding box regression).

  2. Knowl 2 — SiamMask Multi-Task Loss Formulations

    equation

    SiamMask trains candidate window mask prediction jointly with classification and localization tasks using multi-task objectives.

    For a target mask of size w×hw \times h (w=h=63w = h = 63), let yn∈{+1,−1}y_n \in \{+1, -1\} denote the positive or negative ground-truth label of the nn-th candidate Response of a candidate Window (RoW), and let cnij∈{+1,−1}c_n^{ij} \in \{+1, -1\} denote the ground-truth binary segmentation label at pixel (i,j)(i, j) in the candidate window. The mask loss Lmask(θ,ϕ)\mathcal{L}_{\text{mask}}(\theta, \phi) over network backbone parameters θ\theta and mask head parameters ϕ\phi is a binary logistic regression loss computed solely over positive candidate windows (yn=1y_n = 1):

    Lmask(θ,ϕ)=∑n(1+yn2wh∑i=1w∑j=1hlog⁡(1+e−cnijmnij))\mathcal{L}_{\text{mask}}(\theta, \phi) = \sum_{n} \left( \frac{1 + y_n}{2wh} \sum_{i=1}^w \sum_{j=1}^h \log\left(1 + e^{-c_n^{ij} m_n^{ij}}\right) \right)

    where mnijm_n^{ij} is the predicted logit for pixel (i,j)(i, j) of the nn-th RoW.

    For the two-branch architecture (SiamMask-2B), the total objective balances mask prediction with Siamese similarity loss Lsim\mathcal{L}_{\text{sim}}:

    L2B=λ1Lmask+λ2Lsim\mathcal{L}_{\text{2B}} = \lambda_1 \mathcal{L}_{\text{mask}} + \lambda_2 \mathcal{L}_{\text{sim}}

    For the three-branch architecture (SiamMask), the total objective combines mask prediction, cross-entropy classification score loss Lscore\mathcal{L}_{\text{score}}, and smooth L1L_1 box regression loss Lbox\mathcal{L}_{\text{box}}:

    L3B=λ1Lmask+λ2Lscore+λ3Lbox\mathcal{L}_{\text{3B}} = \lambda_1 \mathcal{L}_{\text{mask}} + \lambda_2 \mathcal{L}_{\text{score}} + \lambda_3 \mathcal{L}_{\text{box}}

    Hyperparameter weights are set to λ1=32\lambda_1 = 32 and λ2=λ3=1\lambda_2 = \lambda_3 = 1. In L3B\mathcal{L}_{\text{3B}}, a RoW is assigned yn=1y_n = 1 if at least one anchor box achieves an Intersection over Union (IoU) ≥0.6\ge 0.6 with the ground-truth bounding box, and yn=−1y_n = -1 otherwise.

  3. Knowl 3 — Online Tracking and Video Object Segmentation Inference Algorithm

    algorithm

    SiamMask executes online, causal inference at test time without sequence-specific fine-tuning or online network updates.

    Input: Video frames (I1,I2,…,IT)(I_1, I_2, \dots, I_T), initial bounding box B1B_1 in I1I_1
    Output: Estimated target segmentation masks (M2,…,MT)(M_2, \dots, M_T) and target bounding boxes (B2,…,BT)(B_2, \dots, B_T)
    Extract exemplar image patch zz of size 127×127127 \times 127 centered on B1B_1 in I1I_1
    Compute exemplar feature map fθ(z)f_\theta(z)
    Set current target center and scale from B1B_1
    for t=2t = 2 to TT do
        Extract search image patch xtx_t of size 255×255255 \times 255 from ItI_t centered at previous target position
        Compute search feature map fθ(xt)f_\theta(x_t)
        Compute cross-correlated response map gθ(z,xt)=fθ(z)⋆dfθ(xt)g_\theta(z, x_t) = f_\theta(z) \star_d f_\theta(x_t)
        Compute classification score map across all RoWs
        Find the RoW index n∗n^* attaining the maximum classification score
        Evaluate mask head mn∗=hϕ(gθn∗(z,xt))m_{n^*} = h_\phi(g_\theta^{n^*}(z, x_t))
        Apply pixel-wise sigmoid and threshold at 0.5 to generate binary mask MtM_t
        if using two-branch variant (SiamMask-2B) then
            Fit axis-aligned enclosing rectangle (Min-max) to MtM_t to obtain box reference BtB_t
        else if using three-branch variant (SiamMask) then
            Extract box prediction BtB_t from highest-scoring anchor output of box regression branch
            Alternatively, compute rotated Minimum Bounding Rectangle (MBR) from MtM_t as BtB_t
        end if
        Update target center location and scale for frame It+1I_{t+1} using BtB_t
    end for
  4. Knowl 4 — Bounding Box Derivation Strategies from Binary Masks

    model/method

    To evaluate binary segmentation models on visual tracking benchmarks requiring bounding box outputs, SiamMask analyzes three distinct box generation strategies from the predicted pixel-level mask:

    1. Min-max (Axis-aligned rectangle): Fits an axis-aligned box defined by the minimum and maximum horizontal and vertical pixel coordinates of the segmented mask.
    2. MBR (Minimum Bounding Rectangle): Computes the minimum area rotated bounding rectangle that fully encloses the binary mask.
    3. Opt (Optimization strategy): Uses the iterative polygon optimization procedure from the VOT-2016 benchmark to compute a rotated bounding box maximizing overlap with the segmented mask.

    Generating rotated bounding boxes via MBR or Opt significantly enhances tracking accuracy over axis-aligned boxes by eliminating background pixels enclosed within tilted or non-rigid targets.

  5. Knowl 5 — Performance of SiamMask on the VOT-2018 Tracking Benchmark

    data/table

    SiamMask sets a new state-of-the-art for real-time trackers on the VOT-2018 benchmark, evaluated under Expected Average Overlap (EAO), Accuracy, Robustness (failure rate), and frame rate in frames per second (fps).

    Metric SiamMask-Opt SiamMask SiamMask-2B DaSiamRPN SiamRPN SA_Siam_R CSRDCF STRCF
    EAO ↑\uparrow 0.387 0.380 0.334 0.326 0.244 0.337 0.263 0.345
    Accuracy ↑\uparrow 0.642 0.609 0.575 0.569 0.490 0.566 0.466 0.523
    Robustness ↓\downarrow 0.295 0.276 0.304 0.337 0.460 0.258 0.318 0.215
    Speed (fps) ↑\uparrow 5 55 60 160 200 32.4 48.9 2.9

    The three-branch SiamMask with MBR box generation achieves an EAO of 0.380 at 55 fps, outperforming DaSiamRPN (0.326 EAO). The two-branch SiamMask-2B achieves 0.334 EAO at 60 fps without bounding box regression branches. SiamMask-Opt achieves the highest EAO of 0.387, but its speed drops to 5 fps due to the iterative bounding box optimization.

  6. Knowl 6 — Representation Accuracy and Theoretical Bounds on VOT-2016

    data/table

    An experiment on randomly cropped search patches from VOT-2016 assesses the representation capability of masks versus bounding boxes, alongside theoretical oracle upper bounds computed using ground-truth object boundaries.

    Method mIOU (%) [email protected] IOU [email protected] IOU
    Fixed a.r. Oracle 73.43 90.15 62.52
    Min-max Oracle 77.70 88.84 65.16
    MBR Oracle 84.07 97.77 80.68
    SiamFC 50.48 56.42 9.28
    SiamRPN 60.02 76.20 32.47
    SiamMask-Min-max 65.05 82.99 43.09
    SiamMask-MBR 67.15 85.42 50.86
    SiamMask-Opt 71.68 90.77 60.47

    The MBR Oracle improves mean IoU by 10.64% over the Fixed aspect-ratio Oracle (84.07% vs 73.43%). SiamMask-MBR achieves 67.15% mIOU and 50.86% [email protected], outperforming SiamRPN by +7.13% mIOU and +18.39% [email protected].

  7. Knowl 7 — Semi-Supervised Video Object Segmentation on DAVIS and YouTube-VOS Benchmarks

    data/table

    SiamMask performs competitive semi-supervised video object segmentation (VOS) online at 55 fps, requiring only a bounding box initialization extracted from the first-frame mask (without online fine-tuning FT\text{FT} or mask input M\text{M}).

    DAVIS-2016 FT M JM↑\mathcal{J}_M \uparrow JO↑\mathcal{J}_O \uparrow JD↓\mathcal{J}_D \downarrow FM↑\mathcal{F}_M \uparrow FO↑\mathcal{F}_O \uparrow FD↓\mathcal{F}_D \downarrow
    OnAVOS ✓ ✓ 86.1 96.1 5.2 84.9 89.7 5.8
    OSMN ×\times ✓ 74.0 87.6 9.0 72.9 84.0 10.6
    RGMP ×\times ✓ 81.5 91.7 10.9 82.0 90.8 10.1
    SiamMask ×\times ×\times 71.7 86.8 3.0 67.8 79.8 2.1
    DAVIS-2017 FT M JM↑\mathcal{J}_M \uparrow JO↑\mathcal{J}_O \uparrow JD↓\mathcal{J}_D \downarrow FM↑\mathcal{F}_M \uparrow FO↑\mathcal{F}_O \uparrow FD↓\mathcal{F}_D \downarrow
    OnAVOS ✓ ✓ 61.6 67.4 27.9 69.1 75.4 26.6
    OSMN ×\times ✓ 52.5 60.9 21.5 57.1 66.1 24.3
    SiamMask ×\times ×\times 54.3 62.8 19.3 58.5 67.5 20.9
    YouTube-VOS FT M Jseen↑\mathcal{J}_{\text{seen}} \uparrow Junseen↑\mathcal{J}_{\text{unseen}} \uparrow Fseen↑\mathcal{F}_{\text{seen}} \uparrow Funseen↑\mathcal{F}_{\text{unseen}} \uparrow Overall O↑\mathcal{O} \uparrow Speed
    OnAVOS ✓ ✓ 60.1 46.6 62.7 51.4 55.2 0.1 fps
    OSMN ×\times ✓ 60.0 40.6 60.1 44.0 51.2 8.0 fps
    SiamMask ×\times ×\times 60.2 45.1 58.2 47.7 52.8 55 fps

    SiamMask is orders of magnitude faster (55 fps vs. 0.1–8.0 fps) while achieving lower performance decay across time (JD=3.0\mathcal{J}_D = 3.0, FD=2.1\mathcal{F}_D = 2.1 on DAVIS-2016).

  8. Knowl 8 — Ablation Analysis of Backbones, Multi-Task Learning, and Mask Refinement

    data/table

    Ablation experiments isolate the effects of the neural network backbone, multi-task learning, and feature refinement on VOT-2018 (EAO) and DAVIS-2016 (region similarity JM\mathcal{J}_M and contour accuracy FM\mathcal{F}_M).

    Model Variant AN RN EAO ↑\uparrow JM↑\mathcal{J}_M \uparrow FM↑\mathcal{F}_M \uparrow Speed (fps)
    SiamFC ✓ 0.188 - - 86
    SiamFC ✓ 0.251 - - 40
    SiamRPN ✓ 0.243 - - 200
    SiamRPN ✓ 0.359 - - 76
    SiamMask-2B w/o R ✓ 0.326 62.3 55.6 43
    SiamMask w/o R ✓ 0.375 68.6 57.8 58
    SiamMask-2B-score ✓ 0.265 - - 40
    SiamMask-box ✓ 0.363 - - 76
    SiamMask-2B ✓ 0.334 67.4 63.5 60
    SiamMask ✓ 0.380 71.7 67.8 55

    Key conclusions:

    1. Upgrading the backbone fθf_\theta from AlexNet (AN) to ResNet-50 (RN) substantially boosts tracking accuracy (SiamRPN EAO rises from 0.243 to 0.359).
    2. Multi-task training alone improves box-only tracking: SiamMask-box achieves 0.363 EAO vs 0.359 for SiamRPN, and SiamMask-2B-score achieves 0.265 EAO vs 0.251 for SiamFC, proving the auxiliary mask loss benefits the shared feature representation.
    3. The refinement module ("R") provides major gains in boundary precision, increasing FM\mathcal{F}_M from 57.8% to 67.8% in SiamMask.
  9. Knowl 9 — SiamMask Implementation and Training Pipeline

    experimental setup

    The SiamMask model implementation and offline training configuration are detailed as follows:

    • Backbone Architecture: A modified ResNet-50 truncated after the 4th stage serves as fθf_\theta. To preserve spatial resolution, the output stride is reduced to 8 by applying stride-1 convolutions in stages 3 and 4, while receptive field size is maintained via atrous (dilated) convolutions. An unshared 1×11 \times 1 convolutional adjust layer projects features to 256 output channels.
    • Input Dimensions: Exemplar patches zz are 127×127×3127 \times 127 \times 3 pixels; search patches xx are 255×255×3255 \times 255 \times 3 pixels.
    • Data Augmentation: Training patches undergo random translations of up to ±8\pm 8 pixels and random scale jittering (2±1/82^{\pm 1/8} for exemplar, 2±1/42^{\pm 1/4} for search).
    • Training Datasets: Jointly trained on MS COCO, ImageNet-VID, and YouTube-VOS datasets.
    • Optimization: Backbone weights are pre-trained on ImageNet-1k. Training is performed with Stochastic Gradient Descent (SGD) for 20 epochs: a 5-epoch linear warmup where learning rate increases from 10−310^{-3} to 5×10−35 \times 10^{-3}, followed by 15 epochs of logarithmic decay down to 5×10−45 \times 10^{-4}.
  10. Knowl 10 — Failure Modes Under Severe Motion Blur and Non-Object Textures

    limitation

    SiamMask exhibits tracking failures primarily under two visual conditions:

    1. Severe motion blur: Extreme frame-to-frame target motion causes edge degradation and feature attenuation in the search patch.
    2. "Non-object" or background texture patterns: Arbitrary or ambiguous image regions lacking well-defined physical boundaries cannot be cleanly delineated.

    These failures arise from dataset distribution bias during offline training: video segmentation datasets like YouTube-VOS and image segmentation datasets like MS COCO exclusively annotate distinct, unambiguously foreground target instances.

Coverage note — None was omitted; all contributed methodology, loss formulations, inference mechanisms, bounding box generation techniques, benchmark results, ablations, and stated failure modes are represented.

References

  1. 1.B. Babenko, M.-H. Yang, and S. Belongie. Visual tracking with online multiple instance learning. In IEEE Conference on Computer Vision and Pattern Recognition, 2009.
  2. 2.L. Bao, B. Wu, and W. Liu. Cnn in mrf: Video object segmentation via inference in a cnn-based higher-order spatio-temporal mrf. In IEEE Conference on Computer Vision and Pattern Recognition, 2018.
  3. 3.L. Bertinetto, J. F. Henriques, J. Valmadre, P. H. S. Torr, and A. Vedaldi. Learning feed-forward one-shot learners. In Advances in Neural Information Processing Systems, 2016.
  4. 4.L. Bertinetto, J. Valmadre, J. F. Henriques, A. Vedaldi, and P. H. Torr. Fully-convolutional siamese networks for object tracking. In European Conference on Computer Vision workshops, 2016.
  5. 5.C. Bibby and I. Reid. Robust Real-Time Visual Tracking using Pixel-Wise Posteriors. In European Conference on Computer Vision, 2008.
  6. 6.D. S. Bolme, J. R. Beveridge, B. A. Draper, and Y. M. Lui. Visual object tracking using adaptive correlation filters. In IEEE Conference on Computer Vision and Pattern Recognition, 2010.
  7. 7.S. Caelles, K.-K. Maninis, J. Pont-Tuset, L. Leal-Taixe, D. Cremers, and L. Van Gool. One-shot video object segmentation. In IEEE Conference on Computer Vision and Pattern Recognition, 2017.
  8. 8.L.-C. Chen, G. Papandreou, I. Kokkinos, K. Murphy, and A. L. Yuille. Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2018.
  9. 9.Y. Chen, J. Pont-Tuset, A. Montes, and L. Van Gool. Blazingly fast video object segmentation with pixel-wise metric learning. In IEEE Conference on Computer Vision and Pattern Recognition, 2018.
  10. 10.J. Cheng, Y.-H. Tsai, W.-C. Hung, S. Wang, and M.-H. Yang. Fast and accurate online video object segmentation via tracking parts. In IEEE Conference on Computer Vision and Pattern Recognition, 2018.
  11. 11.J. Cheng, Y.-H. Tsai, S. Wang, and M.-H. Yang. Segflow: Joint learning for video object segmentation and optical flow. In IEEE International Conference on Computer Vision, 2017.
  12. 12.H. Ci, C. Wang, and Y. Wang. Video object segmentation by learning location-sensitive embeddings. In European Conference on Computer Vision, 2018.
  13. 13.D. Comaniciu, V. Ramesh, and P. Meer. Real-time tracking of non-rigid objects using mean shift. In IEEE Conference on Computer Vision and Pattern Recognition, 2000.
  14. 14.M. Danelljan, G. Bhat, F. S. Khan, M. Felsberg, et al. Eco: Efficient convolution operators for tracking. In IEEE Conference on Computer Vision and Pattern Recognition, 2017.
  15. 15.M. Danelljan, G. Hager, F. S. Khan, and M. Felsberg. Learning spatially regularized correlation filters for visual tracking. In IEEE International Conference on Computer Vision, 2015.
  16. 16.C. Feichtenhofer, A. Pinz, and A. Zisserman. Detect to track and track to detect. In IEEE International Conference on Computer Vision, 2017.
  17. 17.A. He, C. Luo, X. Tian, and W. Zeng. Towards a better match in siamese network based visual object tracker. In European Conference on Computer Vision workshops, 2018.
  18. 18.A. He, C. Luo, X. Tian, and W. Zeng. A twofold siamese network for real-time object tracking. In IEEE Conference on Computer Vision and Pattern Recognition, 2018.
  19. 19.K. He, G. Gkioxari, P. Dollar, and R. Girshick. Mask r-cnn. In IEEE International Conference on Computer Vision, 2017.
  20. 20.K. He, X. Zhang, S. Ren, and J. Sun. Deep residual learning for image recognition. In IEEE Conference on Computer Vision and Pattern Recognition, 2016.
  21. 21.D. Held, S. Thrun, and S. Savarese. Learning to track at 100 fps with deep regression networks. In European Conference on Computer Vision, 2016.
  22. 22.J. F. Henriques, R. Caseiro, P. Martins, and J. Batista. Highspeed tracking with kernelized correlation filters. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2015.
  23. 23.Y.-T. Hu, J.-B. Huang, and A. G. Schwing. Videomatch: Matching based video object segmentation. In European Conference on Computer Vision, 2018.
  24. 24.V. Jampani, R. Gadde, and P. V. Gehler. Video propagation networks. In IEEE Conference on Computer Vision and Pattern Recognition, 2017.
  25. 25.A. Khoreva, R. Benenson, E. Ilg, T. Brox, and B. Schiele. Lucid data dreaming for object tracking. In IEEE Conference on Computer Vision and Pattern Recognition workshops, 2017.
  26. 26.H. Kiani Galoogahi, T. Sim, and S. Lucey. Multi-channel correlation filters. In IEEE International Conference on Computer Vision, 2013.
  27. 27.H. Kiani Galoogahi, T. Sim, and S. Lucey. Correlation filters with limited boundaries. In IEEE Conference on Computer Vision and Pattern Recognition, 2015.
  28. 28.M. Kristan, A. Leonardis, J. Matas, M. Felsberg, R. Pflugfelder, L. Cehovin, T. Voj ır, G. Hager, A. Luke zi c, G. Fernandez, et al. The visual object tracking vot2016 challenge results. In European Conference on Computer Vision, 2016.
  29. 29.M. Kristan, A. Leonardis, J. Matas, M. Felsberg, R. Pfugfelder, L. C. Zajc, T. Vojir, G. Bhat, A. Lukezic, A. Eldesokey, G. Fernandez, and et al. The sixth visual object tracking vot-2018 challenge results. In European Conference on Computer Vision workshops, 2018.
  30. 30.M. Kristan, J. Matas, A. Leonardis, T. Vojıˇr, R. Pflugfelder, G. Fernandez, G. Nebehay, F. Porikli, and L. Cehovin. A novel performance evaluation methodology for single-target trackers. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2016.
  31. 31.B. Li, J. Yan, W. Wu, Z. Zhu, and X. Hu. High performance visual tracking with siamese region proposal network. In IEEE Conference on Computer Vision and Pattern Recognition, 2018.
  32. 32.F. Li, C. Tian, W. Zuo, L. Zhang, and M.-H. Yang. Learning spatial-temporal regularized correlation filters for visual tracking. In IEEE Conference on Computer Vision and Pattern Recognition, 2018.
  33. 33.X. Li and C. C. Loy. Video object segmentation with joint re-identification and attention-aware mask propagation. In European Conference on Computer Vision, 2018.
  34. 34.P. Liang, E. Blasch, and H. Ling. Encoding color information for visual tracking: Algorithms and benchmark. In IEEE Transactions on Image Processing, 2015.
  35. 35.T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollar, and C. L. Zitnick. Microsoft coco: Common objects in context. In European Conference on Computer Vision, 2014.
  36. 36.J. Long, E. Shelhamer, and T. Darrell. Fully convolutional networks for semantic segmentation. In IEEE Conference on Computer Vision and Pattern Recognition, 2015.
  37. 37.A. Lukezic, T. Vojir, L. C. Zajc, J. Matas, and M. Kristan. Discriminative correlation filter with channel and spatial reliability. In IEEE Conference on Computer Vision and Pattern Recognition, 2017.
  38. 38.T. Makovski, G. A. Vazquez, and Y. V. Jiang. Visual learning in multiple-object tracking. PLoS One, 2008.
  39. 39.K.-K. Maninis, S. Caelles, Y. Chen, J. Pont-Tuset, L. Leal-Taixe, D. Cremers, and L. Van Gool. Video object segmentation without temporal information. In IEEE Transactions on Pattern Analysis and Machine Intelligence, 2017.
  40. 40.N. Marki, F. Perazzi, O. Wang, and A. Sorkine-Hornung. Bilateral space video segmentation. In IEEE Conference on Computer Vision and Pattern Recognition, 2016.
  41. 41.O. Miksik, J.-M. Perez-R ua, P. H. Torr, and P. P erez. Roam: a rich object appearance model with application to rotoscoping. In IEEE Conference on Computer Vision and Pattern Recognition, 2017.
  42. 42.M. Mueller, N. Smith, and B. Ghanem. A benchmark and simulator for uav tracking. In European Conference on Computer Vision, 2016.
  43. 43.M. Muller, A. Bibi, S. Giancola, S. Al-Subaihi, and B. Ghanem. Trackingnet: A large-scale dataset and benchmark for object tracking in the wild. In European Conference on Computer Vision, 2018.
  44. 44.F. Perazzi. Video Object Segmentation. PhD thesis, ETH Zurich, 2017.
  45. 45.F. Perazzi, A. Khoreva, R. Benenson, B. Schiele, and A. Sorkine-Hornung. Learning video object segmentation from static images. In IEEE Conference on Computer Vision and Pattern Recognition, 2017.
  46. 46.F. Perazzi, J. Pont-Tuset, B. McWilliams, L. Van Gool, M. Gross, and A. Sorkine-Hornung. A benchmark dataset and evaluation methodology for video object segmentation. In IEEE Conference on Computer Vision and Pattern Recognition, 2017.
  47. 47.F. Perazzi, O. Wang, M. Gross, and A. Sorkine-Hornung. Fully connected object proposals for video segmentation. In IEEE International Conference on Computer Vision, 2015.
  48. 48.P. Perez, C. Hue, J. Vermaak, and M. Gangnet. Color-Based Probabilistic Tracking. In European Conference on Computer Vision, 2002.
  49. 49.P. O. Pinheiro, R. Collobert, and P. Dollar. Learning to segment object candidates. In Advances in Neural Information Processing Systems, 2015.
  50. 50.P. O. Pinheiro, T.-Y. Lin, R. Collobert, and P. Dollar. Learning to refine object segments. In European Conference on Computer Vision, 2016.
  51. 51.J. Pont-Tuset, F. Perazzi, S. Caelles, P. Arbelaez, A. Sorkine-Hornung, and L. Van Gool. The 2017 davis challenge on video object segmentation. arXiv preprint arXiv:1704.00675, 2017.
  52. 52.H. Possegger, T. Mauthner, and H. Bischof. In defense of color-based model-free tracking. In IEEE Conference on Computer Vision and Pattern Recognition, 2015.
  53. 53.S. Ren, K. He, R. Girshick, and J. Sun. Faster r-cnn: Towards real-time object detection with region proposal networks. In Advances in Neural Information Processing Systems, 2015.
  54. 54.O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, et al. Imagenet large scale visual recognition challenge. International Journal of Computer Vision, 2015.
  55. 55.A. W. Smeulders, D. M. Chu, R. Cucchiara, S. Calderara, A. Dehghan, and M. Shah. Visual tracking: An experimental survey. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2014.
  56. 56.R. Tao, E. Gavves, and A. W. Smeulders. Siamese instance search for tracking. In IEEE Conference on Computer Vision and Pattern Recognition, 2016.
  57. 57.Y.-H. Tsai, M.-H. Yang, and M. J. Black. Video segmentation via object flow. In IEEE Conference on Computer Vision and Pattern Recognition, 2016.
  58. 58.J. Valmadre, L. Bertinetto, J. Henriques, A. Vedaldi, and P. H. S. Torr. End-to-end representation learning for correlation filter based tracking. In IEEE Conference on Computer Vision and Pattern Recognition, 2017.
  59. 59.J. Valmadre*, L. Bertinetto*, J. F. Henriques, R. Tao, A. Vedaldi, A. Smeulders, P. H. S. Torr, and E. Gavves*. Long-term tracking in the wild: A benchmark. In European Conference on Computer Vision, 2018.
  60. 60.P. Voigtlaender and B. Leibe. Online adaptation of convolutional neural networks for video object segmentation. In British Machine Vision Conference, 2017.
  61. 61.L. Wen, D. Du, Z. Lei, S. Z. Li, and M.-H. Yang. Jots: Joint online tracking and segmentation. In IEEE Conference on Computer Vision and Pattern Recognition, 2015.
  62. 62.Y. Wu, J. Lim, and M.-H. Yang. Online object tracking: A benchmark. In IEEE Conference on Computer Vision and Pattern Recognition, 2013.
  63. 63.S. Wug Oh, J.-Y. Lee, K. Sunkavalli, and S. Joo Kim. Fast video object segmentation by reference-guided mask propagation. In IEEE Conference on Computer Vision and Pattern Recognition, 2018.
  64. 64.N. Xu, L. Yang, Y. Fan, J. Yang, D. Yue, Y. Liang, B. Price, S. Cohen, and T. Huang. Youtube-vos: Sequence-to-sequence video object segmentation. In European Conference on Computer Vision, 2018.
  65. 65.H. Yang, L. Shao, F. Zheng, L. Wang, and Z. Song. Recent advances and trends in visual tracking: A review. Neurocomputing, 2011.
  66. 66.L. Yang, Y. Wang, X. Xiong, J. Yang, and A. K. Katsaggelos. Efficient video object segmentation via network modulation. In IEEE Conference on Computer Vision and Pattern Recognition, June 2018.
  67. 67.T. Yang and A. B. Chan. Learning dynamic memory networks for object tracking. In European Conference on Computer Vision, 2018.
  68. 68.D. Yeo, J. Son, B. Han, and J. H. Han. Superpixel-based tracking-by-segmentation using markov chains. In IEEE Conference on Computer Vision and Pattern Recognition, 2017.
  69. 69.A. Yilmaz, O. Javed, and M. Shah. Object tracking: A survey. Acm computing surveys (CSUR), 2006.
  70. 70.J. S. Yoon, F. Rameau, J. Kim, S. Lee, S. Shin, and I. S. Kweon. Pixel-level matching for video object segmentation using convolutional neural networks. In IEEE International Conference on Computer Vision, 2017.
  71. 71.Z. Zhu, Q. Wang, B. Li, W. Wu, J. Yan, and W. Hu. Distractor-aware siamese networks for visual object tracking. In European Conference on Computer Vision, 2018.

Citation

MLA
Wang, Q., et al. “Fast Online Object Tracking and Segmentation: A Unifying Approach”. arXiv, 2018, http://arxiv.org/abs/1812.05050v2.
APA
Wang, Q., Zhang, L., Bertinetto, L., Hu, W., & Torr, P. H. S. (2018). Fast Online Object Tracking and Segmentation: A Unifying Approach. arXiv. http://arxiv.org/abs/1812.05050v2
Chicago
Wang, Q., L. Zhang, L. Bertinetto, W. Hu, and P. H. S. Torr. 2018. “Fast Online Object Tracking and Segmentation: A Unifying Approach”. arXiv. http://arxiv.org/abs/1812.05050v2.
Harvard
Wang, Q. et al. (2018) “Fast Online Object Tracking and Segmentation: A Unifying Approach”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1812.05050v2.
Vancouver
1. Wang Q, Zhang L, Bertinetto L, Hu W, Torr PHS (2018) Fast Online Object Tracking and Segmentation: A Unifying Approach. arXiv

BibTeX

@article{wang2018fast,
  title = {Fast Online Object Tracking and Segmentation: A Unifying Approach},
  author = {Wang, Qiang and Zhang, Li and Bertinetto, Luca and Hu, Weiming and Torr, Philip H. S.},
  year = {2018},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1812.05050v2},
  eprint = {1812.05050}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE