End-to-End Representation Learning for Correlation Filter Based Tracking

Jack ValmadreLuca BertinettoJoão F. HenriquesAndrea VedaldiPhilip H. S. Torr

article2017CVPR1,467 citations

Develops a differentiable correlation filter layer for deep networks, enabling end-to-end representation learning that allows lightweight models to achieve high-speed visual tracking with state-of-the-art accuracy.

Listen

Visual object tracking in video requires identifying an unknown object across successive frames after seeing it only once in an initial bounding box. Deep neural networks provide powerful representations for visual similarity, but adapting large networks during runtime is computationally expensive. Conversely, traditional Correlation Filters offer high-speed, per-frame online learning through fast mathematical solutions, yet they historically relied on separate, hand-crafted features or pre-trained neural representations not optimized for the tracking algorithm itself.

The article evaluates whether incorporating a Correlation Filter directly into a deep convolutional architecture and training it end-to-end improves tracking performance and efficiency. It demonstrates how to mathematically propagate errors through the closed-form Correlation Filter optimization during model training, tightly coupling the learned feature representation to the tracking algorithm.

To test this approach, the researchers built an asymmetric architecture termed CFNet, which integrates the filter into a fully-convolutional Siamese network framework. The model was trained offline across more than 3,800 videos containing over 1 million frames from the ImageNet Video dataset. Performance was systematically benchmarked across standard evaluation datasets—including OTB-2013, OTB-50, and OTB-100—measuring tracking accuracy, overlap success rates, and operating speeds against competitive real-time trackers.

The analysis revealed three critical findings. First, end-to-end integration delivers dramatic gains for shallow networks, providing a relative accuracy improvement of 31% for one-layer networks and 13% for two-layer networks over baseline Siamese architectures. Second, for deep networks with three to five layers, adding the filter provides diminishing returns, as deep embeddings alone are sufficiently expressive to match the performance. Third, a lightweight two-layer CFNet matches the accuracy of a five-layer baseline network while using less than 4% of the parameters, requiring only 600 kilobytes of storage, and running at 75 frames per second.

These results establish that end-to-end learning allows ultra-lightweight neural networks to achieve top-tier visual tracking performance. This presents a major operational advantage for low-power and embedded devices where memory capacity and computational budgets are constrained. The findings challenge the conventional practice of combining out-of-the-box classification features with filters, showing that features trained specifically for the tracking formulation perform significantly better.

Organizations developing real-time computer vision systems should consider adopting shallow, end-to-end integrated architectures for deployment on edge hardware to reduce hardware costs and memory footprints without sacrificing accuracy. Future development work should explore incorporating temporal adaptation mechanisms across successive frames and applying this differentiation technique to related domains like few-shot learning and domain adaptation.

The reported benchmarks are validated across established datasets with rigorous hyperparameter tuning on separate validation sets, providing high confidence in the comparative findings. However, the evaluation focused on core architecture efficiency rather than auxiliary enhancements like optical flow or bounding box regression, meaning absolute precision could be further influenced by standard pipeline post-processing.

Cover for End-to-End Representation Learning for Correlation Filter Based Tracking

Abstract

The Correlation Filter is an algorithm that trains a linear template to discriminate between images and their translations. It is well suited to object tracking because its formulation in the Fourier domain provides a fast solution, enabling the detector to be re-trained once per frame. Previous works that use the Correlation Filter, however, have adopted features that were either manually designed or trained for a different task. This work is the first to overcome this limitation by interpreting the Correlation Filter learner, which has a closed-form solution, as a differentiable layer in a deep neural network. This enables learning deep features that are tightly coupled to the Correlation Filter. Experiments illustrate that our method has the important practical benefit of allowing lightweight architectures to achieve state-of-the-art performance at high framerates.

Table of Contents

  • 1 Introduction
  • 2 Related work
  • 3 Method
  • 3.1 Fully-convolutional Siamese networks
  • 3.2 Tracking algorithm
  • 3.3 Correlation Filter networks
  • 3.4 Correlation Filter
  • 4 Experiments
  • 4.1 Evaluation criteria
  • 4.2 Comparison to Siamese baseline
  • 4.3 Feature transfer experiment
  • 4.4 Importance of adaptation
  • 4.5 Comparison with the state-of-the-art
  • 4.6 Speed and practical benefits
  • 5 Conclusion
  • A Implementation details
  • B Back-propagation for the Correlation Filter
  • C Correlation Filter formulation
  • C.1 Kernel linear regression
  • C.2 Single-channel Correlation Filter
  • C.3 Multi-channel Correlation Filter
  • D Adjoint of the differential
  • E Back-propagation for multi-channel case
  • F Hyperparameter optimization
  • G Detailed results on the OTB benchmarks
  • References

Knowls

  1. Knowl 1 — CFNet Asymmetric Siamese Architecture

    model/method

    CFNet is an asymmetric Siamese convolutional neural network for visual object tracking that integrates a differentiable Correlation Filter (CF) layer into deep representation learning. The network takes two image patches as input: a training exemplar x′∈RHx×Wx×3x' \in \mathbb{R}^{H_x \times W_x \times 3} centered on the target object, and a larger search area image z′∈RHz×Wz×3z' \in \mathbb{R}^{H_z \times W_z \times 3}.

    Both input images are mapped through a shared fully convolutional feature extractor fρf_\rho parameterized by learnable weights ρ\rho, yielding feature representations x=fρ(x′)∈Rm×m×Cx = f_\rho(x') \in \mathbb{R}^{m \times m \times C} and z=fρ(z′)∈Rm′×m′×Cz = f_\rho(z') \in \mathbb{R}^{m' \times m' \times C}.

    To mitigate circular boundary artifacts, the exemplar feature map xx is pre-multiplied element-wise by a 2D cosine window. The Correlation Filter block ω(x)\omega(x) then computes a multi-channel discriminative linear template w=ω(x)∈Rk×k×Cw = \omega(x) \in \mathbb{R}^{k \times k \times C} by solving a regularized ridge regression problem in closed form in the Fourier domain. The resulting template ww is cropped to a smaller spatial extent.

    The response map is produced by cross-correlating the learned template ww across the search features fρ(z′)f_\rho(z'), adjusted with learnable scalar scale ss and bias bb parameters: hρ,s,b(x′,z′)=s (ω(fρ(x′))⋆fρ(z′))+bh_{\rho, s, b}(x', z') = s \, (\omega(f_\rho(x')) \star f_\rho(z')) + b where ⋆\star denotes multi-channel spatial cross-correlation. The spatial maximum of hρ,s,b(x′,z′)h_{\rho, s, b}(x', z') corresponds to the predicted target location.

  2. Knowl 2 — Closed-Form Multi-Channel Correlation Filter Layer in the Fourier Domain

    theoretical result

    Given a multi-channel input feature map x=(x1,…,xC)x = (x_1, \dots, x_C) where each channel xp∈Rm×mx_p \in \mathbb{R}^{m \times m}, and a desired scalar response map y∈Rm×my \in \mathbb{R}^{m \times m} (a 2D Gaussian peaked at the target center), the optimal multi-channel Correlation Filter template w=(w1,…,wC)w = (w_1, \dots, w_C) minimizes the regularized least-squares objective over all circular shifts: arg⁡min⁡w12n∥∑p=1Cwp⋆xp−y∥22+λ2∑p=1C∥wp∥22\arg \min_{w} \frac{1}{2n} \left\| \sum_{p=1}^C w_p \star x_p - y \right\|_2^2 + \frac{\lambda}{2} \sum_{p=1}^C \|w_p\|_2^2 where n=m2n = m^2 denotes the number of spatial pixels/circular shifts, ⋆\star denotes circular cross-correlation, and λ>0\lambda > 0 is a quadratic regularization parameter.

    Due to the circulant structure of circular shifts, the dual solution can be solved analytically and computed in the Fourier domain with computational complexity that scales linearly in the number of channels CC: k^=1n∑p=1C(x^p∗∘x^p)+λ1\hat{k} = \frac{1}{n} \sum_{p=1}^C (\hat{x}_p^* \circ \hat{x}_p) + \lambda \mathbf{1} α^=1nk^−1∘y^\hat{\alpha} = \frac{1}{n} \hat{k}^{-1} \circ \hat{y} w^p=α^∗∘x^p∀p∈{1,…,C}\hat{w}_p = \hat{\alpha}^* \circ \hat{x}_p \quad \forall p \in \{1, \dots, C\} where ⋅^\hat{\cdot} denotes the 2D Discrete Fourier Transform (DFT), x^p∗\hat{x}_p^* is the complex conjugate of x^p\hat{x}_p, ∘\circ denotes element-wise Hadamard product, k^−1\hat{k}^{-1} denotes element-wise scalar inversion, and 1∈Rm×m\mathbf{1} \in \mathbb{R}^{m \times m} is a tensor of ones. The spatial template channels are recovered via inverse DFT: wp=F−1(w^p)w_p = \mathcal{F}^{-1}(\hat{w}_p).

  3. Knowl 3 — Closed-Form Back-Propagation for the Multi-Channel Correlation Filter Layer

    theoretical result

    To train deep feature extractors end-to-end through the Correlation Filter layer, gradients must be back-propagated through the closed-form ridge regression solver. By applying Parseval's theorem, which ensures unitary inner product preservation under the Fourier transform, the adjoint of the differential linear maps yields closed-form expressions in the Fourier domain.

    Given the gradient of a scalar loss ℓ\ell with respect to the template Fourier channels, ∇w^pℓ\nabla_{\hat{w}_p} \ell, and optionally with respect to the desired response, ∇y^ℓ\nabla_{\hat{y}} \ell, the gradients with respect to the dual multipliers α^\hat{\alpha}, the kernel signal k^\hat{k}, and the input feature channels x^p\hat{x}_p are given by: ∇α^ℓ=∑p=1Cx^p∘(∇w^pℓ)∗\nabla_{\hat{\alpha}} \ell = \sum_{p=1}^C \hat{x}_p \circ (\nabla_{\hat{w}_p} \ell)^* ∇y^ℓ=1nk^−∗∘∇α^ℓ\nabla_{\hat{y}} \ell = \frac{1}{n} \hat{k}^{-*} \circ \nabla_{\hat{\alpha}} \ell ∇k^ℓ=−k^−∗∘α^∗∘∇α^ℓ\nabla_{\hat{k}} \ell = -\hat{k}^{-*} \circ \hat{\alpha}^* \circ \nabla_{\hat{\alpha}} \ell ∇x^pℓ=α^∘∇w^pℓ+2nx^p∘Re{∇k^ℓ}∀p∈{1,…,C}\nabla_{\hat{x}_p} \ell = \hat{\alpha} \circ \nabla_{\hat{w}_p} \ell + \frac{2}{n} \hat{x}_p \circ \text{Re}\{\nabla_{\hat{k}} \ell\} \quad \forall p \in \{1, \dots, C\} where k^−∗=(k^−1)∗\hat{k}^{-*} = (\hat{k}^{-1})^*, Re{⋅}\text{Re}\{\cdot\} extracts the real component, n=m2n = m^2 is the number of spatial pixels, and ∘\circ is element-wise multiplication. Spatial gradients are recovered by inverse DFT: ∇xpℓ=F−1(∇x^pℓ)\nabla_{x_p} \ell = \mathcal{F}^{-1}(\nabla_{\hat{x}_p} \ell). The computational complexity of the backward pass scales as O(C⋅m2log⁡m)\mathcal{O}(C \cdot m^2 \log m).

  4. Knowl 4 — CFNet Online Visual Tracking Algorithm

    algorithm

    Online visual tracking with CFNet operates by extracting candidate search windows at multiple scales, evaluating the forward score maps, updating object position and scale, and maintaining an adaptive template running average.

    Input: Video frames I1,…,ITI_1, \dots, I_T, initial bounding box B1B_1, CNN fρf_\rho, CF solver ω\omega, search window scale factor 44, scale step factor 1.041.04, scale penalty 0.970.97, scale update rate γs=0.6\gamma_s = 0.6, template update rate η\eta, cosine window WcosW_{cos}, score penalty window WscoreW_{score}
    Output: Bounding boxes B2,…,BTB_2, \dots, B_T
    Extract target patch x1′x'_1 from I1I_1 at B1B_1
    Compute initial template: wcum=ω(fρ(x1′)∘Wcos)w_{cum} = \omega(f_\rho(x'_1) \circ W_{cos})
    Initialize target scale S1=(width1,height1)S_1 = (\text{width}_1, \text{height}_1) from B1B_1
    for t=2t = 2 to TT do
        for each scale scale multiplier σ∈{1/1.04,1.0,1.04}\sigma \in \{1/1.04, 1.0, 1.04\} do
            Extract search patch zt,σ′z'_{t, \sigma} around center Pt−1P_{t-1} of size 4×(St−1⋅σ)4 \times (S_{t-1} \cdot \sigma)
            Compute score map Rσ=s⋅(wcum⋆fρ(zt,σ′))+bR_\sigma = s \cdot (w_{cum} \star f_\rho(z'_{t, \sigma})) + b
            Apply spatial translation penalty: Rσ=Rσ∘WscoreR_\sigma = R_\sigma \circ W_{score}
            if σ≠1.0\sigma \neq 1.0 then
                Apply scale change penalty: Rσ=Rσ×0.97R_\sigma = R_\sigma \times 0.97
            end if
        end for
        Select optimal scale σ∗=arg⁡max⁡σ(max⁡Rσ)\sigma^* = \arg\max_\sigma (\max R_\sigma)
        Find displacement Δp=arg⁡max⁡Rσ∗\Delta p = \arg\max R_{\sigma^*}
        Update center position: Pt=Pt−1+ΔpP_t = P_{t-1} + \Delta p
        Update scale: St=(1−γs)St−1+γs(St−1⋅σ∗)S_t = (1 - \gamma_s) S_{t-1} + \gamma_s (S_{t-1} \cdot \sigma^*)
        Set bounding box Bt=(Pt,St)B_t = (P_t, S_t)
        Extract new exemplar patch xt′x'_t from ItI_t at BtB_t
        Compute current template: wt=ω(fρ(xt′)∘Wcos)w_t = \omega(f_\rho(x'_t) \circ W_{cos})
        Update moving average template: wcum=(1−η)wcum+ηwtw_{cum} = (1 - \eta) w_{cum} + \eta w_t
    end for
  5. Knowl 5 — End-to-End Offline Training Setup for CFNet

    experimental setup

    CFNet is trained offline end-to-end using pairs of images (xi′,zi′)(x'_i, z'_i) sampled from the 3,862 training videos of the ImageNet Video (VID) dataset containing over one million annotated frames.

    Pairs (xi′,zi′)(x'_i, z'_i) are randomly selected from the same video sequence such that their temporal distance does not exceed 100 frames. The exemplar patch xi′x'_i (255×255×3255 \times 255 \times 3) is centered on the target object, and the search image zi′z'_i (255×255×3255 \times 255 \times 3) contains the search region. Each pair is associated with an element-wise label grid ci∈{−1,+1}m×mc_i \in \{-1, +1\}^{m \times m}, where +1+1 represents the true object position and −1-1 represents background positions.

    The feature parameters ρ\rho, scale ss, and bias bb are optimized by minimizing an element-wise logistic loss over the training dataset: arg⁡min⁡ρ,s,b∑i∑ulog⁡(1+exp⁡(−ci[u] hρ,s,b(xi′,zi′)[u]))\arg \min_{\rho, s, b} \sum_i \sum_u \log \left( 1 + \exp\left( -c_i[u] \, h_{\rho, s, b}(x'_i, z'_i)[u] \right) \right)

    Training is performed using stochastic gradient descent (SGD) for 100 epochs, sampling approximately 12 pairs per video per epoch with mini-batches of size 8. Parameters are initialized using Xavier initialization. The network feature extractor uses a reduced total stride of 4 (stride 2 at conv1, stride 2 at pool1) and restricts the final feature map to 32 output channels.

  6. Knowl 6 — Accuracy-Depth Saturation and Ultra-Lightweight Efficiency of CFNet

    empirical result

    Evaluation of CFNet against a fully-convolutional Siamese baseline (SiamFC) across network depths from 1 to 5 convolutional layers demonstrates two primary phenomena:

    1. Substantial accuracy gains at shallow depths: Incorporating the Correlation Filter layer during offline training provides a relative improvement in average IoU overlap of 31% for depth 1 (conv1) and 13% for depth 2 (conv2) over the baseline Siamese architecture.
    2. Depth saturation: At depths 3, 4, and 5, CFNet accuracy plateaus around 48% average overlap, and the performance difference between CFNet and the baseline becomes negligible. The inductive bias of the Correlation Filter becomes redundant when the deep embedding function itself possesses sufficient expressive capacity and representational depth.
    3. Ultra-lightweight efficiency: CFNet with a 2-layer backbone (CFNet-conv2) achieves tracking accuracy comparable to a 5-layer baseline (Baseline-conv5), while containing approximately 30×30\times fewer parameters, requiring only 600 kB of storage (<4% of Baseline-conv5), and operating at 75 frames per second (fps) on a Titan X GPU (compared to 52 fps for Baseline-conv5). The single-layer CFNet-conv1 runs at 83 fps with under 100 kB of parameter storage (<1% of Baseline-conv5).
  7. Knowl 7 — Feature Transfer vs. End-to-End Correlation Filter Training

    empirical result

    Comparing end-to-end trained CFNet against transfer learning configurations illustrates the necessity of training deep representations jointly with the Correlation Filter layer:

    1. CFNet (End-to-End): Features trained directly with the closed-form CF solver in the computational graph.
    2. Baseline+CF (Transfer from Siamese): Features trained using a standard Siamese embedding objective, with a Correlation Filter applied post-hoc during tracking.
    3. ImageNet+CF (Transfer from Classification): Features taken from a network pre-trained on ImageNet object classification and combined with a CF during tracking.

    On a 129-video validation benchmark:

    • At depths 1 and 2, CFNet substantially outperforms Baseline+CF, proving that back-propagating gradients through the inverse convolution problem forces early convolutional layers to learn filters optimized for translation discrimination.
    • At depths 3 through 5, Baseline+CF achieves performance on par with CFNet.
    • ImageNet+CF yields significantly inferior tracking performance across all depths compared to both CFNet and Baseline. This gap worsens at deeper layers (conv4 and conv5) because classification networks learn translation-invariant representations that discard precise spatial localization cues required for correlation tracking.
  8. Knowl 8 — Necessity of Dynamic Online Adaptation vs Static Filter Weights

    empirical result

    To determine whether the performance advantages of CFNet stem from dynamic online adaptation (solving the ridge regression system per exemplar) or merely from the mathematical form of the layer, CFNet was evaluated against a constant variant (CFNet-const). In CFNet-const, the Lagrange multiplier vector α\alpha is treated as a static parameter learned offline and held fixed during test-time tracking, disabling dynamic exemplar-conditioned inverse convolution.

    Across all network depths (1 through 5) on a 129-video validation set, CFNet consistently outperforms CFNet-const. For example, at depth 1, CFNet achieves ~44.8% average IoU overlap compared to ~35.5% for CFNet-const. This establishes that back-propagating through the analytical solution of the inverse convolution problem—enabling the template to dynamically adapt to the exemplar appearance online—is indispensable for the accuracy gains.

  9. Knowl 9 — Tracking Performance on OTB-2013, OTB-50, and OTB-100 Benchmarks

    data/table

    Performance of CFNet variants evaluated on OTB-2013, OTB-50, and OTB-100 benchmarks under One-Pass Evaluation (OPE) and Temporal Robustness Evaluation (TRE), measured by Area Under Curve IoU overlap and precision (at 20 pixels location error threshold):

    OTB-2013 OTB-50 OTB-100
    OPE TRE OPE TRE OPE TRE
    Method speed (fps) IoU prec. IoU prec. IoU prec. IoU prec. IoU prec. IoU prec.
    CFNet-conv1 83 57.8 77.6 58.6 77.6 48.8 65.3 51.0 67.9 53.6 71.3 55.9 72.6
    CFNet-conv2 75 61.1 80.7 64.0 84.8 53.0 70.2 56.5 75.3 56.8 74.8 60.6 79.1
    Baseline+CF-conv3 67 61.0 82.2 63.1 83.9 53.8 72.3 57.4 76.7 58.9 77.7 61.1 79.8
    CFNet-conv5 43 61.1 80.3 62.6 82.5 53.9 73.2 56.6 75.9 58.6 77.7 60.8 78.8
    Baseline-conv5 52 61.8 80.6 64.0 83.7 51.7 68.3 56.1 74.2 58.8 76.9 61.6 79.7
    SiamFC-3s - 60.7 81.0 61.8 82.2 51.6 69.2 55.5 75.2 58.2 77.0 60.5 79.5
    Staple - 60.0 79.3 61.7 80.3 50.9 68.1 54.1 72.6 58.1 78.4 60.4 78.9
    LCT - 61.2 86.2 59.4 81.3 49.2 69.1 49.5 67.4 56.2 76.2 56.9 74.5
    SAMF - - - - - 46.2 63.9 51.4 70.9 53.9 74.6 57.7 77.6
    DSST - 55.4 74.0 56.6 73.8 45.2 60.4 48.4 64.1 51.3 68.0 - -

    The results indicate that CFNet-conv2 matches the performance of the deeper 5-layer Baseline-conv5 and surpasses real-time baselines (SiamFC-3s, Staple, LCT, SAMF, DSST) across benchmarks while operating at high frame rates (75 fps).

  10. Knowl 10 — Tracking Hyperparameter Sensitivity and Optimization via Random Search

    data/table

    Online tracking accuracy depends strongly on non-differentiable tracking hyperparameters: scale step multiplier, multiplicative scale penalty, scale update learning rate, spatial cosine window penalty weight, and template update learning rate. Optimal configurations selected from 300 uniform random search iterations on a 129-video validation set:

    Architecture Avg Overlap Best Overlap Scale Step Scale Penalty Scale L.R. Win. Weight Template L.R.
    CFNet-conv1 44.8% 46.5% 1.0355 0.9825 0.700 0.2375 0.0058
    CFNet-conv2 47.8% 49.5% 1.0575 0.9780 0.520 0.2625 0.0050
    Baseline+CF-conv3 47.7% 49.9% 1.0340 0.9820 0.660 0.2700 0.0080
    CFNet-conv5 46.9% 48.5% 1.0310 0.9815 0.525 0.2000 0.0110
    Baseline-conv5 47.8% 49.2% 1.0470 0.9825 0.680 0.1750 0.0102

    The distributions demonstrate that all tracking architectures benefit from conservative template update rates (η≈0.005−0.011\eta \approx 0.005 - 0.011) and scale search steps around 1.03−1.061.03 - 1.06, with CFNet-conv2 achieving a peak validation overlap of 49.5%, outperforming 5-layer variants.

Coverage note — None was omitted; all primary architectural components, mathematical derivations (forward and backward passes), tracking logic, and empirical analyses from the paper are represented.

References

  1. 1.L. Bertinetto, J. F. Henriques, J. Valmadre, P. H. S. Torr, and A. Vedaldi. Learning feed-forward one-shot learners. In NIPS, pages 523–531, 2016. 2
  2. 2.L. Bertinetto, J. Valmadre, S. Golodetz, O. Miksik, and P. H. S. Torr. Staple: Complementary learners for real-time tracking. In CVPR, pages 1401–1409, 2016. 2, 7
  3. 3.L. Bertinetto, J. Valmadre, J. F. Henriques, A. Vedaldi, and P. H. S. Torr. Fully-convolutional Siamese networks for object tracking. In ECCV Workshops, pages 850–865, 2016. 1, 2, 3, 5, 7, 8
  4. 4.D. S. Bolme, J. R. Beveridge, B. A. Draper, and Y. M. Lui. Visual object tracking using adaptive correlation filters. In CVPR, 2010. 2, 3, 5
  5. 5.K. Chen and W. Tao. Once for all: A two-flow convolutional neural network for visual tracking. arXiv preprint arXiv:1604.07507, 2016. 1
  6. 6.M. Danelljan, G. Hager, F. Khan, and M. Felsberg. Accurate scale estimation for robust visual tracking. In BMVC, 2014. 2, 3, 7
  7. 7.M. Danelljan, G. Hager, F. Shahbaz Khan, and M. Felsberg. Convolutional features for correlation filter based visual tracking. In ICCV Workshops, pages 58–66, 2015. 2, 3, 6, 8
  8. 8.M. Danelljan, G. Hager, F. Shahbaz Khan, and M. Felsberg. Learning spatially regularized correlation filters for visual tracking. In ICCV, pages 4310–4318, 2015. 2, 3
  9. 9.M. Danelljan, A. Robinson, F. S. Khan, and M. Felsberg. Beyond correlation filters: Learning continuous convolution operators for visual tracking. In ECCV, pages 472–488, 2016. 2, 6, 8
  10. 10.J. A. Fernandez and B. Vijayakumar. Zero-aliasing correlation filters. In International Symposium on Image and Signal Processing and Analysis 2013, pages 101–106, 2013. 2
  11. 11.S. Gould, B. Fernando, A. Cherian, P. Anderson, R. S. Cruz, and E. Guo. On differentiating parameterized argmin and argmax problems with application to bi-level optimization. arXiv preprint arXiv:1607.05447, 2016. 2
  12. 12.A. Griewank and A. Walther. Evaluating derivatives: Principles and techniques of algorithmic differentiation. SIAM, 2008. 10
  13. 13.D. Held, S. Thrun, and S. Savarese. Learning to track at 100 fps with deep regression networks. In ECCV, pages 749–765. Springer, 2016. 1, 2
  14. 14.J. F. Henriques, R. Caseiro, P. Martins, and J. Batista. High-speed tracking with kernelized correlation filters. IEEE TPAMI, 37(3):583–596, 2015. 2, 3, 4, 5, 9
  15. 15.C. Ionescu, O. Vantzos, and C. Sminchisescu. Matrix back-propagation for deep networks with structured layers. In ICCV, pages 2965–2973, 2015. 2, 4, 10
  16. 16.H. Kiani Galoogahi, T. Sim, and S. Lucey. Correlation filters with limited boundaries. In CVPR, pages 4630–4638, 2015. 2, 3
  17. 17.M. Kristan, A. Leonardis, J. Matas, M. Felsberg, R. Pflugfelder, L. Čehovin, T. Vojíř, G. Häger, A. Lukežič, G. Fernández, et al. The Visual Object Tracking VOT2016 challenge results. In ECCV Workshops. Springer, 2016. 5
  18. 18.L. Leal-Taixé, C. Canton-Ferrer, and K. Schindler. Learning by tracking: Siamese CNN for robust target association. In CVPR Workshops, pages 33–40, 2016. 1
  19. 19.Y. Li and J. Zhu. A scale adaptive kernel correlation filter tracker with feature integration. In ECCV, pages 254–265, 2014. 2, 7
  20. 20.P. Liang, E. Blasch, and H. Ling. Encoding color information for visual tracking: Algorithms and benchmark. IEEE Transactions on Image Processing, 24(12):5630–5644, 2015. 5
  21. 21.J. Long, E. Shelhamer, and T. Darrell. Fully convolutional networks for semantic segmentation. In CVPR, pages 3431–3440, 2015. 2
  22. 22.C. Ma, J.-B. Huang, X. Yang, and M.-H. Yang. Hierarchical convolutional features for visual tracking. In ICCV, pages 3074–3082, 2015. 2, 3, 6
  23. 23.C. Ma, X. Yang, C. Zhang, and M.-H. Yang. Long-term correlation tracking. In CVPR, pages 5388–5396, 2015. 2, 3, 7
  24. 24.D. Maclaurin, D. Duvenaud, and R. P. Adams. Gradient-based hyperparameter optimization through reversible learning. In ICML, 2015. 2
  25. 25.I. Murray. Differentiation of the Cholesky decomposition. arXiv preprint arXiv:1602.07527, 2016. 2
  26. 26.H. Nam and B. Han. Learning multi-domain convolutional neural networks for visual tracking. In CVPR 2016, pages 4293–4302, 2016. 1, 6, 7
  27. 27.A. Rodriguez, V. N. Boddeti, B. V. K. V. Kumar, and A. Mahalanobis. Maximum margin correlation filter: A new approach for localization and classification. IEEE Transactions on Image Processing, 22(2):631–643, 2013. 2
  28. 28.O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, A. C. Berg, and L. Fei-Fei. ImageNet Large Scale Visual Recognition Challenge. International Journal of Computer Vision (IJCV), 115(3):211–252, 2015. 6, 8
  29. 29.R. Tao, E. Gavves, and A. W. M. Smeulders. Siamese instance search for tracking. In CVPR, pages 1420–1429, 2016. 1, 2, 7
  30. 30.J. Valmadre, S. Sridharan, and S. Lucey. Learning detectors quickly with stationary statistics. In ACCV, pages 99–114. Springer, 2014. 3
  31. 31.O. Vinyals, C. Blundell, T. Lillicrap, D. Wierstra, et al. Matching networks for one shot learning. In NIPS, pages 3630–3638, 2016. 2
  32. 32.N. Wang, S. Li, A. Gupta, and D.-Y. Yeung. Transferring rich feature hierarchies for robust visual tracking. arXiv preprint arXiv:1501.04587, 2015. 1, 2, 6
  33. 33.Y. Wu, J. Lim, and M.-H. Yang. Online object tracking: A benchmark. In CVPR, pages 2411–2418, 2013. 5, 7
  34. 34.Y. Wu, J. Lim, and M.-H. Yang. Object tracking benchmark. TPAMI, 37(9):1834–1848, 2015. 5, 7
  35. 35.M. D. Zeiler, D. Krishnan, G. W. Taylor, and R. Fergus. Deconvolutional networks. In CVPR, pages 2528–2535, 2010. 2
  36. 36.M. Zhai, M. J. Roshtkhari, and G. Mori. Deep learning of appearance models for online object tracking. arXiv preprint arXiv:1607.02568, 2016. 1, 6
  37. 37.S. Zheng, S. Jayasumana, B. Romera-Paredes, V. Vineet, Z. Su, D. Du, C. Huang, and P. H. S. Torr. Conditional random fields as recurrent neural networks. In ICCV, pages 1529–1537, 2015. 2

Citation

MLA
Valmadre, J., et al. “End-to-end Representation Learning for Correlation Filter Based Tracking”. arXiv, 2017, http://arxiv.org/abs/1704.06036v1.
APA
Valmadre, J., Bertinetto, L., Henriques, J. F., Vedaldi, A., & Torr, P. H. S. (2017). End-to-end representation learning for Correlation Filter based tracking. arXiv. http://arxiv.org/abs/1704.06036v1
Chicago
Valmadre, J., L. Bertinetto, J. F. Henriques, A. Vedaldi, and P. H. S. Torr. 2017. “End-to-end Representation Learning for Correlation Filter Based Tracking”. arXiv. http://arxiv.org/abs/1704.06036v1.
Harvard
Valmadre, J. et al. (2017) “End-to-end representation learning for Correlation Filter based tracking”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1704.06036v1.
Vancouver
1. Valmadre J, Bertinetto L, Henriques JF, Vedaldi A, Torr PHS (2017) End-to-end representation learning for Correlation Filter based tracking. arXiv

BibTeX

@article{valmadre2017end,
  title = {End-to-end representation learning for Correlation Filter based tracking},
  author = {Valmadre, Jack and Bertinetto, Luca and Henriques, João F. and Vedaldi, Andrea and Torr, Philip H. S.},
  year = {2017},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1704.06036v1},
  eprint = {1704.06036}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE