Optical Flow Estimation Using a Spatial Pyramid Network

Anurag RanjanMichael J. Black

article2016CVPR1,434 citations

Introduces SPyNet, a lightweight optical flow architecture that combines classical coarse-to-fine spatial pyramids with deep learning to exceed FlowNet accuracy while reducing model parameters by 96%.

Listen

Estimating optical flow—the pattern of apparent motion of objects across video frames—is a critical computer vision capability for autonomous systems, robotics, mobile devices, and video processing. Classical approaches rely on hand-crafted assumptions and iterative mathematical optimization, which can be computationally slow and struggle with complex real-world conditions. While recent deep learning models like FlowNet attempt to learn motion directly, they rely on massive network architectures that demand substantial memory and computational resources, limiting their practicality for real-time and embedded hardware.

The article demonstrates and evaluates a hybrid framework called Spatial Pyramid Network (SPyNet), which combines classical coarse-to-fine spatial pyramid principles with compact deep convolutional neural networks. The objective is to produce a model that achieves state-of-the-art flow accuracy and faster runtimes while drastically reducing parameter count and memory consumption.

The evaluation uses standard computer vision benchmarks—including MPI Sintel, KITTI, and Middlebury—with training conducted primarily on the synthetic Flying Chairs dataset and fine-tuning applied to specific target environments. Rather than forcing a single neural network to resolve both large-scale displacements and minute sub-pixel shifts simultaneously, the approach employs a multi-level image pyramid. Large motions are handled structurally at coarser resolutions via image warping, allowing small five-layer convolutional networks at each pyramid level to focus exclusively on estimating small, residual flow corrections of less than a few pixels.

The findings show that SPyNet reduces model size by approximately 96% compared to FlowNet, utilizing only 1.2 million parameters (9.7 MB of storage) versus over 32 million parameters. In terms of runtime, the model operates at 0.069 seconds per frame, outperforming FlowNet variants (0.080 to 0.150 seconds per frame) while offering superior or comparable accuracy across standard benchmarks. Specifically, after fine-tuning, SPyNet achieves significantly lower error rates on Middlebury and KITTI datasets and performs especially well near motion boundaries and across small-to-moderate velocity ranges. Additionally, visual analysis reveals that the network naturally learns structured, biologically plausible spatio-temporal filters, unlike the unstructured filters seen in unconstrained end-to-end models.

These results demonstrate that reintroducing well-engineered classical vision concepts into deep learning pipelines can yield substantial gains in computational efficiency, storage footprint, and execution speed. For technical leaders and engineering teams, this makes high-accuracy optical flow practical for low-power edge processors, drones, and mobile GPU hardware without incurring heavy engineering overhead.

Moving forward, development teams should explore deploying this compact architecture into real-time mobile and robotic pipelines. Next steps should also include augmenting the architecture with long-range feature matching or channel constancy techniques to overcome the primary limitation of spatial pyramids—namely, capturing the motion of small, fast-moving, or thin objects that disappear at coarse pyramid resolutions. Readers can place high confidence in the demonstrated performance gains across standard benchmarks, though further training on realistic datasets beyond synthetic chairs is recommended before deploying the model in specialized production environments.

Cover for Optical Flow Estimation Using a Spatial Pyramid Network

Abstract

We learn to compute optical flow by combining a classical spatial-pyramid formulation with deep learning. This estimates large motions in a coarse-to-fine approach by warping one image of a pair at each pyramid level by the current flow estimate and computing an update to the flow. Instead of the standard minimization of an objective function at each pyramid level, we train one deep network per level to compute the flow update. Unlike the recent FlowNet approach, the networks do not need to deal with large motions; these are dealt with by the pyramid. This has several advantages. First, our Spatial Pyramid Network (SPyNet) is much simpler and 96% smaller than FlowNet in terms of model parameters. This makes it more efficient and appropriate for embedded applications. Second, since the flow at each pyramid level is small (< 1 pixel), a convolutional approach applied to pairs of warped images is appropriate. Third, unlike FlowNet, the learned convolution filters appear similar to classical spatio-temporal filters, giving insight into the method and how to improve it. Our results are more accurate than FlowNet on most standard benchmarks, suggesting a new direction of combining classical flow methods with deep learning.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Spatial Pyramid Network
  • 3.1 Spatial Sampling
  • 3.2 Inference
  • 3.3 Training and Network Architecture
  • 4 Experiments
  • 5 Analysis
  • 6 Discussion and Future Work
  • 7 Conclusions
  • References

Knowls

  1. Knowl 1 — SPyNet Coarse-to-Fine Optical Flow Formulation

    model/method

    The Spatial Pyramid Network (SPyNet) estimates optical flow between two frames I1I^1 and I2I^2 by combining a classical coarse-to-fine pyramid framework with deep convolutional networks.

    Let d(⋅)d(\cdot) be a spatial downsampling operator that decimates an image by a factor of 2, u(⋅)u(\cdot) be a spatial upsampling operator by a factor of 2 using bilinear interpolation, and w(I,V)w(I, V) denote the warping of an image II according to flow field VV via bilinear interpolation.

    A KK-level spatial pyramid is formed where Ik1I_k^1 and Ik2I_k^2 denote the image pair downsampled to the kk-th level for k∈{0,1,…,K}k \in \{0, 1, \dots, K\}, with level 0 being the coarsest (smallest) resolution and level KK being the full resolution.

    At each level kk, a convolutional neural network GkG_k computes a residual flow field vkv_k:

    vk=Gk(Ik1,w(Ik2,u(Vk−1)),u(Vk−1))v_k = G_k\left(I_k^1, w(I_k^2, u(V_{k-1})), u(V_{k-1})\right)

    where u(Vk−1)u(V_{k-1}) is the upsampled flow field from the previous coarser level (for k=0k = 0, u(V−1)u(V_{-1}) is set to 0 everywhere). The total accumulated flow VkV_k at level kk is updated as:

    Vk=u(Vk−1)+vkV_k = u(V_{k-1}) + v_k

    The network GkG_k receives an 8-channel input tensor consisting of the reference image Ik1I_k^1 (3 RGB channels), the warped second image w(Ik2,u(Vk−1))w(I_k^2, u(V_{k-1})) (3 RGB channels), and the upsampled flow field u(Vk−1)u(V_{k-1}) (2 channels for horizontal and vertical displacement). The output is a 2-channel residual flow vkv_k.

  2. Knowl 2 — SPyNet Per-Level ConvNet Architecture and Model Size

    model/method

    Each level network GkG_k in the SPyNet pyramid is a 5-layer convolutional network designed to predict small sub-pixel residual updates.

    The architectural specifications for each GkG_k are:

    • Number of convolutional layers: 5
    • Spatial filter kernel size: 7×77 \times 7 for all layers
    • Output feature map channels across layers: {32,64,32,16,2}\{32, 64, 32, 16, 2\}
    • Non-linear activations: Rectified Linear Units (ReLU) following each convolutional layer except the final output layer, which outputs linear 2D flow components (vx,vy)(v^x, v^y)
    • Input dimensions: 8 channels (33 for Ik1I_k^1, 33 for w(Ik2,u(Vk−1))w(I_k^2, u(V_{k-1})), and 22 for u(Vk−1)u(V_{k-1}))

    Each sub-network GkG_k contains 240,050 trainable parameters. For a 5-level pyramid network (K=4K = 4), the entire SPyNet model comprises 1,200,250 parameters, requiring 9.7 MB of disk storage. This is approximately 96% smaller than FlowNetS (32,070,472 parameters) and FlowNetC (32,561,032 parameters).

  3. Knowl 3 — Sequential Pyramid Training via Residual Flow Estimation

    algorithm

    The convnet models {G0,G1,…,GK}\{G_0, G_1, \dots, G_K\} are trained independently and sequentially from coarse to fine resolutions.

    Input: Training image pairs (I1,I2)(I^1, I^2) and ground truth optical flow fields V^\hat{V}
    Output: Trained sub-networks {G0,G1,…,GK}\{G_0, G_1, \dots, G_K\}
    Initialize G0G_0 randomly
    Train G0G_0 on coarsest resolution (I01,I02)(I_0^1, I_0^2) with target residual v^0=V^0\hat{v}_0 = \hat{V}_0
    for k=1k = 1 to KK do
        Initialize weights of GkG_k with the trained weights of Gk−1G_{k-1}
        Downsample images to obtain Ik1,Ik2I_k^1, I_k^2 and downsample flow to obtain V^k\hat{V}_k
        Compute coarse flow Vk−1V_{k-1} using frozen upstream networks {G0,…,Gk−1}\{G_0, \dots, G_{k-1}\}
        Compute target residual flow: v^k=V^k−u(Vk−1)\hat{v}_k = \hat{V}_k - u(V_{k-1})
        Train GkG_k to predict residual vkv_k minimizing LEPE(vk,v^k)\mathcal{L}_{EPE}(v_k, \hat{v}_k)
    end for

    The loss function minimized by each network GkG_k is the average End Point Error (EPE):

    LEPE(vk,v^k)=1mknk∑x,y(vkx(x,y)−v^kx(x,y))2+(vky(x,y)−v^ky(x,y))2\mathcal{L}_{EPE}(v_k, \hat{v}_k) = \frac{1}{m_k n_k} \sum_{x,y} \sqrt{\left(v_k^x(x,y) - \hat{v}_k^x(x,y)\right)^2 + \left(v_k^y(x,y) - \hat{v}_k^y(x,y)\right)^2}

    where mk×nkm_k \times n_k denotes the image spatial resolution at pyramid level kk, vk=(vkx,vky)v_k = (v_k^x, v_k^y) is the predicted residual flow, and v^k=(v^kx,v^ky)\hat{v}_k = (\hat{v}_k^x, \hat{v}_k^y) is the ground-truth residual flow.

  4. Knowl 4 — SPyNet Training Setup and Data Augmentation Protocol

    experimental setup

    SPyNet sub-networks {G0,…,G4}\{G_0, \dots, G_4\} are trained on the synthetic Flying Chairs dataset with input resolutions starting at 24×3224 \times 32 pixels for G0G_0 and doubling at each pyramid level up to 384×512384 \times 512 pixels for G4G_4.

    Hyperparameters and optimization configuration:

    • Optimizer: Adam (β1=0.9\beta_1 = 0.9, β2=0.999\beta_2 = 0.999)
    • Batch size: 32
    • Epoch size: 4000 iterations per epoch
    • Learning rate schedule: 10−410^{-4} for the first 60 epochs, reduced to 10−510^{-5} until convergence
    • Training duration: 3 days for G0G_0, and 1 day each for {G1,G2,G3,G4}\{G_1, G_2, G_3, G_4\} on a single NVIDIA Titan X GPU

    Data augmentation pipeline applied during training:

    1. Random spatial scaling by a factor uniformly sampled from [1,2][1, 2]
    2. Random spatial rotation sampled uniformly from [−17∘,17∘][-17^\circ, 17^\circ]
    3. Random cropping to match the resolution of level kk
    4. Additive zero-mean Gaussian noise with standard deviation sampled uniformly from N(0,0.1)\mathcal{N}(0, 0.1)
    5. Color jitter with additive brightness, contrast, and saturation perturbations sampled from N(0,0.4)\mathcal{N}(0, 0.4)
    6. Channel normalization using ImageNet dataset mean and standard deviation
  5. Knowl 5 — Optical Flow Benchmark Performance and Execution Speed

    data/table

    The table below evaluates average End Point Error (EPE in pixels) across standard optical flow benchmarks (MPI Sintel Clean and Final passes, KITTI, Middlebury, and Flying Chairs) and runtime per frame (measured on Flying Chairs excluding disk I/O). Models marked +ft are fine-tuned on benchmark-specific training splits. SPyNet+ft* includes additional fine-tuning data from synthetic driving scenes.

    Method Sintel Clean Sintel Final KITTI Middlebury Flying Chairs Time (s)
    train test train test train test train test test
    Classic+NLP 4.13 6.73 5.90 8.29 - - 0.22 0.32 3.93 102
    FlowNetS 4.50 7.42 5.45 8.43 8.26 - 1.09 - 2.71 0.080
    FlowNetC 4.31 7.28 5.87 8.81 9.35 - 1.15 - 2.19 0.150
    SPyNet 4.12 6.69 5.57 8.43 9.12 - 0.33 0.58 2.63 0.069
    FlowNetS+ft 3.66 6.96 4.44 7.76 7.52 9.1 0.98 - 3.04 0.080
    FlowNetC+ft 3.78 6.85 5.28 8.51 8.79 - 0.93 - 2.27 0.150
    SPyNet+ft 3.17 6.64 4.32 8.36 8.25 10.1 0.33 0.58 3.07 0.069
    SPyNet+ft* - - - - 3.36 4.1 - - - -

    Without fine-tuning, SPyNet achieves the lowest EPE among convnet models on Sintel Clean test (6.69) and Flying Chairs (2.63 vs 2.71 for FlowNetS). On Middlebury, SPyNet achieves an EPE of 0.58 on the test set, substantially outperforming FlowNet architectures (0.93 to 1.15 on train). SPyNet is also the fastest network evaluated at 0.069 seconds per frame.

  6. Knowl 6 — Performance Breakdown by Velocity and Motion Boundary Distance

    data/table

    The table below compares End Point Error (EPE) on the MPI Sintel benchmark partitioned by pixel displacement speed ranges (s0−10s_{0-10}, s10−40s_{10-40}, s40+s_{40+} in pixels per frame) and distance ranges from motion boundaries (d0−10d_{0-10}, d10−60d_{10-60}, d60−140d_{60-140} in pixels).

    Method Sintel Final Sintel Clean
    d0−10d_{0-10} d10−60d_{10-60} d60−140d_{60-140} s0−10s_{0-10} s10−40s_{10-40} s40+s_{40+} d0−10d_{0-10} d10−60d_{10-60} d60−140d_{60-140} s0−10s_{0-10} s10−40s_{10-40} s40+s_{40+}
    FlowNetS+ft 7.25 4.61 2.99 1.87 5.83 43.24 5.99 3.56 2.19 1.42 3.81 40.10
    FlowNetC+ft 7.19 4.62 3.30 2.30 6.17 40.78 5.57 3.18 1.99 1.62 3.97 33.37
    SPyNet+ft 6.69 4.37 3.29 1.39 5.53 49.71 5.50 3.12 1.71 0.83 3.34 43.44

    SPyNet+ft achieves lower EPE than FlowNet variants across small velocities (s0−10s_{0-10}: 1.39 on Final, 0.83 on Clean) and medium velocities (s10−40s_{10-40}: 5.53 on Final, 3.34 on Clean), as well as near motion boundaries (d0−10d_{0-10}: 6.69 on Final, 5.50 on Clean). Conversely, FlowNet models perform better on very large displacements (s40+s_{40+}: 40.78 on Final for FlowNetC+ft vs 49.71 for SPyNet+ft).

  7. Knowl 7 — Spatio-Temporal Filter Structures Learned in SPyNet

    empirical result

    Analysis of the first-layer convolutional filters in SPyNet sub-networks reveals properties that closely parallel classical spatio-temporal motion filters:

    1. Spatial derivative-like structure: Spatial filter weights operating on RGB channels resemble classical Gaussian derivative filters across multiple orientations and scales, as well as second-order derivative and Gabor filter profiles. The filters exhibit equal sensitivity across color channels, appearing largely grayscale.
    2. Temporal derivative structure: Subtracting the spatial filter weights corresponding to the reference image I1I^1 from those of the warped second image w(I2,⋅)w(I^2, \cdot) reveals an explicit temporal derivative structure.
    3. Evolution across pyramid levels: Filters initialized from coarser levels become sharper and increase in contrast as resolution increases towards level G4G_4, while higher-order frequency filters become increasingly defined.
    4. Contrast with full-frame networks: Because image warping at each pyramid level restricts the input displacement to small sub-pixel ranges (<1< 1 pixel), SPyNet learns structured, biologically plausible spatio-temporal filter geometries, in contrast to the high-frequency or random-appearing filters learned by full-frame networks like FlowNet.
  8. Knowl 8 — Fast-Motion and Coarse-to-Fine Limitations in SPyNet

    limitation

    SPyNet inherits fundamental limitations of classical coarse-to-fine spatial pyramid representations:

    1. Fast-moving small and thin objects: Objects that have spatial dimensions smaller than the pyramid downsampling factor disappear at coarse pyramid levels. When such objects undergo large motions, the coarse levels fail to detect the displacement, leading to unrecoverable errors at finer pyramid levels.
    2. Large displacements (>40> 40 pixels/frame): As demonstrated by performance breakdowns on Sintel, SPyNet incurs higher EPE on large velocities (s40+s_{40+}: 49.71 on Final) compared to end-to-end architectures with explicit correlation layers such as FlowNetC (40.78 on Final).
    3. Occlusion reasoning: Two-frame feedforward residual estimation lacks temporal history across multiple frames, making it difficult for the network to resolve occlusions and disocclusions without multi-frame context or explicit matching modules.

Coverage note — None was omitted; all core methodology, network architecture details, training procedures, quantitative benchmark results, analytical findings on learned filters, and stated model limitations are covered.

References

  1. 1.E. H. Adelson, C. H. Anderson, J. R. Bergen, P. J. Burt, and J. M. Ogden. Pyramid methods in image processing. RCA engineer, 29(6):33–41, 1984.
  2. 2.E. H. Adelson and J. R. Bergen. Spatiotemporal energy models for the perception of motion. J. Opt. Soc. Am. A, 2(2):284–299, Feb. 1985.
  3. 3.A. Ahmadi and I. Patras. Unsupervised convolutional neural networks for motion estimation. arXiv preprint arXiv:1601.06087, 2016.
  4. 4.S. Baker, D. Scharstein, J. Lewis, S. Roth, M. J. Black, and R. Szeliski. A database and evaluation methodology for optical flow. International Journal of Computer Vision, 92(1):1–31, 2011.
  5. 5.L. Bao, Q. Yang, and H. Jin. Fast edge-preserving PatchMatch for large displacement optical flow. Image Processing, IEEE Transactions on, 23(12):4996–5006, Dec 2014.
  6. 6.J. Barron, D. J. Fleet, and S. S. Beauchemin. Performance of optical flow techniques. Int. J. Comp. Vis. (IJCV), 12(1):43–77, 1994.
  7. 7.M. J. Black and P. Anandan. A framework for the robust estimation of optical flow. In Computer Vision, 1993. Proceedings., Fourth International Conference on, pages 231–236. IEEE, 1993.
  8. 8.T. Brox, C. Bregler, and J. Malik. Large displacement optical flow. In Computer Vision and Pattern Recognition, 2009. CVPR 2009. IEEE Conference on, pages 41–48. IEEE, 2009.
  9. 9.T. Brox, A. Bruhn, N. Papenberg, and J. Weickert. High accuracy optical flow estimation based on a theory for warping. In European conference on computer vision, pages 25–36. Springer, 2004.
  10. 10.P. J. Burt and E. H. Adelson. The Laplacian pyramid as a compact image code. IEEE Transactions on Communications, COM-34(4):532–540, 1983.
  11. 11.D. J. Butler, J. Wulff, G. B. Stanley, and M. J. Black. A naturalistic open source movie for optical flow evaluation. In A. Fitzgibbon et al. (Eds.), editor, European Conf. on Computer Vision (ECCV), Part IV, LNCS 7577, pages 611–625. Springer-Verlag, Oct. 2012.
  12. 12.C. Cadieu and B. A. Olshausen. Learning transformational invariants from natural movies. In Advances in neural information processing systems, pages 209–216, 2008.
  13. 13.L.-C. Chen, G. Papandreou, I. Kokkinos, K. Murphy, and A. L. Yuille. Semantic image segmentation with deep convolutional nets and fully connected crfs. arXiv preprint arXiv:1412.7062, 2014.
  14. 14.M. Chessa, N. K. Medathati, G. S. Masson, F. Solari, and P. Kornprobst. Decoding mt motion response for optical flow estimation: An experimental evaluation. In Signal Processing Conference (EUSIPCO), 2015 23rd European, pages 2241–2245. IEEE, 2015.
  15. 15.S. D., S. Roth, J. Lewis, and M. Black. Learning optical flow. In ECCV, pages 83–97, 2008.
  16. 16.E. L. Denton, S. Chintala, R. Fergus, et al. Deep generative image models using a laplacian pyramid of adversarial networks. In Advances in neural information processing systems, pages 1486–1494, 2015.
  17. 17.A. Dosovitskiy, P. Fischery, E. Ilg, C. Hazirbas, V. Golkov, P. van der Smagt, D. Cremers, T. Brox, et al. Flownet: Learning optical flow with convolutional networks. In 2015 IEEE International Conference on Computer Vision (ICCV), pages 2758–2766. IEEE, 2015.
  18. 18.W. T. Freeman, E. C. Pasztor, and O. T. Carmichael. Learning low-level vision. International Journal of Computer Vision, 40(1):25–47, 2000.
  19. 19.A. Geiger, P. Lenz, and R. Urtasun. Are we ready for autonomous driving? the KITTI vision benchmark suite. In Conference on Computer Vision and Pattern Recognition (CVPR), 2012.
  20. 20.F. Glazer. Hierarchical motion detection. PhD thesis, University of Massachusetts, Amherst, MA, 1987. COINS TR 87–02.
  21. 21.F. Guney and A. Geiger. Deep discrete flow. In Asian Conference on Computer Vision (ACCV), 2016.
  22. 22.S. Hauberg, A. Feragen, R. Enficiaud, and M. Black. Scalable robust principal component analysis using Grassmann averages. IEEE Trans. Pattern Analysis and Machine Intelligence (PAMI), Dec. 2015.
  23. 23.K. He, X. Zhang, S. Ren, and J. Sun. Deep residual learning for image recognition. arXiv preprint arXiv:1512.03385, 2015.
  24. 24.D. J. Heeger. Model for the extraction of image flow. J. Opt. Soc. Am, 4(8):1455–1471, Aug. 1987.
  25. 25.B. K. Horn and B. G. Schunck. Determining optical flow. 1981 Technical Symposium East, pages 319–331, 1981.
  26. 26.D. Kingma and J. Ba. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014.
  27. 27.T. Kroeger, R. Timofte, D. Dai, and L. V. Gool. Fast optical flow using dense inverse search. In Computer Vision – ECCV, 2016.
  28. 28.J. Long, E. Shelhamer, and T. Darrell. Fully convolutional networks for semantic segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 3431–3440, 2015.
  29. 29.R. Memisevic and G. E. Hinton. Learning to represent spatial transformations with factored higher-order boltzmann machines. Neural Computation, 22(6):1473–1492, 2010.
  30. 30.N.Mayer, E.Ilg, P.Hausser, P.Fischer, D.Cremers, A.Dosovitskiy, and T.Brox. A large dataset to train convolutional networks for disparity, optical flow, and scene flow estimation. In IEEE International Conference on Computer Vision and Pattern Recognition (CVPR), 2016. arXiv:1512.02134.
  31. 31.B. A. Olshausen. Learning sparse, overcomplete representations of time-varying natural images. In Image Processing, 2003. ICIP 2003. Proceedings. 2003 International Conference on, volume 1, pages I–41. IEEE, 2003.
  32. 32.J. Revaud, P. Weinzaepfel, Z. Harchaoui, and C. Schmid. EpicFlow: Edge-Preserving Interpolation of Correspondences for Optical Flow. In Computer Vision and Pattern Recognition, 2015.
  33. 33.S. Roth and M. J. Black. Fields of experts. International Journal of Computer Vision, 82(2):205–229, 2009.
  34. 34.O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, A. C. Berg, and L. Fei-Fei. ImageNet Large Scale Visual Recognition Challenge. International Journal of Computer Vision (IJCV), 115(3):211–252, 2015.
  35. 35.L. Sevilla-Lara, D. Sun, E. G. Learned-Miller, and M. J. Black. Optical flow estimation with channel constancy. In Computer Vision – ECCV 2014, volume 8689 of Lecture Notes in Computer Science, pages 423–438. Springer International Publishing, Sept. 2014.
  36. 36.E. P. Simoncelli and D. J. Heeger. A model of neuronal responses in visual area MT. Vision Res., 38(5):743–761, 1998.
  37. 37.F. Solari, M. Chessa, N. K. Medathati, and P. Kornprobst. What can we expect from a v1-mt feedforward architecture for optical flow estimation? Signal Processing: Image Communication, 39:342–354, 2015.
  38. 38.D. Sun, S. Roth, and M. J. Black. A quantitative analysis of current practices in optical flow estimation and the principles behind them. International Journal of Computer Vision, 106(2):115–137, 2014.
  39. 39.N. Sundaram, T. Brox, and K. Keutzer. Dense point trajectories by gpu-accelerated large displacement optical flow. In European conference on computer vision, pages 438–451. Springer, 2010.
  40. 40.C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich. Going deeper with convolutions. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 1–9, 2015.
  41. 41.G. W. Taylor, R. Fergus, Y. LeCun, and C. Bregler. Convolutional learning of spatio-temporal features. In European conference on computer vision, pages 140–153. Springer, 2010.
  42. 42.D. Tran, L. Bourdev, R. Fergus, L. Torresani, and M. Paluri. Deep End2End Voxel2Voxel prediction. In The 3rd Workshop on Deep Learning in Computer Vision, 2016.
  43. 43.L. van Hateren and J. Ruderman. Independent component analysis of natural image sequences yields spatio-temporal filters similar to simple cells in primary visual cortex. Proceedings: Biological Sciences, 265(1412):23152320, 1998.
  44. 44.P. Weinzaepfel, J. Revaud, Z. Harchaoui, and C. Schmid. Deepflow: Large displacement optical flow with deep matching. In Proceedings of the IEEE International Conference on Computer Vision, pages 1385–1392, 2013.
  45. 45.M. Werlberger, W. Trobin, T. Pock, A. Wedel, D. Cremers, and H. Bischof. Anisotropic Huber-L1 optical flow. In BMVC, London, UK, September 2009.
  46. 46.J. Wulff and M. J. Black. Efficient sparse-to-dense optical flow estimation using a learned basis and layers. In 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 120–130. IEEE, 2015.
  47. 47.J. J. Yu, A. W. Harley, and K. G. Derpanis. Back to basics: Unsupervised learning of optical flow via brightness constancy and motion smoothness. arXiv preprint arXiv:1608.05842, 2016.

Citation

MLA
Ranjan, A., and M. J. Black. “Optical Flow Estimation Using a Spatial Pyramid Network”. arXiv, 2016, http://arxiv.org/abs/1611.00850v2.
APA
Ranjan, A., & Black, M. J. (2016). Optical Flow Estimation using a Spatial Pyramid Network. arXiv. http://arxiv.org/abs/1611.00850v2
Chicago
Ranjan, A., and M. J. Black. 2016. “Optical Flow Estimation Using a Spatial Pyramid Network”. arXiv. http://arxiv.org/abs/1611.00850v2.
Harvard
Ranjan, A. and Black, M.J. (2016) “Optical Flow Estimation using a Spatial Pyramid Network”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1611.00850v2.
Vancouver
1. Ranjan A, Black MJ (2016) Optical Flow Estimation using a Spatial Pyramid Network. arXiv

BibTeX

@article{ranjan2016optical,
  title = {Optical Flow Estimation using a Spatial Pyramid Network},
  author = {Ranjan, Anurag and Black, Michael J.},
  year = {2016},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1611.00850v2},
  eprint = {1611.00850}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE