SIFT Flow: Dense Correspondence across Scenes and Its Applications

Ce LiuJenny YuenAntonio Torralba

article2011TPAMI1,785 citations

Introduces SIFT flow, a technique that matches dense local feature descriptors across different complex scenes to establish semantically meaningful pixel-level correspondences, enabling tasks such as motion prediction from a single still image and cross-scene information transfer.

Listen

Aligning visual data across distinct scenes remains a major challenge in computer vision because images depicting different physical environments often vary significantly in object appearance, scale, viewpoint, and spatial arrangement. Traditional alignment techniques excel at matching identical scenes across frames or specific object instances, but they fail when required to establish semantically meaningful, pixel-level correspondences across entirely different scenes.

The article introduces and evaluates SIFT flow, an algorithm designed to achieve dense, pixel-to-pixel correspondence across diverse scenes by matching local structural descriptors within a large-scale database framework.

To accomplish this, the authors adapt optical flow concepts by calculating Scale-Invariant Feature Transform (SIFT) descriptors at every pixel rather than relying on raw brightness values. The algorithm matches these dense descriptors across nearest-neighbor images retrieved from a database of over 100,000 video frames, utilizing a decoupled dual-layer belief propagation method with a coarse-to-fine optimization scheme to enforce spatial smoothness and manage large spatial displacements.

The evaluation yielded several key findings. First, the coarse-to-fine matching approach dramatically improves computational efficiency, reducing the processing time for a standard image pair from approximately 127 minutes to 31 seconds while achieving equal or lower energy solutions. Second, a user validation study revealed that dense SIFT flow matches human perceptual alignment significantly better than direct feature matching without spatial regularization. Third, when applied to predicting motion from a single static image, the framework ranked the correct plausible motion first in over 50% of general scenes and in 66% of specialized street scenes. Finally, the method successfully transferred moving objects across scenes and outperformed standard sparse-feature techniques in challenging same-scene satellite image registration, reducing alignment error from 0.030 to 0.021 while remaining competitive in face recognition benchmarks with limited training data.

These findings demonstrate that dense scene alignment enables effective nonparametric data transfer—including motion, labels, and geometry—from repository images to novel query images without requiring explicit object recognition models. This capability significantly reduces the need for extensive task-specific training data in applications ranging from image synthesis and animation to satellite surveillance and biometric identification.

Organizations seeking to implement dense scene alignment should adopt the coarse-to-fine optimization pipeline to balance quality and computational load. For production environments requiring faster throughput, developing hardware-accelerated implementations on graphical processing units represents the primary path forward to reduce runtime further.

The primary limitation of this approach is its fundamental dependence on database density; if a query image has no semantically similar counterpart in the reference repository, the alignment will produce incorrect correspondences. While confidence is high in the algorithm's effectiveness for densely sampled domains, users should exercise caution when applying the method to rare or highly atypical imagery without expanding the underlying image library.

Cover for SIFT Flow: Dense Correspondence across Scenes and Its Applications

Abstract

While image alignment has been studied in different areas of computer vision for decades, aligning images depicting different scenes remains a challenging problem. Analogous to optical flow where an image is aligned to its temporally adjacent frame, we propose SIFT flow, a method to align an image to its nearest neighbors in a large image corpus containing a variety of scenes. The SIFT flow algorithm consists of matching densely sampled, pixel-wise SIFT features between two images, while preserving spatial discontinuities. The SIFT features allow robust matching across different scene/object appearances, whereas the discontinuity-preserving spatial model allows matching of objects located at different parts of the scene. Experiments show that the proposed approach robustly aligns complex scene pairs containing significant spatial differences. Based on SIFT flow, we propose an alignment-based large database framework for image analysis and synthesis, where image information is transferred from the nearest neighbors to a query image according to the dense scene correspondence. This framework is demonstrated through concrete applications, such as motion field prediction from a single image, motion synthesis via object transfer, satellite image registration and face recognition.

Table of Contents

  • 1 INTRODUCTION
  • 2 RELATED WORK
  • 3 THE SIFT FLOW ALGORITHM
  • 3.1 Dense SIFT descriptors and visualization
  • 3.2 Matching Objective
  • 3.3 Coarse-to-fine matching scheme
  • 3.4 Neighborhood of SIFT flow
  • 4 EXPERIMENTS ON VIDEO RETRIEVAL
  • 4.1 Results of video retrieval
  • 4.2 Evaluation of the dense scene alignment
  • 5 DENSE SCENE ALIGNMENT APPLICATIONS
  • 5.1 Predicting motion fields from a single image
  • 5.2 Quantitative evaluation of motion prediction
  • 5.3 Motion synthesis via object transfer
  • 6 EXPERIMENTS ON IMAGE ALIGNMENT AND FACE RECOGNITION
  • 6.1 Same-scene image registration
  • 6.2 Face recognition
  • 7 DISCUSSIONS
  • I. Scene alignment
  • II. Sparse vs. Dense correspondence
  • III. SIFT flow vs. optical flow
  • 8 CONCLUSION
  • 9 ACKNOWLEDGMENTS
  • REFERENCES

Knowls

  1. Knowl 1 — SIFT Flow Objective Function

    equation

    The SIFT flow objective function estimates a dense, discrete displacement field w(p)=(u(p),v(p))∈Z2\mathbf{w}(\mathbf{p}) = (u(\mathbf{p}), v(\mathbf{p})) \in \mathbb{Z}^2 between a source SIFT image s1\mathbf{s}_1 and a target SIFT image s2\mathbf{s}_2 across an image grid coordinate system p=(x,y)\mathbf{p} = (x, y). The objective function is defined as:

    E(w)=∑pmin⁡(∥s1(p)−s2(p+w(p))∥1, t)+∑pη(∣u(p)∣+∣v(p)∣)+∑(p,q)∈E(min⁡(α∣u(p)−u(q)∣, d)+min⁡(α∣v(p)−v(q)∣, d))E(\mathbf{w}) = \sum_{\mathbf{p}} \min\left( \|\mathbf{s}_1(\mathbf{p}) - \mathbf{s}_2(\mathbf{p} + \mathbf{w}(\mathbf{p}))\|_1, \, t \right) + \sum_{\mathbf{p}} \eta \left( |u(\mathbf{p})| + |v(\mathbf{p})| \right) + \sum_{(\mathbf{p}, \mathbf{q}) \in \mathcal{E}} \left( \min(\alpha |u(\mathbf{p}) - u(\mathbf{q})|, \, d) + \min(\alpha |v(\mathbf{p}) - v(\mathbf{q})|, \, d) \right)

    where:

    • s1(p),s2(p+w(p))∈R128\mathbf{s}_1(\mathbf{p}), \mathbf{s}_2(\mathbf{p} + \mathbf{w}(\mathbf{p})) \in \mathbb{R}^{128} are 128-dimensional dense SIFT descriptors extracted at pixel locations p\mathbf{p} and p+w(p)\mathbf{p} + \mathbf{w}(\mathbf{p}).
    • E\mathcal{E} denotes the set of 4-connected spatial neighbor pairs in the image lattice.
    • t>0t > 0 is a scalar truncation threshold on the L1L_1 descriptor data term to account for outliers and occlusions.
    • η≥0\eta \ge 0 is a regularization coefficient penalizing displacement magnitudes when visual evidence is ambiguous.
    • α>0\alpha > 0 is the spatial smoothness weight.
    • d>0d > 0 is a scalar truncation threshold on the smoothness term to preserve sharp flow discontinuities at object and scene boundaries.
    • u(p)u(\mathbf{p}) and v(p)v(\mathbf{p}) take discrete values from a set of LL possible displacement states.
  2. Knowl 2 — Dense SIFT Image Representation

    definition

    A SIFT image is a dense, per-pixel feature map where each pixel on an image grid is represented by a 128-dimensional local gradient vector. For every pixel location p=(x,y)\mathbf{p}=(x, y), a 16×1616 \times 16 pixel neighborhood centered at p\mathbf{p} is partitioned into a 4×44 \times 4 array of spatial cells. Within each cell, gradient orientations are computed and accumulated into an 8-bin histogram. Concatenating the orientation histograms across all 4×44 \times 4 cells produces a local descriptor of dimension 4×4×8=1284 \times 4 \times 8 = 128 at each pixel, capturing local image structure, edge orientation, and context while preserving spatial boundaries.

  3. Knowl 3 — Coarse-to-Fine SIFT Flow Matching Algorithm

    algorithm

    The coarse-to-fine SIFT flow algorithm computes dense correspondence between two SIFT images s1\mathbf{s}_1 and s2\mathbf{s}_2 across a multi-resolution pyramid {s(k)}k=1K\{\mathbf{s}^{(k)}\}_{k=1}^K, reducing the optimization time complexity from O(h4)O(h^4) to O(h2log⁡h)O(h^2 \log h) for an h×hh \times h image.

    Input: Dense SIFT images s1,s2\mathbf{s}_1, \mathbf{s}_2, number of pyramid levels KK (e.g., K=3K=3), top window size m×mm \times m, fine search window size n×nn \times n (with n=11n=11), parameters α,d,η,t\alpha, d, \eta, t.
    Output: Optimal dense displacement field w(p1)=(u(p1),v(p1))\mathbf{w}(\mathbf{p}_1) = (u(\mathbf{p}_1), v(\mathbf{p}_1)) for all pixels p1\mathbf{p}_1 in s1(1)\mathbf{s}_1^{(1)}.
    1. Construct pyramids {s1(k)}k=1K\{\mathbf{s}_1^{(k)}\}_{k=1}^K and {s2(k)}k=1K\{\mathbf{s}_2^{(k)}\}_{k=1}^K by iterative smoothing and downsampling.
    2. Initialize top level k←Kk \leftarrow K with search windows centered at cK=pK\mathbf{c}_K = \mathbf{p}_K of size m×mm \times m.
    3. Run dual-layer sequential belief propagation (BP-S) to compute w(pK)\mathbf{w}(\mathbf{p}_K) minimizing the SIFT flow energy.
    4. for k=K−1k = K-1 down to 1 do
           a. For each pixel pk\mathbf{p}_k, interpolate the displacement from level k+1k+1 to define the search window centroid ck\mathbf{c}_k.
           b. Set search window size to n×nn \times n centered at ck\mathbf{c}_k.
           c. Update displacement penalty η←2η\eta \leftarrow 2 \eta while keeping α\alpha and dd constant.
           d. Execute dual-layer BP-S using generalized distance transforms across neighboring nodes with shifted search window offsets.
           e. Obtain updated flow vectors w(pk)\mathbf{w}(\mathbf{p}_k).
       end for
    5. return w(p1)\mathbf{w}(\mathbf{p}_1).
  4. Knowl 4 — Decoupled Dual-Layer Belief Propagation for SIFT Flow

    model/method

    The discrete MRF energy of SIFT flow is optimized using a dual-layer loopy belief propagation architecture. The horizontal flow field u(p)u(\mathbf{p}) and vertical flow field v(p)v(\mathbf{p}) are decoupled into two distinct 2D grid layers sharing the same spatial coordinates, linked together only by the data term at corresponding locations. Message passing operates by first updating intra-layer spatial messages along horizontal and vertical neighbors separately within each layer, followed by updating inter-layer messages between u(p)u(\mathbf{p}) and v(p)v(\mathbf{p}).

    Because the pairwise potential functions are truncated L1L_1 norms, distance transform techniques are generalized to compute spatial messages across neighboring nodes with offset search windows. For state space size LL per displacement component, this decoupling reduces the per-iteration message passing time complexity from O(L4)O(L^4) to O(L2)O(L^2), with sequential belief propagation (BP-S) applied to accelerate convergence.

  5. Knowl 5 — Alignment-Based Large Database Framework for Image Analysis and Synthesis

    model/method

    The alignment-based large database framework operates on the premise that dense sampling of the space of world images allows semantically meaningful scene alignment between different 3D scenes sharing common structures, directly analogous to dense temporal sampling in video optical flow.

    Given an input query image qq and a large database of images or video frames D\mathcal{D}:

    1. Candidate Retrieval: A fast indexing mechanism (such as spatial pyramid visual word histogram matching or global GIST matching) retrieves the top NN nearest neighbors from D\mathcal{D}.
    2. Dense Alignment via SIFT Flow: SIFT flow is computed between qq and each retrieved candidate, re-ranking the candidates by their minimized SIFT flow matching energy E(w)E(\mathbf{w}).
    3. Information Transfer: Associated metadata, annotations, depth maps, labels, or temporal motion fields belonging to the optimal matches are warped onto the coordinate frame of qq using the estimated dense pixel displacement fields.
  6. Knowl 6 — Motion Field Prediction from a Single Image

    model/method

    SIFT flow enables hallucinating plausible temporal motion fields for a static single image by transferring dynamic motion information from visually similar video sequences. Given a still query image, the top matching video frames are retrieved from a video database using quantized SIFT visual-word spatial histograms. Dense correspondence between the static query and the retrieved database frames is computed via SIFT flow. The temporal motion fields of the retrieved video frames (estimated between consecutive frames in the source video) are then warped into the query image coordinate frame.

    To quantitatively evaluate predicted motion against distracters, motion fields are binarized into grids M,N∈{0,1}H×WM, N \in \{0, 1\}^{H \times W} representing moving (11) versus static (00) pixels. The similarity between two motion segmentations is measured by:

    S(M,N)=∑(x,y)∈G[M(x,y)=N(x,y)]S(M, N) = \sum_{(x, y) \in G} [M(x, y) = N(x, y)]

    When evaluating against 11 random motion distracters across 720 test scenes, the warped motion field inferred via SIFT flow ranked 1st in over 50% of all video scenes and achieved 66% top-1 rank precision on street/car scenes.

  7. Knowl 7 — Motion Synthesis via Moving Object Transfer

    algorithm

    Motion synthesis animates a still image qq by detecting and compositing moving objects from video sequences D\mathcal{D} with matching scene layouts.

    Input: Static query image qq, video database D\mathcal{D}, motion magnitude threshold TT.
    Output: Synthesized video sequence G={g1,g2,…,gK}G = \{g_1, g_2, \dots, g_K\}.
    1. Query database D\mathcal{D} using SIFT-based scene matching to retrieve candidate video frames F={fi}i=1KF = \{f_i\}_{i=1}^K.
    2. for each frame fi∈Ff_i \in F do
           a. Compute temporal motion field mim_i between consecutive frames fif_i and fi+1f_{i+1}.
           b. Perform motion segmentation to obtain the binary foreground mask zi=[∣mi∣>T]z_i = [|m_i| > T].
           c. Align fif_i and mask ziz_i to query image qq using SIFT flow.
           d. Blend the moving foreground into the query background using Poisson image editing:
              gi=PoissonBlend(q,fi,zi)g_i = \text{PoissonBlend}(q, f_i, z_i).
       end for
    3. return G={g1,g2,…,gK}G = \{g_1, g_2, \dots, g_K\}.
  8. Knowl 8 — Face Recognition Performance Under Limited Training Samples

    data/table

    SIFT flow was applied to face recognition on the ORL database (40 subjects, 10 images per subject with pose and expression variations) to establish dense alignment between query test images and database templates before nearest neighbor classification. When very few training images are available per subject, dense alignment compensates for pose and expression differences.

    Method 1 Train Sample 2 Train Samples 3 Train Samples
    S-LDA N/A 17.1±2.717.1 \pm 2.7 8.1±1.88.1 \pm 1.8
    SIFT flow 28.4±3.028.4 \pm 3.0 16.6±2.216.6 \pm 2.2 8.9±2.18.9 \pm 2.1

    The table reports test error rates (mean ±\pm standard deviation in percent) across 1, 2, and 3 training samples per subject comparing SIFT flow against Spatially Smooth Subspace Learning (S-LDA). SIFT flow achieves performance comparable to the state-of-the-art without requiring explicit facial feature point detection or component-specific alignment models.

  9. Knowl 9 — Multi-Temporal Satellite Image Registration Performance

    empirical result

    SIFT flow was evaluated on multi-temporal satellite image registration of the Martian surface using image pairs captured four years apart under severe variations in illumination, sensor conditions, and surface appearances. Comparing SIFT flow against a baseline of sparse SIFT keypoint detection and matching followed by interpolation to a dense displacement field:

    • Sparse SIFT feature matching with interpolation achieved a pixel-wise Mean Absolute Error (MAE) of 0.0300.030.
    • Dense SIFT flow achieved an MAE of 0.0210.021, reducing registration error by 30%30\% and resolving local image distortions and stitching discontinuities.
  10. Knowl 10 — Perceptual Alignment Accuracy Relative to Human Ground Truth

    empirical result

    In a user study with 11 human subjects annotating 10 preselected sparse correspondence points across complex cross-scene image pairs, SIFT flow was evaluated against direct minimum L1L_1-norm SIFT matching without spatial regularization. Performance was assessed by computing Pr⁡(∃zi,∥zi−w(p)∥≤r)\Pr(\exists z_i, \|z_i - \mathbf{w}(\mathbf{p})\| \le r), the probability that at least one human annotation ziz_i lies within radius rr (in pixels) of the computed flow vector w(p)\mathbf{w}(\mathbf{p}). Across all tested radii rr, SIFT flow achieved significantly higher correspondence agreement with human perceptual annotations than independent minimum L1L_1 SIFT matching, demonstrating that the spatial smoothness and small displacement regularizers are essential for meaningful semantic matching.

Coverage note — None was omitted; all primary contributions, algorithm formulations, mathematical models, and quantitative applications (motion prediction, motion synthesis, satellite image registration, and face recognition) are covered.

References

  1. 1.S. Avidan. Ensemble tracking. IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), 29(2):261–271, 2007.
  2. 2.S. Baker, D. Scharstein, J. P. Lewis, S. Roth, M. J. Black, and R. Szeliski. A database and evaluation methodology for optical flow. In Proc. ICCV, 2007.
  3. 3.J. L. Barron, D. J. Fleet, and S. S. Beauchemin. Systems and experiment performance of optical flow techniques. International Journal of Computer Vision (IJCV), 12(1):43–77, 1994.
  4. 4.S. Belongie, J. Malik, and J. Puzicha. Shape context: A new descriptor for shape matching and object recognition. In Advances in Neural Information Processing Systems (NIPS), 2000.
  5. 5.S. Belongie, J. Malik, and J. Puzicha. Shape matching and object recognition using shape contexts. IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), 24(4):509–522, 2002.
  6. 6.A. Berg, T. Berg., and J. Malik. Shape matching and object recognition using low distortion correspondence. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2005.
  7. 7.J. R. Bergen, P. Anandan, K. J Hanna, and R. Hingorani. Hierarchical model-based motion estimation. In European Conference on Computer Vision (ECCV), pages 237–252, 1992.
  8. 8.M. J. Black and P. Anandan. The robust estimation of multiple motions: parametric and piecewise-smooth flow fields. Computer Vision and Image Understanding, 63(1):75–104, January 1996.
  9. 9.Y. Boykov, O. Veksler, and R. Zabih. Fast approximate energy minimization via graph cuts. IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), 23(11):1222–1239, 2001.
  10. 10.T. Brox, C. Bregler, and J. Malik. Large displacement optical flow. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2009.
  11. 11.T. Brox, A. Bruhn, N. Papenberg, and J. Weickert. High accuracy optical flow estimation based on a theory for warping. In European Conference on Computer Vision (ECCV), pages 25–36, 2004.
  12. 12.A. Bruhn, J. Weickert, and C. Schn¨orr. Lucas/Kanade meets Horn/Schunk: combining local and global optical flow methods. International Journal of Computer Vision (IJCV), 61(3):211–231, 2005.
  13. 13.D. Cai, X. He, Y. Hu, J. Han, and T. Huang. Learning a spatially smooth subspace for face recognition. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2007.
  14. 14.C. Carson, S. Belongie, H. Greenspan, and J. Malik. Blobworld: Color- and Texture-Based Image Segmentation Using EM and Its Application to Image Querying and Classification. IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), 24(8):1026–1038, 2002.
  15. 15.T. F. Cootes, G. J. Edwards, and C. J. Taylor. Active appearance models. In European Conference on Computer Vision (ECCV), volume 2, pages 484–498, 1998.
  16. 16.N. Cornelis and L. V. Gool. Real-time connectivity constrained depth map computation using programmable graphics hardware. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 1099–1104, 2005.
  17. 17.N. Dalal and B. Triggs. Histograms of oriented gradients for human detection. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2005.
  18. 18.L. Fei-Fei and P. Perona. A bayesian hierarchical model for learning natural scene categories. In cvpr, volume 2, pages 524–531, 2005.
  19. 19.P. Felzenszwalb and D. Huttenlocher. Pictorial structures for object recognition. International Journal of Computer Vision (IJCV), 61(1), 2005.
  20. 20.P. F. Felzenszwalb and D. P. Huttenlocher. Efficient belief propagation for early vision. International Journal of Computer Vision (IJCV), 70(1):41–54, 2006.
  21. 21.D. J. Fleet, A. D. Jepson, and M. R. M. Jenkin. Phase-based disparity measurement. Computer Vision, Graphics and Image Processing (CVGIP), 53(2):198–210, 1991.
  22. 22.W. T. Freeman, E. C. Pasztor, and O. T. Carmichael. Learning low-level vision. International Journal of Computer Vision (IJCV), 40(1):25–47, 2000.
  23. 23.M. M. Gorkani and R. W. Picard. Texture orientation for sorting photos at a glance. In IEEE International Conference on Pattern Recognition (ICPR), volume 1, pages 459–464, 1994.
  24. 24.K. Grauman and T. Darrell. Pyramid match kernels: Discriminative classification with sets of image features. In IEEE International Conference on Computer Vision (ICCV), 2005.
  25. 25.W. E. L. Grimson. Computational experiments with a feature based stereo algorithm. IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), 7(1):17–34, 1985.
  26. 26.M. J. Hannah. Computer Matching of Areas in Stereo Images. PhD thesis, Stanford University, 1974.
  27. 27.C. Harris and M. Stephens. A combined corner and edge detector. In Proceedings of the 4th Alvey Vision Conference, pages 147–151, 1988.
  28. 28.J. Hays and A. A Efros. Scene completion using millions of photographs. ACM SIGGRAPH, 26(3), 2007.
  29. 29.B. K. P. Horn and B. G. Schunck. Determing optical flow. Artificial Intelligence, 17:185–203, 1981.
  30. 30.D. G. Jones and J. Malik. A computational framework for determining stereo correspondence from a set of linear spatial filters. In European Conference on Computer Vision (ECCV), pages 395–410, 1992.
  31. 31.V. Kolmogorov and R. Zabih. Computing visual correspondence with occlusions using graph cuts. In IEEE International Conference on Computer Vision (ICCV), pages 508–515, 2001.
  32. 32.S. Lazebnik, C. Schmid, and J. Ponce. Beyond bags of features: Spatial pyramid matching for recognizing natural scene categories. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), volume II, pages 2169–2178, 2006.
  33. 33.C. Liu, W. T. Freeman, and E. H. Adelson. Analysis of contour motions. In Advances in Neural Information Processing Systems (NIPS), 2006.
  34. 34.C. Liu, W. T. Freeman, E. H. Adelson, and Y. Weiss. Human-assisted motion annotation. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2008.
  35. 35.C. Liu, J. Yuen, and A. Torralba. Nonparametric scene parsing: Label transfer via dense scene alignment. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2009.
  36. 36.C. Liu, J. Yuen, A. Torralba, J. Sivic, and W. T. Freeman. SIFT flow: dense correspondence across different scenes. In European Conference on Computer Vision (ECCV), 2008.
  37. 37.D. G. Lowe. Object recognition from local scale-invariant features. In IEEE International Conference on Computer Vision (ICCV), pages 1150–1157, Kerkyra, Greece, 1999.
  38. 38.B. Lucas and T. Kanade. An iterative image registration technique with an application to stereo vision. In Proceedings of the International Joint Conference on Artificial Intelligence, pages 674–679, 1981.
  39. 39.A. Oliva and A. Torralba. Modeling the shape of the scene: a holistic representation of the spatial envelope. International Journal of Computer Vision (IJCV), 42(3):145–175, 2001.
  40. 40.P. P´erez, M. Gangnet, and A. Blake. Poisson image editing. ACM SIGGRAPH, 22(3):313–318, 2003.
  41. 41.C. Rother, T. Minka, A. Blake, and V. Kolmogorov. Cosegmentation of image pairs by histogram matching - incorporating a global constraint into mrfs. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), volume 1, pages 993–1000, 2006.
  42. 42.B. C. Russell, A. Torralba, C. Liu, R. Fergus, and W. T. Freeman. Object recognition by scene alignment. In Advances in Neural Information Processing Systems (NIPS), 2007.
  43. 43.B. C. Russell, A. Torralba, K. P. Murphy, and W. T. Freeman. LabelMe: a database and web-based tool for image annotation. International Journal of Computer Vision (IJCV), 77(1-3):157–173, 2008.
  44. 44.F. Samaria and A. Harter. Parameterization of a stochastic model for human face identification. In IEEE Workshop on Applications of Computer Vision, 1994.
  45. 45.D. Scharstein and R. Szeliski. A taxonomy and evaluation of dense two-frame stereo correspondence algorithms. International Journal of Computer Vision (IJCV), 47(1):7–42, 2002.
  46. 46.C. Schmid, R. Mohr, and C. Bauckhage. Evaluation of interest point detectors. International Journal of Computer Vision (IJCV), 37(2):151–172, 2000.
  47. 47.A. Shekhovtsov, I. Kovtun, and V. Hlavac. Efficient MRF deformation model for non-rigid image matching. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2007.
  48. 48.J. Sivic and A. Zisserman. Video Google: a text retrieval approach to object matching in videos. In IEEE International Conference on Computer Vision (ICCV), 2003.
  49. 49.J. Sun, N. Zheng, , and H. Shum. Stereo matching using belief propagation. IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), 25(7):787–800, 2003.
  50. 50.M. J. Swain and D. H. Ballard. Color indexing. International Journal of Computer Vision (IJCV), 7(1), 1991.
  51. 51.R. Szeliski. Image alignment and stiching: A tutorial. Foundations and Trends in Computer Graphics and Computer Vision, 2(1), 2006.
  52. 52.R. Szeliski, R. Zabih, D. Scharstein, O. Veksler, V. Kolmogorov, A. Agarwala, M. Tappen, and C. Rother. A comparative study of energy minimization methods for markov random fields with smoothness-based priors. IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), 30(6):1068–1080, 2008.
  53. 53.A. Torralba, R. Fergus, and W. T. Freeman. 80 million tiny images: a large dataset for non-parametric object and scene recognition. IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), 30(11):1958–1970, 2008.
  54. 54.P. Viola and W. Wells III. Alignment by maximization of mutual information. In IEEE International Conference on Computer Vision (ICCV), pages 16–23, 1995.
  55. 55.Y. Weiss. Interpreting images by propagating bayesian beliefs. In Advances in Neural Information Processing Systems (NIPS), pages 908–915, 1997.
  56. 56.Y. Weiss. Smoothness in layers: Motion segmentation using nonparametric mixture estimation. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 520–527, 1997.
  57. 57.J. Winn and N. Jojic. Locus: Learning object classes with unsupervised segmentation. In IEEE International Conference on Computer Vision (ICCV), pages 756–763, 2005.
  58. 58.G. Yang, C. V. Stewart, M. Sofka, and C.-L. Tsai. Registration of challenging image pairs: Initialization, estimation, and decision. IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), 29(11):1973–1989, 2007.

Citation

MLA
Ce Liu, et al. “SIFT Flow: Dense Correspondence Across Scenes and Its Applications”. IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 33, no. 5, 2011, pp. 978–94, https://doi.org/10.1109/TPAMI.2010.147.
APA
Ce Liu, Yuen, J., & Torralba, A. (2011). SIFT Flow: Dense Correspondence across Scenes and Its Applications. IEEE Transactions on Pattern Analysis and Machine Intelligence, 33(5), 978–994. https://doi.org/10.1109/TPAMI.2010.147
Chicago
Ce Liu, J. Yuen, and A. Torralba. 2011. “SIFT Flow: Dense Correspondence Across Scenes and Its Applications”. IEEE Transactions on Pattern Analysis and Machine Intelligence 33 (5): 978–94. https://doi.org/10.1109/TPAMI.2010.147.
Harvard
Ce Liu, Yuen, J. and Torralba, A. (2011) “SIFT Flow: Dense Correspondence across Scenes and Its Applications”, IEEE Transactions on Pattern Analysis and Machine Intelligence, 33(5), pp. 978–994. Available at: https://doi.org/10.1109/TPAMI.2010.147.
Vancouver
1. Ce Liu, Yuen J, Torralba A (2011) SIFT Flow: Dense Correspondence across Scenes and Its Applications. IEEE Transactions on Pattern Analysis and Machine Intelligence 33:978–994

BibTeX

@article{Ce_Liu_2011, title={SIFT Flow: Dense Correspondence across Scenes and Its Applications}, volume={33}, ISSN={2160-9292}, url={http://dx.doi.org/10.1109/TPAMI.2010.147}, DOI={10.1109/tpami.2010.147}, number={5}, journal={IEEE Transactions on Pattern Analysis and Machine Intelligence}, publisher={Institute of Electrical and Electronics Engineers (IEEE)}, author={Ce Liu and Yuen, Jenny and Torralba, Antonio}, year={2011}, month=May, pages={978–994} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF