GMFlow: Learning Optical Flow via Global Matching

Haofei XuJing ZhangJianfei CaiHamid RezatofighiDacheng Tao

article2022CVPR573 citations

Proposes a global matching framework that reformulates optical flow estimation via Transformer-based feature matching and attention propagation, outperforming RAFT on the Sintel benchmark while running faster with only a single refinement step.

Listen

Optical flow estimation—the task of tracking pixel-level motion across consecutive video frames—is a foundational technology for computer vision applications such as autonomous driving and video processing. Prevailing deep learning architectures typically estimate flow by applying convolutional neural networks over localized search regions. However, this local formulation struggles with large pixel displacements and requires iterative, multi-step refinement loops (such as the 31 iterations used by the state-of-the-art RAFT model) to approximate global movement. These sequential iterations significantly increase runtime and computing overhead, limiting real-time operational deployment.

The article demonstrates that optical flow can be reformulated as a direct global feature-matching problem rather than a local regression task. The primary objective is to evaluate whether an attention-based matching architecture, termed GMFlow, can achieve superior accuracy and computational efficiency over iterative convolutional methods while using minimal refinement steps.

The authors developed the GMFlow architecture comprising three core modules: an attention-based feature enhancer (Transformer) incorporating cross-attention and local window mechanisms, a differentiable global matching layer that computes all pairwise correlations via matrix multiplication, and a self-attention propagation layer that fills in occluded or out-of-boundary regions using visual structure cues. A single residual refinement step at higher resolution reuses these shared components. The system was trained on standard synthetic datasets (FlyingChairs and FlyingThings3D) and benchmarked against leading methods across synthetic and real-world datasets, including Sintel and KITTI.

Key findings show that GMFlow fundamentally outperforms traditional multi-step pipelines. First, on the challenging Sintel clean benchmark, GMFlow with only one refinement step achieved an average endpoint error of 1.08 pixels, outperforming RAFT with 31 refinements (1.41 pixels) while executing significantly faster (66 ms versus 91 ms on modern hardware). Second, the global formulation resolves large-displacement motion far more effectively, reducing error on large movements (over 40 pixels) from 40.48 to 8.97 pixels prior to refinement. Third, the self-attention propagation mechanism substantially reduced error in unmatched and occluded regions from 15.54 to 10.39 pixels on Sintel clean. Finally, the framework enables bidirectional motion estimation at inference time by transposing the correlation matrix, eliminating redundant computational passes.

These findings suggest that global matching offers a practical path toward high-accuracy, low-latency motion tracking. By eliminating heavy iterative loops, the framework is well-suited to modern hardware acceleration and scalable deployment in latency-sensitive environments, potentially reducing operational compute costs. However, evaluations on the real-world KITTI dataset revealed lower performance than convolutional baselines, indicating that attention mechanisms require larger, more diverse datasets to overcome the domain gap between synthetic training data and real-world environments.

Organizations evaluating this approach for practical deployment should prioritize its use in latency-critical and high-motion video systems while taking targeted next steps. Engineering teams should fine-tune the architecture using diverse datasets (such as Virtual KITTI, AutoFlow, or TartanAir) to address cross-domain gaps and improve occlusion handling before integrating it into safety-critical domains like autonomous vehicle navigation.

arXiv: 2111.13680
Cover for GMFlow: Learning Optical Flow via Global Matching

Abstract

Learning-based optical flow estimation has been dominated with the pipeline of cost volume with convolutions for flow regression, which is inherently limited to local correlations and thus is hard to address the long-standing challenge of large displacements. To alleviate this, the state-of-the-art framework RAFT gradually improves its prediction quality by using a large number of iterative refinements, achieving remarkable performance but introducing linearly increasing inference time. To enable both high accuracy and efficiency, we completely revamp the dominant flow regression pipeline by reformulating optical flow as a global matching problem, which identifies the correspondences by directly comparing feature similarities. Specifically, we propose a GMFlow framework, which consists of three main components: a customized Transformer for feature enhancement, a correlation and softmax layer for global feature matching, and a self-attention layer for flow propagation. We further introduce a refinement step that reuses GMFlow at higher feature resolution for residual flow prediction. Our new framework outperforms 31-refinements RAFT on the challenging Sintel benchmark, while using only one refinement and running faster, suggesting a new paradigm for accurate and efficient optical flow estimation. Code is available at https://github.com/haofeixu/gmflow.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Methodology
  • 3.1. Formulation
  • 3.2. Feature Enhancement
  • 3.3. Flow Propagation
  • 3.4. Refinement
  • 3.5. Training Loss
  • 4. Experiments
  • 4.1. Methodology Comparison
  • 4.2. Ablations
  • 4.3. Comparison with RAFT
  • 4.4. Comparison on Benchmarks
  • 4.5. Limitation and Discussion
  • 5. Conclusion
  • References

Knowls

  1. Knowl 1 — Global matching formulation for dense optical flow

    equation

    GMFlow reformulates optical flow between consecutive frames I1I_1 and I2I_2 as dense global feature matching. A weight-sharing convolutional backbone produces feature maps that are flattened into mathbf F_1,mathbf F_2 ewcommand{f}{mathbf}mathbf F_2\in\mathbb{R}^{HW\times D}, where H×WH\times W is the feature-map resolution and DD is the feature dimension. Let mathbf G\in\mathbb{R}^{HW\times 2} contain the 2D feature-grid coordinates, measured in feature-pixel units. Every location in the first frame is compared with every location in the second frame:

    C=F1F2⊤D∈RHW×HW,M=softmax⁡q(C),G^=MG,V=G^−G.\mathbf C=\frac{\mathbf F_1\mathbf F_2^\top}{\sqrt{D}}\in\mathbb{R}^{HW\times HW},\qquad \mathbf M=\operatorname{softmax}_{q}(\mathbf C),\qquad \widehat{\mathbf G}=\mathbf M\mathbf G,\qquad \mathbf V=\widehat{\mathbf G}-\mathbf G.

    Here Cpq\mathbf C_{pq} is the similarity between source location pp in I1I_1 and candidate location qq in I2I_2, and the softmax is applied over all candidate locations qq for each source location pp. The resulting M\mathbf M is a differentiable matching distribution, mathbf V\in\mathbb{R}^{HW\times 2} is the predicted flow, and the weighted coordinate average gives sub-pixel correspondence estimates. Because the same global correlation contains similarities in both directions, backward flow can be obtained by transposing the correlation matrix at inference without a second network forward pass; the two directions can then be used for forward-backward consistency occlusion detection.

  2. Knowl 2 — Transformer feature enhancement with shifted local attention

    model/method

    The independently extracted feature maps are enhanced by a symmetric Transformer so that their descriptors encode both within-image context and cross-image dependencies. Fixed 2D sine-cosine positional encodings mathbf P are added because the feature sets otherwise lack explicit spatial coordinates. With the first argument interpreted as queries and the second as keys and values, the enhanced features are

    F^1=T(F1+P,F2+P),F^2=T(F2+P,F1+P).\widehat{\mathbf F}_1=\mathcal T(\mathbf F_1+\mathbf P,\mathbf F_2+\mathbf P),\qquad \widehat{\mathbf F}_2=\mathcal T(\mathbf F_2+\mathbf P,\mathbf F_1+\mathbf P).

    The Transformer uses six stacked blocks, each containing self-attention, cross-attention, and a feed-forward network, applied symmetrically to both frames. To avoid full-image quadratic attention, the H×WH\times W feature map is divided into a fixed number K×KK\times K of adaptive windows rather than fixed-size windows. The implementation uses K=2K=2, so each window has size (H/2)×(W/2)(H/2)\times(W/2); the partition is shifted by half a window between consecutive partitions to connect neighboring windows. On the Things validation set, the complete six-block model obtains EPE 6.676.67, whereas removing cross-attention, positional encoding, the feed-forward network, or self-attention gives EPE 10.8410.84, 8.388.38, 8.718.71, and 7.047.04, respectively. Cross-attention therefore provides the largest measured contribution, while positional encoding and self-attention also improve matching.

  3. Knowl 3 — Self-attention flow propagation for unmatched pixels

    equation

    Softmax matching assumes that a corresponding visible location exists in the second frame, which fails for occluded or out-of-boundary pixels. GMFlow first computes an initial flow from the enhanced features and then propagates flow through feature self-similarity:

    V^=softmax⁡q(F^1F^2⊤D)G−G,\widehat{\mathbf V}=\operatorname{softmax}_{q}\left(\frac{\widehat{\mathbf F}_1\widehat{\mathbf F}_2^{\top}}{\sqrt D}\right)\mathbf G-\mathbf G, V~=softmax⁡q(F^1F^1⊤D)V^.\widetilde{\mathbf V}=\operatorname{softmax}_{q}\left(\frac{\widehat{\mathbf F}_1\widehat{\mathbf F}_1^{\top}}{\sqrt D}\right)\widehat{\mathbf V}.

    Here mathbf G\in\mathbb{R}^{HW\times2} is the feature-coordinate grid, mathbf F_1,mathbf F_2\in\mathbb{R}^{HW\times D} are the enhanced feature matrices, mathbf V is the initial flow at every source location, and mathbf V is the propagated flow. The second softmax averages initial flow vectors from feature-similar locations within the first frame, allowing reliable matched pixels to provide estimates for difficult unmatched pixels. On Sintel, propagation reduces unmatched-pixel EPE from 15.5415.54 to 10.3910.39 on the clean pass and from 19.5019.50 to 15.5215.52 on the final pass; overall EPE changes from 2.282.28 to 1.891.89 and from 3.443.44 to 3.133.13, respectively.

  4. Knowl 4 — Higher-resolution residual refinement

    model/method

    GMFlow first predicts flow using features at 1/81/8 of the input image resolution and optionally performs one refinement at 1/41/4 resolution. The coarse flow is upsampled to 1/41/4 resolution, the second-frame feature is warped using this estimate, and the refinement network predicts only the residual flow. The same GMFlow operations are reused locally: the Transformer partitions the 1/41/4-resolution feature map into 8×88\times8 windows, each corresponding to 1/321/32 of the original image resolution; matching searches a 9×99\times9 local window around the warped correspondence; and flow propagation uses a 3×33\times3 local self-attention window. The Transformer and flow-propagation weights are shared between the global 1/81/8 stage and the local 1/41/4 refinement stage. Backbone weights are also shared when producing the two resolutions, reducing parameters and improving generalization.

  5. Knowl 5 — Exponentially weighted supervision of flow predictions

    equation

    All intermediate and final flow predictions are supervised with an ell_1 loss. If mathbf V_i is the iith predicted flow, mathbf V_{\mathrm{gt}} is the ground-truth flow, and NN is the total number of predictions, the training objective is

    L=∑i=1NγN−i∥Vi−Vgt∥1,γ=0.9.L=\sum_{i=1}^{N}\gamma^{N-i}\left\|\mathbf V_i-\mathbf V_{\mathrm{gt}}\right\|_1, \qquad \gamma=0.9.

    The norm sums absolute flow-component errors over the spatial domain. Because gamma^{N-i} gives larger weights to later predictions, the final refinements receive stronger supervision while earlier predictions remain directly trained.

  6. Knowl 6 — Training and evaluation protocol

    experimental setup

    GMFlow uses a convolutional backbone matching RAFT's backbone design but produces 128-dimensional features instead of RAFT's 256-dimensional features. It uses six Transformer blocks and RAFT's convex upsampling to recover full-resolution flow. Training uses AdamW: 100,000 iterations on FlyingChairs with batch size 16 and learning rate 4×10−44\times10^{-4}, followed by FlyingThings3D fine-tuning with batch size 8 and learning rate 2×10−42\times10^{-4}. Ablations use 200,000 Things iterations, while the final Things model uses 800,000 iterations. The principal metrics are end-point error, the mean Euclidean flow error; motion-specific EPE is reported for ground-truth magnitudes of 00–1010, 1010–4040, and more than 4040 pixels, denoted s0−10s_{0-10}, s10−40s_{10-40}, and s40+s_{40+}. KITTI additionally uses F1-all, the percentage of outlier pixels.

  7. Knowl 7 — Global matching improves large-displacement estimation

    empirical result

    On models trained on FlyingChairs and FlyingThings3D, the proposed Transformer-plus-global-softmax formulation is substantially stronger than a cost-volume-plus-convolution baseline with a similar comparison protocol. With six Transformer blocks, GMFlow obtains Things validation EPE 6.676.67 and s40+=17.37s_{40+}=17.37, Sintel training-clean EPE 2.282.28 and s40+=13.89s_{40+}=13.89, and Sintel training-final EPE 3.443.44 and s40+=21.02s_{40+}=21.02, using 4.2 million parameters. An 18-block cost-volume model obtains EPE 8.678.67, 2.612.61, and 3.943.94 on those three evaluations, with corresponding s40+s_{40+} values 23.4323.43, 15.9115.91, and 24.5824.58, while using 15.7 million parameters. Replacing global matching with a 9×99\times9 local search increases Things validation EPE from 6.676.67 to 19.8819.88 and s40+s_{40+} from 17.3717.37 to 61.0661.06, despite nearly identical inference time: 52.652.6 ms for global matching versus 52.952.9 ms for local matching. The comparison indicates that the large-motion improvement comes from the global search rather than merely from using a stronger feature processor.

  8. Knowl 8 — Accuracy and efficiency relative to iterative RAFT

    empirical result

    When trained on FlyingChairs and FlyingThings3D and evaluated at Sintel resolution 436×1024436\times1024, GMFlow reaches stronger accuracy with fewer sequential refinements than RAFT. With no refinement, GMFlow obtains Things validation EPE 3.483.48, Sintel-clean EPE 1.501.50, Sintel-final EPE 2.962.96, and corresponding s40+s_{40+} values 8.978.97, 8.268.26, and 17.7017.70; its inference time is 5757 ms on a V100 and 2626 ms on an A100. With one refinement, these values become EPE 2.802.80, 1.081.08, and 2.482.48, s40+s_{40+} values 7.317.31, 6.266.26, and 15.6715.67, and inference time 151151 ms on a V100 and 6666 ms on an A100. RAFT with 31 refinements obtains EPE 4.254.25, 1.411.41, and 2.692.69, s40+s_{40+} values 11.6311.63, 8.838.83, and 17.4517.45, and inference time 170170 ms on a V100 and 9191 ms on an A100. Thus one GMFlow refinement outperforms the 31-refinement RAFT model on these evaluations while using 4.7 million rather than 5.3 million parameters and requiring less sequential computation.

  9. Knowl 9 — Performance on the Sintel test benchmark

    empirical result

    After fine-tuning on mixed training data, GMFlow obtains Sintel test EPE 1.741.74 on the clean pass and 2.902.90 on the final pass. For clean Sintel, its EPE on matched pixels is 0.650.65 and on unmatched pixels is 10.5610.56; for final Sintel, the corresponding values are 1.321.32 and 15.8015.80. Among methods restricted to the two input frames at inference, GMFlow is better than FlowNet2, PWC-Net+, HD3, VCN, DICL, and RAFT, whose all-pixel clean/final EPE values are respectively 4.16/5.744.16/5.74, 3.45/4.603.45/4.60, 4.79/4.674.79/4.67, 2.81/4.402.81/4.40, 2.63/3.602.63/3.60, and 1.94/3.181.94/3.18. RAFT with an additional previous-frame initialization obtains 1.61/2.861.61/2.86, and GMA with that initialization obtains 1.39/2.471.39/2.47; these methods use multi-frame information unavailable to the two-frame GMFlow evaluation.

  10. Knowl 10 — Occlusion and domain-gap limitations on KITTI

    limitation

    GMFlow remains weaker than RAFT on real-world KITTI data, particularly in occluded regions. After fine-tuning on KITTI 2015, F1-all/F1 on non-occluded pixels is 9.32/3.809.32/3.80 for GMFlow versus 5.10/3.075.10/3.07 for RAFT; the larger gap on all pixels indicates that occlusion handling is a principal weakness. The same limitation appears under synthetic-to-real generalization: after training only on FlyingChairs and FlyingThings3D, GMFlow obtains KITTI EPE 7.777.77 and F1-all 23.4023.40, compared with RAFT's 5.325.32 and 17.4617.46. Adding Virtual KITTI 2 reduces the gap, giving GMFlow 2.852.85 EPE and 10.7710.77 F1-all versus RAFT's 2.452.45 and 7.907.90. The authors therefore identify unreliable occluded-region predictions and poor generalization across large training-test domain gaps as remaining limitations.

Coverage note — Secondary numerical ablations on training length, window-count speed trade-offs, multi-scale weight sharing, and qualitative visualizations were omitted because they validate the included design choices without adding independent load-bearing contributions.

References

  1. 1.Christian Bailer, Bertram Taetz, and Didier Stricker. Flow fields: Dense correspondence fields for highly accurate large displacement optical flow estimation. In ICCV, pages 4015–4023, 2015. 2
  2. 2.Thomas Brox, Christoph Bregler, and Jitendra Malik. Large displacement optical flow. In ICCV, pages 41–48. IEEE, 2009. 2
  3. 3.Thomas Brox, Andrés Bruhn, Nils Papenberg, and Joachim Weickert. High accuracy optical flow estimation based on a theory for warping. In ECCV, pages 25–36. Springer, 2004. 2
  4. 4.Daniel J Butler, Jonas Wulff, Garrett B Stanley, and Michael J Black. A naturalistic open source movie for optical flow evaluation. In ECCV, pages 611–625. Springer, 2012. 2, 5, 8
  5. 5.Yohann Cabon, Naila Murray, and Martin Humenberger. Virtual kitti 2. arXiv preprint arXiv:2001.10773, 2020. 8
  6. 6.Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. End-to-end object detection with transformers. In ECCV, pages 213–229. Springer, 2020. 4
  7. 7.Zhuoyuan Chen, Hailin Jin, Zhe Lin, Scott Cohen, and Ying Wu. Large displacement optical flow from nearest neighbor fields. In CVPR, pages 2443–2450, 2013. 2
  8. 8.Stephane d’Ascoli, Hugo Touvron, Matthew Leavitt, Ari Morcos, Giulio Biroli, and Levent Sagun. Convit: Improving vision transformers with soft convolutional inductive biases. ICML, 2021. 8
  9. 9.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition at scale. ICLR, 2021. 7, 8
  10. 10.Alexey Dosovitskiy, Philipp Fischer, Eddy Ilg, Philip Hausser, Caner Hazirbas, Vladimir Golkov, Patrick Van Der Smagt, Daniel Cremers, and Thomas Brox. Flownet: Learning optical flow with convolutional networks. In ICCV, pages 2758–2766, 2015. 1, 5
  11. 11.Adrien Gaidon, Qiao Wang, Yohann Cabon, and Eleonora Vig. Virtual worlds as proxy for multi-object tracking analysis. In CVPR, pages 4340–4349, 2016. 8
  12. 12.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In CVPR, pages 770–778, 2016. 5
  13. 13.Asmaa Hosni, Christoph Rhemann, Michael Bleyer, Carsten Rother, and Margrit Gelautz. Fast cost-volume filtering for visual correspondence and beyond. TPAMI, 35(2):504–511, 2012. 1
  14. 14.Yinlin Hu, Rui Song, and Yunsong Li. Efficient coarse-to-fine patchmatch for large displacement optical flow. In CVPR, pages 5704–5712, 2016. 2
  15. 15.Junhwa Hur and Stefan Roth. Iterative residual refinement for joint optical flow and occlusion estimation. In CVPR, pages 5754–5763, 2019. 1, 2, 6
  16. 16.Eddy Ilg, Nikolaus Mayer, Tonmoy Saikia, Margret Keuper, Alexey Dosovitskiy, and Thomas Brox. Flownet 2.0: Evolution of optical flow estimation with deep networks. In CVPR, pages 2462–2470, 2017. 1, 2, 5, 8
  17. 17.Shihao Jiang, Dylan Campbell, Yao Lu, Hongdong Li, and Richard Hartley. Learning to estimate hidden motions with global motion aggregation. In ICCV, pages 9772–9781, October 2021. 1, 2, 4, 8
  18. 18.Wei Jiang, Eduard Trulls, Jan Hosang, Andrea Tagliasacchi, and Kwang Moo Yi. Cotr: Correspondence transformer for matching across images. ICCV, 2021. 2
  19. 19.Alex Kendall, Hayk Martirosyan, Saumitro Dasgupta, Peter Henry, Ryan Kennedy, Abraham Bachrach, and Adam Bry. End-to-end learning of geometry and context for deep stereo regression. In ICCV, pages 66–75, 2017. 3
  20. 20.Daniel Kondermann, Rahul Nair, Katrin Honauer, Karsten Krispin, Jonas Andrulis, Alexander Brock, Burkhard Gussefeld, Mohsen Rahimimoghaddam, Sabine Hofmann, Claus Brenner, et al. The hci benchmark suite: Stereo and flow ground truth with uncertainties for urban autonomous driving. In CVPR Workshops, pages 19–28, 2016. 8
  21. 21.Yanghao Li, Yuntao Chen, Naiyan Wang, and Zhaoxiang Zhang. Scale-aware trident networks for object detection. In ICCV, pages 6054–6063, 2019. 4
  22. 22.Zhaoshuo Li, Xingtong Liu, Nathan Drenkow, Andy Ding, Francis X Creighton, Russell H Taylor, and Mathias Unberath. Revisiting stereo depth estimation from a sequence-to-sequence perspective with transformers. In ICCV, pages 6197–6206, 2021. 3
  23. 23.Tsung-Yi Lin, Piotr Dollar, Ross Girshick, Kaiming He, Bharath Hariharan, and Serge Belongie. Feature pyramid networks for object detection. In CVPR, pages 2117–2125, 2017. 4
  24. 24.Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. ICCV, 2021. 4, 7
  25. 25.Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101, 2017. 5
  26. 26.Zhaoyang Lv, Kihwan Kim, Alejandro Troccoli, Deqing Sun, James M Rehg, and Jan Kautz. Learning rigidity in dynamic scenes with a moving camera for 3d motion field estimation. In ECCV, pages 468–484, 2018. 8
  27. 27.Nikolaus Mayer, Eddy Ilg, Philip Hausser, Philipp Fischer, Daniel Cremers, Alexey Dosovitskiy, and Thomas Brox. A large dataset to train convolutional networks for disparity, optical flow, and scene flow estimation. In CVPR, pages 4040–4048, 2016. 5, 8
  28. 28.Simon Meister, Junhwa Hur, and Stefan Roth. Unflow: Unsupervised learning of optical flow with a bidirectional census loss. In AAAI, 2018. 6
  29. 29.Moritz Menze and Andreas Geiger. Object scene flow for autonomous vehicles. In CVPR, pages 3061–3070, 2015. 5, 8
  30. 30.Federico Perazzi, Jordi Pont-Tuset, Brian McWilliams, Luc Van Gool, Markus Gross, and Alexander Sorkine-Hornung. A benchmark dataset and evaluation methodology for video object segmentation. In CVPR, pages 724–732, 2016. 8
  31. 31.Anurag Ranjan and Michael J Black. Optical flow estimation using a spatial pyramid network. In CVPR, pages 4161–4170, 2017. 1
  32. 32.Jerome Revaud, Philippe Weinzaepfel, Zaid Harchaoui, and Cordelia Schmid. Epicflow: Edge-preserving interpolation of correspondences for optical flow. In CVPR, pages 1164–1172, 2015. 2
  33. 33.Stephan R Richter, Zeeshan Hayder, and Vladlen Koltun. Playing for benchmarks. In ICCV, pages 2213–2222, 2017. 8
  34. 34.Paul-Edouard Sarlin, Daniel DeTone, Tomasz Malisiewicz, and Andrew Rabinovich. Superglue: Learning feature matching with graph neural networks. In CVPR, pages 4938–4947, 2020. 1, 2, 3, 4
  35. 35.Deqing Sun, Daniel Vlasic, Charles Herrmann, Varun Jampani, Michael Krainin, Huiwen Chang, Ramin Zabih, William T Freeman, and Ce Liu. Autoflow: Learning a better training set for optical flow. In CVPR, pages 10093–10102, 2021. 8
  36. 36.Deqing Sun, Xiaodong Yang, Ming-Yu Liu, and Jan Kautz. Pwc-net: Cnns for optical flow using pyramid, warping, and cost volume. In CVPR, pages 8934–8943, 2018. 1, 2, 5
  37. 37.Deqing Sun, Xiaodong Yang, Ming-Yu Liu, and Jan Kautz. Models matter, so does training: An empirical study of cnns for optical flow estimation. TPAMI, 42(6):1408–1423, 2019. 8
  38. 38.Jiaming Sun, Zehong Shen, Yuang Wang, Hujun Bao, and Xiaowei Zhou. Loftr: Detector-free local feature matching with transformers. In CVPR, pages 8922–8931, 2021. 1, 2, 3, 4
  39. 39.Zachary Teed and Jia Deng. Raft: Recurrent all-pairs field transforms for optical flow. In ECCV, pages 402–419. Springer, 2020. 1, 2, 5, 7, 8
  40. 40.Prune Truong, Martin Danelljan, and Radu Timofte. Glu-net: Global-local universal network for dense flow and correspondences. In CVPR, pages 6258–6268, 2020. 2
  41. 41.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In NeurIPS, pages 5998–6008, 2017. 1, 2, 3, 4
  42. 42.Jianyuan Wang, Yiran Zhong, Yuchao Dai, Kaihao Zhang, Pan Ji, and Hongdong Li. Displacement-invariant matching cost learning for accurate optical flow estimation. NeurIPS, 33, 2020. 2, 8
  43. 43.Qianqian Wang, Xiaowei Zhou, Bharath Hariharan, and Noah Snavely. Learning feature descriptors using camera pose supervision. In ECCV, pages 757–774. Springer, 2020. 1, 3
  44. 44.Wenshan Wang, Delong Zhu, Xiangwei Wang, Yaoyu Hu, Yuheng Qiu, Chen Wang, Yafei Hu, Ashish Kapoor, and Sebastian Scherer. Tartanair: A dataset to push the limits of visual slam. 2020. 8
  45. 45.Philippe Weinzaepfel, Jerome Revaud, Zaid Harchaoui, and Cordelia Schmid. Deepflow: Large displacement optical flow with deep matching. In ICCV, pages 1385–1392, 2013. 2
  46. 46.Haofei Xu, Jiaolong Yang, Jianfei Cai, Juyong Zhang, and Xin Tong. High-resolution optical flow from 1d attention and correlation. In ICCV, pages 10498–10507, 2021. 1, 2, 8
  47. 47.Haofei Xu and Juyong Zhang. Aanet: Adaptive aggregation network for efficient stereo matching. In CVPR, pages 1959–1968, 2020. 3
  48. 48.Yufei Xu, Qiming Zhang, Jing Zhang, and Dacheng Tao. Vitae: Vision transformer advanced by exploring intrinsic inductive bias. NeurIPS, 2021. 8
  49. 49.Gengshan Yang and Deva Ramanan. Volumetric correspondence networks for optical flow. NeurIPS, 32:794–805, 2019. 8
  50. 50.Zhichao Yin, Trevor Darrell, and Fisher Yu. Hierarchical discrete distribution decomposition for match density estimation. In CVPR, pages 6044–6053, 2019. 8
  51. 51.Feihu Zhang, Oliver J. Woodford, Victor Adrian Prisacariu, and Philip H.S. Torr. Separable flow: Learning motion cost volumes for optical flow estimation. In ICCV, pages 10807–10817, October 2021. 1, 2
  52. 52.Qiming Zhang, Yufei Xu, Jing Zhang, and Dacheng Tao. Vitaev2: Vision transformer advanced by exploring inductive bias for image recognition and beyond. arXiv preprint arXiv:2202.10108, 2022. 8

Citation

MLA
Xu, H., et al. “GMFlow: Learning Optical Flow via Global Matching”. arXiv, 2021, http://arxiv.org/abs/2111.13680v4.
APA
Xu, H., Zhang, J., Cai, J., Rezatofighi, H., & Tao, D. (2021). GMFlow: Learning Optical Flow via Global Matching. arXiv. http://arxiv.org/abs/2111.13680v4
Chicago
Xu, H., J. Zhang, J. Cai, H. Rezatofighi, and D. Tao. 2021. “GMFlow: Learning Optical Flow via Global Matching”. arXiv. http://arxiv.org/abs/2111.13680v4.
Harvard
Xu, H. et al. (2021) “GMFlow: Learning Optical Flow via Global Matching”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2111.13680v4.
Vancouver
1. Xu H, Zhang J, Cai J, Rezatofighi H, Tao D (2021) GMFlow: Learning Optical Flow via Global Matching. arXiv

BibTeX

@article{xu2021gmflow,
  title = {GMFlow: Learning Optical Flow via Global Matching},
  author = {Xu, Haofei and Zhang, Jing and Cai, Jianfei and Rezatofighi, Hamid and Tao, Dacheng},
  year = {2021},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2111.13680v4},
  eprint = {2111.13680}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE