Rethinking Optical Flow from Geometric Matching Consistent Perspective

Qiaole DongChenjie CaoYanwei Fu

article2023CVPR65 citations

Proposes MatchFlow, an optical flow architecture that leverages geometric image matching on real-world static scenes as a pre-training task to improve feature correspondences and boost cross-dataset generalization on benchmarks like Sintel and KITTI.

Listen

Optical flow estimation—the tracking of pixel movement between video frames—is vital for applications such as autonomous driving, video enhancement, and action recognition. However, standard deep learning models frequently struggle with small, fast-moving objects, occlusions, and textureless regions. This deficiency occurs largely because models are trained from scratch on synthetic datasets that lack realistic geometric consistency, limiting their ability to learn robust visual correspondence across scenes.

The article evaluates whether pre-training feature extractors on Geometric Image Matching (GIM) using large-scale real-world static scene data significantly improves optical flow accuracy and cross-dataset generalization. The authors propose MatchFlow, a deep learning architecture that integrates GIM pre-training with attention-driven feature matching to establish a stronger foundation for motion estimation.

The approach introduces a two-stage training strategy. First, a Feature Matching Extractor utilizing ResNet-16 and stacked QuadTree attention blocks is pre-trained on real-world multi-view images from MegaDepth to learn fundamental geometric correlations. Second, the model is refined on standard optical flow benchmark datasets (FlyingChairs, FlyingThings3D, Sintel, and KITTI) using iterative recurrent decoders. The evaluation assesses generalization performance and endpoint error across established academic benchmarks.

The findings confirm that geometric pre-training substantially enhances flow estimation. The full MatchFlow model achieved state-of-the-art performance on the Sintel benchmark, delivering an 11.5% error reduction on the Sintel clean pass test set and a 10.1% error reduction on the KITTI test set compared to the GMA baseline. Ablations demonstrate that the greatest accuracy gains occur in non-occluded regions, confirming that foundational matching is heavily improved. Furthermore, the 15.4-million-parameter model maintains modest computational demands, processing frames in 126 milliseconds.

These results demonstrate that incorporating real-world static scene matching into the training pipeline resolves key representation bottlenecks without requiring prohibitive computational overhead. For technical leaders and engineering teams, adopting pre-trained geometric correspondence models can improve tracking precision and edge-case reliability in computer vision deployments.

Organizations developing motion estimation systems should adopt geometric image matching pre-training as the initial phase of their model training pipelines. Future development should focus on expanding pre-training datasets to include more complex motion blur and dynamic lighting variations to eliminate edge-case errors, as well as optimizing attention mechanisms to further lower inference latency for real-time edge environments.

arXiv: 2303.08384
Cover for Rethinking Optical Flow from Geometric Matching Consistent Perspective

Abstract

Optical flow estimation is a challenging problem remaining unsolved. Recent deep learning based optical flow models have achieved considerable success. However, these models often train networks from the scratch on standard optical flow data, which restricts their ability to robustly and geometrically match image features. In this paper, we propose a rethinking to previous optical flow estimation. We particularly leverage Geometric Image Matching (GIM) as a pre-training task for the optical flow estimation (MatchFlow) with better feature representations, as GIM shares some common challenges as optical flow estimation, and with massive labeled real-world data. Thus, matching static scenes helps to learn more fundamental feature correlations of objects and scenes with consistent displacements. Specifically, the proposed MatchFlow model employs a QuadTree attention-based network pre-trained on MegaDepth to extract coarse features for further flow regression. Extensive experiments show that our model has great cross-dataset generalization. Our method achieves 11.5% and 10.1% error reduction from GMA on Sintel clean pass and KITTI test set. At the time of anonymous submission, our MatchFlow(G) enjoys state-of-the-art performance on Sintel clean and final pass compared to published approaches with comparable computation and memory footprint. Codes and models will be released in https://github.com/DQiaole/MatchFlow.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Methodology
  • 3.1. Geometric Image Matching Pre-training
  • 3.2. Optical Flow Refinement
  • 4. Experiments
  • 4.1. Sintel
  • 4.2. KITTI
  • 4.3. Ablation Studies
  • 4.4. Where are Gains Coming from?
  • 4.5. Parameters, Timing, and GPU Memory
  • 4.6. Failure Cases and Limitations
  • 5. Conclusion
  • References

Knowls

  1. Knowl 1 — Geometric matching as optical-flow pretraining

    model/method

    MatchFlow reformulates optical-flow learning as a two-stage curriculum. The first stage trains a Feature Matching Extractor (FME) on geometric image matching (GIM): pairs of static real-world scenes whose displacement is caused by viewpoint changes and whose correspondence labels are obtained from depth and camera poses. The second stage fine-tunes the FME jointly with an Optical Flow Estimator (OFE) on increasingly difficult motion data: FlyingChairs, FlyingThings3D, Sintel, and KITTI.

    GIM is used as a simpler precursor to optical flow because it provides large-scale real image pairs with consistent scene displacement, appearance changes, and large geometric displacements. The resulting representation is intended to learn low-level feature correspondences before the network must model independent 3D motion of multiple objects.

    For two consecutive images I1I_1 and I2I_2, the optical-flow field is a pair of per-pixel displacement functions (f1,f2)(f^1,f^2) mapping pixel (u,v)(u,v) in I1I_1 to

    (u,v)↦(u+f1(u,v),  v+f2(u,v)).(u,v)\mapsto\bigl(u+f^1(u,v),\;v+f^2(u,v)\bigr).

    The full model using global motion aggregation is called MatchFlow(G); the otherwise corresponding model without global motion aggregation is called MatchFlow(R).

  2. Knowl 2 — QuadTree Feature Matching Extractor

    model/method

    The FME extracts matching features at one-eighth of the input resolution using ResNet-16 feature stages followed by eight stacked, interleaved QuadTree attention blocks. The attention blocks alternate self-attention within each image and cross-attention between the two images, allowing the extractor to represent both local detail and long-range correspondence. Sparse GIM labels are available mainly in non-occluded regions, so the FME is pretrained only on the matching task rather than on dense optical-flow supervision.

    For feature maps F1F_1 and F2F_2 from the reference and matching images, respectively, the FME forms an all-pairs 4D correlation volume by inner products:

    C(i,j)=⟨F1(i),F2(j)⟩,C∈RH×W×H×W.C(i,j)=\langle F_1(i),F_2(j)\rangle, \qquad C\in\mathbb{R}^{H\times W\times H\times W}.

    Here ii and jj index spatial locations in the two one-eighth-resolution feature maps, HH and WW are their height and width, and F1(i),F2(j)F_1(i),F_2(j) are feature vectors. During optical-flow refinement, multi-scale versions of this correlation volume are queried by the recurrent estimator.

  3. Knowl 3 — Dual-softmax objective for GIM pretraining

    equation

    The GIM pretraining objective converts the correlation volume into a correspondence probability with dual softmax. For a correlation score C(i,j)C(i,j) between location ii in the first image and location jj in the second image, the probability is

    Pc(i,j)=softmax⁡ ⁣(C(i,⋅)τ)j softmax⁡ ⁣(C(⋅,j)τ)i,\mathcal{P}_c(i,j)=\operatorname{softmax}\!\left(\frac{C(i,\cdot)}{\tau}\right)_j\,\operatorname{softmax}\!\left(\frac{C(\cdot,j)}{\tau}\right)_i,

    where τ=0.1\tau=0.1 is the temperature, the first softmax is over candidate locations in the second image, and the second softmax is over candidate locations in the first image. Let Mcgt\mathcal{M}_c^{gt} be the set of ground-truth matching location pairs obtained from GIM depth and camera-pose annotations at one-eighth resolution. The main pretraining loss is

    Lc=−1∣Mcgt∣∑(i,j)∈Mcgtlog⁡Pc(i,j).\mathcal{L}_c=-\frac{1}{|\mathcal{M}_c^{gt}|}\sum_{(i,j)\in\mathcal{M}_c^{gt}}\log \mathcal{P}_c(i,j).

    An additional ℓ2\ell_2 loss at one-half input resolution supplies finer-grained supervision. The objective therefore trains the FME to concentrate correlation around geometrically correct, non-occluded matches before optical-flow fine-tuning.

  4. Knowl 4 — Recurrent optical-flow refinement with optional global aggregation

    model/method

    After FME pretraining, MatchFlow jointly fine-tunes the FME and an OFE based on the RAFT-style recurrent estimator. Given two images, the OFE repeatedly looks up features from the multi-scale correlation volume and uses a convolutional GRU to decode motion and context features into residual flow updates. The final flow is the sum of all residual updates.

    MatchFlow(G) additionally uses a global motion aggregation module. It extracts 2D context features and propagates reliable motion learned in non-occluded regions to ambiguous or occluded regions using context self-similarity. MatchFlow(R) omits this module and corresponds to the RAFT-style variant. Thus, GIM pretraining primarily improves correspondence in visible regions, while global aggregation is used to extend reliable motion into occlusions.

  5. Knowl 5 — Iterative optical-flow supervision

    equation

    Let fgtf_{gt} be the dense ground-truth optical-flow field, let fif_i be the partially summed flow after recurrent update ii, and let NN be the total number of updates. MatchFlow trains the recurrent estimator with a weighted sum of pixelwise ℓ1\ell_1 errors:

    L=∑i=1NγN−i∥fgt−fi∥1.\mathcal{L}=\sum_{i=1}^{N}\gamma^{N-i}\left\|f_{gt}-f_i\right\|_1.

    The discount factor is γ=0.8\gamma=0.8 for FlyingChairs and FlyingThings3D, and γ=0.85\gamma=0.85 for Sintel and KITTI. Later predictions receive larger weight because the exponent N−iN-i is smaller for later iterations.

  6. Knowl 6 — Training schedule and computational configuration

    experimental setup

    The FME is pretrained on MegaDepth by randomly sampling 36,800 image pairs per epoch for 30 epochs. The resulting model is then trained for 120,000 iterations on FlyingChairs with batch size 8 and another 120,000 iterations on FlyingThings3D with batch size 6. It is subsequently fine-tuned using Sintel, KITTI, and HD1K data; the notation C+T+S+K+HC+T+S+K+H denotes this combined fine-tuning regime. Training uses PyTorch on two RTX 3090 GPUs with a one-cycle learning-rate schedule: the maximum learning rate is 2.5×10−42.5\times10^{-4} on FlyingChairs and 1.25×10−41.25\times10^{-4} on the remaining datasets.

    For KITTI, whose aspect ratio is difficult for QuadTree attention, the final evaluation uses resolution adjustment plus tiling: each image is split into two smaller overlapping sub-images, predictions in the overlap are weighted and averaged, and bilinear interpolation is used instead of zero-padding. The models use 12 recurrent GRU updates at inference. MatchFlow(R) has 14.8M parameters, 110 ms average inference time, 53 h training time, and 22.1 GB test memory; MatchFlow(G) has 15.4M parameters, 126 ms inference time, 58 h training time, and 23.6 GB test memory. These measurements were obtained on the reported Sintel/FlyingThings3D configurations and are substantially lighter than the reported FlowFormer configuration.

  7. Knowl 7 — Cross-dataset optical-flow performance

    empirical result

    GIM pretraining substantially improves generalization from synthetic training data to real benchmark domains. After training on FlyingChairs plus FlyingThings3D, MatchFlow(G) obtains Sintel training AEPE of 1.03 on the clean pass and 2.45 on the final pass, compared with GMA's 1.30 and 2.74. These are reductions of 20.8% and 10.6%. MatchFlow(R) obtains 1.14 and 2.61, improving over RAFT by 20.3% and 3.7%.

    On KITTI-15 training data under the same synthetic-only training regime, MatchFlow(G) obtains EPE 4.08 and Fl-all 15.6%, while MatchFlow(R) obtains EPE 4.19 and Fl-all 13.6%. Fl-all is the percentage of flow vectors whose endpoint error exceeds either 3 pixels or 5% of the ground-truth flow magnitude.

    After fine-tuning with the additional Sintel, KITTI, and HD1K data, MatchFlow(G) obtains Sintel test AEPE of 1.16 on clean and 2.37 on final, and KITTI-15 test Fl-all of 4.63%. The corresponding MatchFlow(R) values are 1.33, 2.64, and 4.72%. Relative to GMA's 1.39 clean AEPE, 2.47 final AEPE, and 5.15% KITTI Fl-all, MatchFlow(G) reduces error by 11.5% on Sintel clean and 10.1% on KITTI, while achieving 2.37 AEPE on Sintel final.

  8. Knowl 8 — Ablation evidence for GIM pretraining and QuadTree attention

    empirical result

    Ablations trained on FlyingChairs plus FlyingThings3D show that the improvement is not explained by merely adding parameters. In the 15.4M-parameter architecture, removing GIM pretraining gives FlyingThings3D test AEPE of 2.82 on clean and 2.56 on final, and Sintel training AEPE of 1.27 and 2.84. GIM pretraining reduces these values to 2.12, 2.07, 1.03, and 2.45, respectively. It also reduces KITTI training EPE from 4.12 to 4.08, although KITTI Fl-all changes from 14.4% to 15.6% in this ablation.

    The number of interleaved QuadTree blocks is also critical. With zero, four, and eight blocks, FlyingThings3D clean/final AEPE is respectively 2.97/2.81, 2.65/2.40, and 2.12/2.07; Sintel clean/final AEPE is 1.36/2.98, 1.18/2.79, and 1.03/2.45. Replacing QuadTree attention with comparable-parameter Linear attention or Global-Local Attention gives FlyingThings3D clean/final AEPE of 2.82/2.56 and 2.66/2.55, respectively, versus 2.12/2.07 with QuadTree attention. Adding match initialization or a matching loss during optical-flow training after GIM pretraining does not provide a consistent improvement.

  9. Knowl 9 — Where the optical-flow gains occur

    data/table

    Using models trained on FlyingThings3D, the paper separates Sintel pixels into non-occluded regions (Noc), occluded regions (Occ), in-frame occlusions (Occ-in), and out-of-frame occlusions (Occ-out). MatchFlow(G) improves over GMA in every category, but the largest relative improvements generally occur in non-occluded areas, supporting the claim that GIM pretraining improves direct feature correspondence.

    On Sintel clean, GMA versus MatchFlow(G) AEPE is 0.58 versus 0.44 in Noc, 10.58 versus 8.51 in Occ, 7.68 versus 6.42 in Occ-in, 12.52 versus 9.63 in Occ-out, and 1.30 versus 1.03 overall. The corresponding relative improvements are 24.1%, 19.6%, 16.4%, 23.1%, and 20.8%.

    On Sintel final, the values are 1.72 versus 1.46 in Noc, 17.33 versus 15.00 in Occ, 14.96 versus 12.31 in Occ-in, 16.44 versus 15.32 in Occ-out, and 2.74 versus 2.45 overall, corresponding to improvements of 15.1%, 13.4%, 17.7%, 6.8%, and 10.6%. On the Albedo pass, the overall AEPE falls from 1.15 to 0.92, a 20.0% improvement, with Noc AEPE falling from 0.48 to 0.37, a 22.9% improvement.

    Normalized correlation volumes computed over 100 Sintel final-pass images also show a sharper peak at the ground-truth displacement for MatchFlow(G): the central correlation value is 2.0 times higher than GMA in Noc regions and 1.6 times higher in Occ regions.

  10. Knowl 10 — Comparison with alternative real-world representation pretraining

    empirical result

    The paper compares MegaDepth-based GIM pretraining with other real-world or generic representation strategies. With the same optical-flow training baseline on FlyingChairs plus FlyingThings3D, the baseline obtains Sintel clean/final AEPE of 1.27/2.84 and KITTI EPE/Fl-all of 4.12/14.4%. Fine-tuning with virtual MegaDepth flow data as in Depthstill gives Sintel 2.10/3.47 and KITTI 3.41/10.9. Initializing from DINO features gives 1.52/3.08 and 5.50/19.2; initializing from Twins-SVT gives 1.15/2.73 and 4.98/16.8.

    The proposed MegaDepth GIM pretraining gives the best Sintel values, 1.03/2.45, and KITTI values of 4.08/15.6%. Depthstill performs better on KITTI but substantially worse on Sintel, whereas the proposed representation provides stronger cross-dataset generalization to the more complex Sintel motion patterns.

  11. Knowl 11 — Failure modes and stated limitation

    limitation

    MatchFlow does not generalize reliably to all extreme appearance and motion conditions. In reported Sintel final-pass failures, motion blur causes the model to miss part of a weapon, and the motion of a shadow is incorrectly estimated. The paper conjectures that insufficient training exposure to blur-heavy data limits the effectiveness of the attention-based feature extractor in these cases. It therefore identifies training with more diverse final-type data as a possible way to improve robustness.

Coverage note — Exhaustive rows for every competing method, qualitative example images, and the detailed KITTI resolution-ablation matrix were omitted because they corroborate the main quantitative comparisons rather than add separate load-bearing contributions.

References

  1. 1.Filippo Aleotti, Matteo Poggi, and Stefano Mattoccia. Learning optical flow from still images. In 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 15196–15206, 2021. 3, 6, 7
  2. 2.Yoshua Bengio, Jer´ ome Louradour, Ronan Collobert, and Jason Weston. Curriculum learning. In Proceedings of the 26th annual international conference on machine learning, pages 41–48, 2009. 2, 3
  3. 3.Michael J Black and Padmanabhan Anandan. A framework for the robust estimation of optical flow. In 1993 (4th) International Conference on Computer Vision, pages 231–236. IEEE, 1993. 2
  4. 4.Michael J Black and Paul Anandan. The robust estimation of multiple motions: Parametric and piecewise-smooth flow fields. Computer vision and image understanding, 63(1):75–104, 1996. 2
  5. 5.Thomas Brox, Andres Bruhn, Nils Papenberg, and Joachim Weickert. High accuracy optical flow estimation based on a theory for warping. In European conference on computer vision, pages 25–36. Springer, 2004. 2
  6. 6.Andres Bruhn, Joachim Weickert, and Christoph Schnorr. Lucas/kanade meets horn/schunck: Combining local and global optic flow methods. International journal of computer vision, 61(3):211–231, 2005. 2
  7. 7.Daniel J Butler, Jonas Wulff, Garrett B Stanley, and Michael J Black. A naturalistic open source movie for optical flow evaluation. In European conference on computer vision, pages 611–625. Springer, 2012. 2
  8. 8.Chenjie Cao, Xinlin Ren, and Yanwei Fu. Mvsformer: Learning robust image representations via transformers and temperature-based depth for multi-view stereo. arXiv preprint arXiv:2208.02541, 2022. 5
  9. 9.Mathilde Caron, Hugo Touvron, Ishan Misra, HervA© Jegou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. Emerging properties in self-supervised vision transformers. In 2021 IEEE/CVF International Conference on Computer Vision (ICCV), pages 9630–9640, 2021. 6, 7
  10. 10.Lucy Chai, Michael Gharbi, Eli Shechtman, Phillip Isola, and Richard Zhang. Any-resolution training for high-resolution image synthesis. In Shai Avidan, Gabriel Brostow, Moustapha Cisse, Giovanni Maria Farinella, and Tal Hassner, editors, Computer Vision – ECCV 2022, pages 170–188, Cham, 2022. Springer Nature Switzerland. 5
  11. 11.Hongkai Chen, Zixin Luo, Lei Zhou, Yurun Tian, Mingmin Zhen, Tian Fang, David Mckinnon, Yanghai Tsin, and Long Quan. Aspanformer: Detector-free image matching with adaptive span transformer. arXiv preprint arXiv:2208.14201, 2022. 3, 5
  12. 12.Christopher B Choy, JunYoung Gwak, Silvio Savarese, and Manmohan Chandraker. Universal correspondence network. Advances in neural information processing systems, 29, 2016. 3
  13. 13.Xiangxiang Chu, Zhi Tian, Yuqing Wang, Bo Zhang, Haibing Ren, Xiaolin Wei, Huaxia Xia, and Chunhua Shen. Twins: Revisiting the design of spatial attention in vision transformers. In M. Ranzato, A. Beygelzimer, Y. Dauphin, P.S. Liang, and J. Wortman Vaughan, editors, Advances in Neural Information Processing Systems, volume 34, pages 9355–9366. Curran Associates, Inc., 2021. 6
  14. 14.Angela Dai, Angel X. Chang, Manolis Savva, Maciej Halber, Thomas Funkhouser, and Matthias Nießner. Scannet: Richly-annotated 3d reconstructions of indoor scenes. In Proc. Computer Vision and Pattern Recognition (CVPR), IEEE, 2017. 1, 3
  15. 15.Qiaole Dong, Chenjie Cao, and Yanwei Fu. Incremental transformer structure enhanced image inpainting with masking positional encoding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 11358–11368, 2022. 5
  16. 16.Alexey Dosovitskiy, Philipp Fischer, Eddy Ilg, Philip Hausser, Caner Hazirbas, Vladimir Golkov, Patrick Van Der Smagt, Daniel Cremers, and Thomas Brox. Flownet: Learning optical flow with convolutional networks. In Proceedings of the IEEE international conference on computer vision, pages 2758–2766, 2015. 1, 2, 3
  17. 17.Chen Gao, Ayush Saraf, Jia-Bin Huang, and Johannes Kopf. Flow-edge guided video completion. In European Conference on Computer Vision, pages 713–729. Springer, 2020. 1
  18. 18.Andreas Geiger, Philip Lenz, and Raquel Urtasun. Are we ready for autonomous driving? the kitti vision benchmark suite. In 2012 IEEE conference on computer vision and pattern recognition, pages 3354–3361. IEEE, 2012. 2
  19. 19.Berthold KP Horn and Brian G Schunck. Determining optical flow. Artificial intelligence, 17(1-3):185–203, 1981. 2
  20. 20.Zhaoyang Huang, Xiaoyu Shi, Chao Zhang, Qiang Wang, Ka Chun Cheung, Hongwei Qin, Jifeng Dai, and Hongsheng Li. Flowformer: A transformer architecture for optical flow. arXiv preprint arXiv:2203.16194, 2022. 1, 2, 5, 6, 7, 8
  21. 21.Tak-Wai Hui, Xiaoou Tang, and Chen Change Loy. Liteflownet: A lightweight convolutional neural network for optical flow estimation. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 8981–8989, 2018. 2
  22. 22.Eddy Ilg, Nikolaus Mayer, Tonmoy Saikia, Margret Keuper, Alexey Dosovitskiy, and Thomas Brox. Flownet 2.0: Evolution of optical flow estimation with deep networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2462–2470, 2017. 1, 2, 3, 4, 6
  23. 23.Andrew Jaegle, Sebastian Borgeaud, Jean-Baptiste Alayrac, Carl Doersch, Catalin Ionescu, David Ding, Skanda Koppula, Daniel Zoran, Andrew Brock, Evan Shelhamer, et al. Perceiver io: A general architecture for structured inputs & outputs. arXiv preprint arXiv:2107.14795, 2021. 5, 6
  24. 24.Huaizu Jiang, Deqing Sun, Varun Jampani, Ming-Hsuan Yang, Erik Learned-Miller, and Jan Kautz. Super slomo: High quality estimation of multiple intermediate frames for video interpolation. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 9000–9008, 2018. 1
  25. 25.Shihao Jiang, Dylan Campbell, Yao Lu, Hongdong Li, and Richard Hartley. Learning to estimate hidden motions with global motion aggregation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 9772–9781, 2021. 1, 2, 3, 4, 6, 7, 8
  26. 26.Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and François Fleuret. Transformers are rnns: Fast autoregressive transformers with linear attention. In International Conference on Machine Learning, pages 5156–5165. PMLR, 2020. 3, 5
  27. 27.Zhengqi Li and Noah Snavely. Megadepth: Learning single-view depth prediction from internet photos. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2041–2050, 2018. 1, 3, 4
  28. 28.David G Lowe. Distinctive image features from scale-invariant keypoints. International journal of computer vision, 60(2):91–110, 2004. 2
  29. 29.Ao Luo, Fan Yang, Xin Li, and Shuaicheng Liu. Learning optical flow with kernel patch attention. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8906–8915, 2022. 2, 5, 6
  30. 30.Ao Luo, Fan Yang, Kunming Luo, Xin Li, Haoqiang Fan, and Shuaicheng Liu. Learning optical flow with adaptive graph reasoning. In Proceedings of the AAAI Conference on Artificial Intelligence (AAAI), 2022. 2, 6
  31. 31.Nikolaus Mayer, Eddy Ilg, Philip Hausser, Philipp Fischer, Daniel Cremers, Alexey Dosovitskiy, and Thomas Brox. A large dataset to train convolutional networks for disparity, optical flow, and scene flow estimation. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 4040–4048, 2016. 2, 3
  32. 32.Namuk Park and Songkuk Kim. How do vision transformers work? arXiv preprint arXiv:2202.06709, 2022. 5
  33. 33.Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al. Pytorch: An imperative style, high-performance deep learning library. Advances in neural information processing systems, 32, 2019. 4
  34. 34.Anurag Ranjan and Michael J Black. Optical flow estimation using a spatial pyramid network. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 4161–4170, 2017. 1
  35. 35.Jerome Revaud, Philippe Weinzaepfel, Cesar De Souza, Noe Pion, Gabriela Csurka, Yohann Cabon, and Martin Humenberger. R2d2: repeatable and reliable detector and descriptor. arXiv preprint arXiv:1906.06195, 2019. 2
  36. 36.Jerome Revaud, Philippe Weinzaepfel, Zaid Harchaoui, and Cordelia Schmid. Epicflow: Edge-preserving interpolation of correspondences for optical flow. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 1164–1172, 2015. 2
  37. 37.Ignacio Rocco, Mircea Cimpoi, Relja Arandjelovic, Akihiko Torii, Tomas Pajdla, and Josef Sivic. Neighbourhood consensus networks. Advances in neural information processing systems, 31, 2018. 3, 5
  38. 38.Paul-Edouard Sarlin, Daniel DeTone, Tomasz Malisiewicz, and Andrew Rabinovich. Superglue: Learning feature matching with graph neural networks. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 4938–4947, 2020. 2
  39. 39.Manolis Savva, Angel X Chang, and Pat Hanrahan. Semantically-enriched 3d models for common-sense knowledge. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops, pages 24–31, 2015. 3
  40. 40.Johannes L Schonberger, Enliang Zheng, Jan-Michael Frahm, and Marc Pollefeys. Pixelwise view selection for unstructured multi-view stereo. In European Conference on Computer Vision, pages 501–518. Springer, 2016. 1
  41. 41.Leslie N Smith and Nicholay Topin. Super-convergence: Very fast training of neural networks using large learning rates. In Artificial intelligence and machine learning for multi-domain operations applications, volume 11006, pages 369–386. SPIE, 2019. 4
  42. 42.Xiuchao Sui, Shaohua Li, Xue Geng, Yan Wu, Xinxing Xu, Yong Liu, Rick Goh, and Hongyuan Zhu. Craft: Cross-attentional flow transformer for robust optical flow. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 17602–17611, 2022. 1, 2, 3, 4, 6
  43. 43.Deqing Sun, Daniel Vlasic, Charles Herrmann, Varun Jampani, Michael Krainin, Huiwen Chang, Ramin Zabih, William T Freeman, and Ce Liu. Autoflow: Learning a better training set for optical flow. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10093–10102, 2021. 6
  44. 44.Deqing Sun, Xiaodong Yang, Ming-Yu Liu, and Jan Kautz. Pwc-net: Cnns for optical flow using pyramid, warping, and cost volume. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 8934–8943, 2018. 1, 2, 6
  45. 45.Deqing Sun, Xiaodong Yang, Ming-Yu Liu, and Jan Kautz. Models matter, so does training: An empirical study of cnns for optical flow estimation. IEEE transactions on pattern analysis and machine intelligence, 42(6):1408–1423, 2019. 6
  46. 46.Jiaming Sun, Zehong Shen, Yuang Wang, Hujun Bao, and Xiaowei Zhou. Loftr: Detector-free local feature matching with transformers. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 8922–8931, 2021. 2, 3, 4
  47. 47.Shangkun Sun, Yuanqi Chen, Yu Zhu, Guodong Guo, and Ge Li. Skflow: Learning optical flow with super kernels. arXiv preprint arXiv:2205.14623, 2022. 2, 6
  48. 48.Shuyang Sun, Zhanghui Kuang, Lu Sheng, Wanli Ouyang, and Wei Zhang. Optical flow guided feature: A fast and robust motion representation for video action recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 1390–1399, 2018. 1
  49. 49.Shitao Tang, Jiahui Zhang, Siyu Zhu, and Ping Tan. Quadtree attention for vision transformers. arXiv preprint arXiv:2201.02767, 2022. 2, 3
  50. 50.Zachary Teed and Jia Deng. Raft: Recurrent all-pairs field transforms for optical flow. In European conference on computer vision, pages 402–419. Springer, 2020. 1, 2, 3, 4, 5, 6
  51. 51.Prune Truong, Martin Danelljan, Luc V Gool, and Radu Timofte. Gocor: Bringing globally optimized correspondence volumes into your neural network. Advances in Neural Information Processing Systems, 33:14278–14290, 2020. 3
  52. 52.Prune Truong, Martin Danelljan, and Radu Timofte. Glunet: Global-local universal network for dense flow and correspondences. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 6258–6268, 2020. 1, 3
  53. 53.Michał Tyszkiewicz, Pascal Fua, and Eduard Trulls. Disk: Learning local features with policy gradient. Advances in Neural Information Processing Systems, 33:14254–14265, 2020. 3, 5
  54. 54.Haofei Xu, Jing Zhang, Jianfei Cai, Hamid Rezatofighi, and Dacheng Tao. Gmflow: Learning optical flow via global matching. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8121–8130, 2022. 1, 3, 6
  55. 55.Christopher Zach, Thomas Pock, and Horst Bischof. A duality based approach for realtime tv-l 1 optical flow. In Joint pattern recognition symposium, pages 214–223. Springer, 2007. 2
  56. 56.Feihu Zhang, Oliver J Woodford, Victor Adrian Prisacariu, and Philip HS Torr. Separable flow: Learning motion cost volumes for optical flow estimation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 10807–10817, 2021. 6
  57. 57.Shiyu Zhao, Long Zhao, Zhixing Zhang, Enyu Zhou, and Dimitris Metaxas. Global matching with overlapping attention for optical flow estimation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 17592–17601, 2022. 2, 3, 4, 5, 6
  58. 58.Zihua Zheng, Ni Nie, Zhi Ling, Pengfei Xiong, Jiangyu Liu, Hao Wang, and Jiankun Li. Dip: Deep inverse patch-match for high-resolution optical flow. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8925–8934, 2022. 2, 6

Citation

MLA
Dong, Q., et al. “Rethinking Optical Flow from Geometric Matching Consistent Perspective”. arXiv, 2023, http://arxiv.org/abs/2303.08384v1.
APA
Dong, Q., Cao, C., & Fu, Y. (2023). Rethinking Optical Flow from Geometric Matching Consistent Perspective. arXiv. http://arxiv.org/abs/2303.08384v1
Chicago
Dong, Q., C. Cao, and Y. Fu. 2023. “Rethinking Optical Flow from Geometric Matching Consistent Perspective”. arXiv. http://arxiv.org/abs/2303.08384v1.
Harvard
Dong, Q., Cao, C. and Fu, Y. (2023) “Rethinking Optical Flow from Geometric Matching Consistent Perspective”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2303.08384v1.
Vancouver
1. Dong Q, Cao C, Fu Y (2023) Rethinking Optical Flow from Geometric Matching Consistent Perspective. arXiv

BibTeX

@article{dong2023rethinking,
  title = {Rethinking Optical Flow from Geometric Matching Consistent Perspective},
  author = {Dong, Qiaole and Cao, Chenjie and Fu, Yanwei},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2303.08384v1},
  eprint = {2303.08384}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/