FlowDiffuser: Advancing Optical Flow Estimation with Diffusion Models

Ao LuoXin LiFan YangJiangyu LiuHaoqiang FanShuaicheng Liu

article2024CVPR52 citations

Reformulates optical flow estimation as a conditional generative task by introducing a diffusion model with a specialized recurrent denoising decoder that boosts cross-dataset generalization on benchmarks like Sintel and KITTI.

Listen

Optical flow estimation calculates pixel displacement between consecutive video frames, serving as a critical foundation for modern computer vision applications such as autonomous driving and video processing. Prevailing systems predominantly frame this as a direct regression problem, which maps image pairs straight to motion vectors. However, this established paradigm exhibits serious vulnerabilities when applied to complex real-world conditions, including severe motion blur, occlusions, and illumination shifts. Standard generative diffusion models offer a potential alternative by modeling motion distributions directly, but they introduce prohibitive computational bottlenecks and lack task-specific designs for optical flow.

The article introduces FlowDiffuser, an optical flow framework that reformulates optical flow estimation as a conditional generative diffusion task. The objective is to demonstrate that progressively removing noise from an initial flow field using a specialized recurrent decoder significantly improves estimation accuracy and model generalization while remaining computationally efficient.

The researchers designed an architecture centered on a Conditional Recurrent Denoising Decoder that integrates a Hidden State Denoising strategy. Rather than using conventional UNet diffusion decoders that restart each step from scratch, FlowDiffuser connects intermediate latent states within a recurrent update loop. The framework was evaluated across standard synthetic and real-world benchmark datasets, including Sintel and KITTI-2015, under two-frame evaluation protocols against leading competitive baselines.

The experimental findings show that FlowDiffuser achieves leading performance across standard benchmarks, securing an average rank of 1.1 across evaluated metrics. In generalization tests on the KITTI dataset, it reduced the outlier error rate (F1-all) to 11.8% and achieved an end-point error of 3.61. In online testing on Sintel, it surpassed prominent models such as MatchFlow and EMD-Flow by 13.2% and 20.9% in error reduction. Furthermore, integrating the conditional denoising decoder into established base architectures, such as RAFT and GMA, delivered performance improvements of 4.3% to 8.7% with minimal parameter overhead. Ablation experiments also confirmed that the framework achieves optimal accuracy with only three denoising steps, avoiding the high latency typical of standard diffusion networks.

These results indicate that generative formulation provides superior motion modeling under challenging conditions where standard regression fails. By combining task-specific recurrent decoding with diffusion modeling, teams can deploy models that are both robust and computationally practical for latency-sensitive applications. Practitioners can adopt this denoising decoder as a modular upgrade to boost existing optical flow pipelines without overhauling underlying feature extraction architectures.

Organizations developing motion estimation systems should evaluate the open-source FlowDiffuser framework for integration into production pipelines. Future efforts should explore scaling the architecture to multi-frame video inputs and optimizing hardware deployment for real-time edge processing. Although the approach demonstrates strong benchmark performance, users should note that the evaluation is constrained to standard two-frame datasets and synthetic pre-training setups, warranting pilot testing on target domain-specific sensor data before deployment in safety-critical environments.

Cover for FlowDiffuser: Advancing Optical Flow Estimation with Diffusion Models

Abstract

Optical flow estimation, a process of predicting pixel-wise displacement between consecutive frames, has commonly been approached as a regression task in the age of deep learning. Despite notable advancements, this de facto paradigm unfortunately falls short in generalization performance when trained on synthetic or constrained data. Pioneering a paradigm shift, we reformulate optical flow estimation as a conditional flow generation challenge, unveiling FlowDiffuser — a new family of optical flow models that could have stronger learning and generalization capabilities. FlowDiffuser estimates optical flow through a ‘noise-to-flow’ strategy, progressively eliminating noise from randomly generated flows conditioned on the provided pairs. To optimize accuracy and efficiency, our FlowDiffuser incorporates a novel Conditional Recurrent Denoising Decoder (Conditional-RDD), streamlining the flow estimation process. It incorporates a unique Hidden State Denoising (HSD) paradigm, effectively leveraging the information from previous time steps. Moreover, FlowDiffuser can be easily integrated into existing flow networks, leading to significant improvements in performance metrics compared to conventional implementations. Experiments on challenging benchmarks, including Sintel and KITTI, demonstrate the effectiveness of our FlowDiffuser with superior performance to existing state-of-the-art models. Code is available at https://github.com/LA30/FlowDiffuser.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Method
  • 3.1. Preliminaries
  • 3.2. Diffusion Model for Optical Flow
  • 3.3. Conditional-RDD with Hidden State Denoising
  • 4. Experiments
  • 4.1. Implementation Details
  • 4.2. Benchmarking on Optical Flow Datasets
  • 4.3. Ablation Study
  • 5. Conclusion
  • References

Knowls

  1. Knowl 1 — Conditional flow generation replaces direct regression

    model/method

    FlowDiffuser reformulates two-frame optical-flow estimation as conditional generation rather than direct regression. Given consecutive images I1I_1 and I2I_2, the model transforms a noisy flow field fnf_n into a clean flow field f0f_0 using learned parameters Θ\Theta:

    f0=PΘ(fn∣I1,I2).f_0=P_\Theta(f_n\mid I_1,I_2).

    During training, fnf_n is obtained by adding Gaussian noise to the ground-truth flow. During inference, fnf_n is initialized from a standard Gaussian distribution, and the model progressively removes noise while conditioning on features extracted from I1I_1 and I2I_2. This formulation is intended to represent multiple plausible motion trajectories and to improve robustness to complex motion, occlusion, blur, and brightness changes.

  2. Knowl 2 — Conditional recurrent denoising decoder

    model/method

    FlowDiffuser uses a RAFT-like encoder–decoder architecture whose decoder performs one conditional denoising update at a time. Two image encoders produce basic features x1x_1 and x2x_2, while a context encoder produces xcx_c. A 4D correlation volume xcvx_{cv} is constructed from x1x_1 and x2x_2 using dot products. At timestep tt, the decoder receives the current noisy flow ftf_t, the correlation volume xcvx_{cv}, the context feature xcx_c, and the recurrent hidden feature xhx_h, and predicts a less noisy flow:

    ft−1=Pθ(ft∣xcv,xc,xh).f_{t-1}=P_\theta(f_t\mid x_{cv},x_c,x_h).

    The decoder looks up the correlation pyramid using the current flow, encodes the resulting motion features, updates a recurrent GRU state, and applies a flow head. A timestep embedding is supplied so that the decoder can distinguish noise levels at different diffusion timesteps. The resulting Conditional Recurrent Denoising Decoder can be inserted into RAFT-like decoders such as RAFT, GMA, KPA-Flow, and SKFlow.

  3. Knowl 3 — Embedding Enhancement for timestep conditioning

    model/method

    FlowDiffuser introduces an Embedding Enhancement module to strengthen the effect of diffusion timestep information. Let xm=FME(xcv,f)x_m=F_{ME}(x_{cv},f) be motion features obtained by encoding the correlation lookup at flow ff, and let A\mathcal{A} be an attentive motion-aggregation operator. For timestep tt, the time embedding et=T(t)e_t=T(t) is split into a scale vector esce_{sc} and a shift vector eshe_{sh}. The module computes

    xa=A(xc,xm),xe=C1(xa)esc+esh,xo=C2(xe) τ+xa,\begin{aligned} x_a&=\mathcal{A}(x_c,x_m),\\ x_e&=\mathcal{C}_1(x_a)e_{sc}+e_{sh},\\ x_o&=\mathcal{C}_2(x_e)\,\tau+x_a, \end{aligned}

    where C1\mathcal{C}_1 and C2\mathcal{C}_2 are convolutional blocks consisting of a 3×33\times3 convolution, GELU activation, and group normalization, and τ\tau is a learnable scalar or channel-wise weight. The enhanced feature xox_o is passed to the GRU and flow head. The design is compatible with attentive aggregation modules such as GMA and KPA.

  4. Knowl 4 — Forward diffusion for optical-flow fields

    equation

    During training, FlowDiffuser generates a noisy flow ftf_t from a ground-truth flow f0f_0 using a Gaussian diffusion process. For timestep t∈{0,1,…,T}t\in\{0,1,\ldots,T\}, the conditional distribution is

    q(ft∣f0)=N(ft∣αˉt f0,(1−αˉt)I),q(f_t\mid f_0)=\mathcal{N}\left(f_t\mid\sqrt{\bar\alpha_t}\,f_0,(1-\bar\alpha_t)\mathbf{I}\right),

    where I\mathbf{I} is the identity covariance over all flow pixels and channels, βs\beta_s is the Gaussian noise-variance schedule, αs=1−βs\alpha_s=1-\beta_s, and αˉt=∏s=1tαs\bar\alpha_t=\prod_{s=1}^{t}\alpha_s. Before diffusion, the ground-truth flow is normalized using the image height and width to lie in [−1,1][-1,1] and then rescaled to [−b,b][-b,b]. The default scale factor is b=0.5b=0.5.

  5. Knowl 5 — Reverse denoising with a DDIM-style update

    equation

    At inference time, FlowDiffuser starts from a random Gaussian flow and repeatedly applies a non-Markovian reverse denoising update. Let fθ(t)f_\theta^{(t)} be the decoder's estimate of the clean flow f0f_0 from the current noisy flow ftf_t, let ϵt\epsilon_t be standard Gaussian noise with the same shape as the flow, let σt≥0\sigma_t\geq0 control reverse-process stochasticity, and let αt\alpha_t denote the cumulative signal coefficient used by the reverse schedule. The update is

    ft−1=αt−1fθ(t)+1−αt−1−σt2 ϵ~t+σtϵt,f_{t-1}=\sqrt{\alpha_{t-1}}f_\theta^{(t)}+\sqrt{1-\alpha_{t-1}-\sigma_t^2}\,\widetilde\epsilon_t+\sigma_t\epsilon_t,

    where the estimated noise is

    ϵ~t=ft−αtfθ(t)1−αt.\widetilde\epsilon_t=\frac{f_t-\sqrt{\alpha_t}f_\theta^{(t)}}{\sqrt{1-\alpha_t}}.

    Setting σt=0\sigma_t=0 for every timestep gives the deterministic DDIM-style process used by the paper's stable inference configuration. Repeated updates produce the trajectory fT→⋯→f0f_T\rightarrow\cdots\rightarrow f_0.

  6. Knowl 6 — Hidden State Denoising

    model/method

    Hidden State Denoising (HSD) reduces the instability and computational cost of applying diffusion updates with a recurrent optical-flow decoder. A subnetwork GG predicts a latent intermediate flow state gθ(t)g_\theta^{(t)} at timestep tt. This state is produced by training the recurrent decoder with more internal iterations than are used at inference, so gθ(t)g_\theta^{(t)} is a lower-iteration approximation to the decoder's full clean-flow prediction fθ(t)f_\theta^{(t)}.

    HSD substitutes gθ(t)g_\theta^{(t)} for fθ(t)f_\theta^{(t)} in the reverse denoising update and applies a striding factor λ\lambda to the resulting flow:

    fˉt−1=λft−1.\bar f_{t-1}=\lambda f_{t-1}.

    The strided intermediate state is then used in the next denoising step. The paper uses λ=0.2\lambda=0.2 and K=3K=3 reverse denoising steps by default. HSD is designed to reduce error accumulation from directly predicting long-range changes from highly noisy flows while retaining the recurrent decoder's motion-refinement capability.

  7. Knowl 7 — Signal-prediction training objective

    equation

    Rather than training the decoder to predict the diffusion noise ϵt\epsilon_t, FlowDiffuser trains it to predict the flow signal itself. Let fgtf_{\mathrm{gt}} be the ground-truth optical flow, let c\mathbf{c} denote the conditioning information from the image pair and its encoded correlation/context features, let tt be a sampled diffusion timestep, and let fˉ0′\bar f'_0 be the final conditional denoising result produced with HSD. The training loss is

    L=Ef0∼q(f0∣c), t∼[1,T][∥fgt−fˉ0′∥1].\mathcal{L}=\mathbb{E}_{f_0\sim q(f_0\mid\mathbf{c}),\,t\sim[1,T]}\left[\left\|f_{\mathrm{gt}}-\bar f'_0\right\|_1\right].

    Thus, optimization directly penalizes the L1 difference between the denoised flow and the ground-truth flow while the input flow at each training timestep is generated by the forward diffusion process.

  8. Knowl 8 — Training and evaluation configuration

    experimental setup

    The default FlowDiffuser uses Twins-SVT image encoders and a RAFT-based recurrent decoder with N=12N=12 internal iterations. Its diffusion settings are scale factor b=0.5b=0.5, HSD striding factor λ=0.2\lambda=0.2, and K=3K=3 reverse denoising steps. Training uses batch size 6, AdamW, and a one-cycle learning-rate schedule. The model is first pretrained on FlyingChairs and FlyingThings and then fine-tuned on Sintel, KITTI-2015, and HD1K. The authors evaluate both the standard Chairs-plus-Things configuration, denoted C+T, and the AutoFlow-plus-Things configuration, denoted AF+T. Evaluation and online testing use a single GPU with batch size 1.

  9. Knowl 9 — State-of-the-art benchmark performance

    data/table

    The quantitative benchmark comparison reported on page 6 evaluates two-frame optical-flow models on Sintel and KITTI-2015. Sintel Clean and Final report average endpoint error (EPE, in pixels) on the clean and final rendering passes. KITTI reports EPE and F1-all, the percentage of outlier pixels. Train-set metrics measure generalization, test-set metrics measure online performance, and lower values are better. The table below reproduces the proposed model and the strongest average-rank competitor shown in the paper.

    Could not parse LaTeX table

    FlowDiffuser obtains the best value on 6 of the 7 listed metrics and the best overall average rank, including Sintel-train EPE of 0.860.86, KITTI-train EPE of 3.613.61, KITTI-train F1-all of 11.8%11.8\%, Sintel-test Clean EPE of 1.021.02, and KITTI-test F1-all of 4.17%4.17\%. The asterisk denotes models trained without the standard C+T setup.

  10. Knowl 10 — Component ablations and compatibility with existing decoders

    empirical result

    The ablations reported on pages 7–8 show that the proposed denoising components improve a discriminative baseline with little parameter overhead. On AF+T training, the baseline obtains Sintel Clean/Final EPE of 0.98/2.430.98/2.43 and KITTI EPE/F1-all of 4.18/14.54.18/14.5, while FlowDiffuser without HSD obtains 0.93/2.310.93/2.31 and 3.92/13.93.92/13.9. Adding HSD improves these values to 0.86/2.190.86/2.19 and 3.61/11.83.61/11.8; the parameter count is 14.914.9M for the baseline and 16.316.3M for the FlowDiffuser variants. Increasing reverse denoising steps from K=1K=1 to K=2,3,4K=2,3,4 gives (0.96,2.39,4.43,15.7)(0.96,2.39,4.43,15.7), (0.89,2.25,3.86,12.4)(0.89,2.25,3.86,12.4), (0.86,2.19,3.61,11.8)(0.86,2.19,3.61,11.8), and (0.86,2.20,3.59,11.7)(0.86,2.20,3.59,11.7) for Sintel Clean, Sintel Final, KITTI EPE, and KITTI F1-all, respectively; consequently, K=3K=3 is selected as the default because K=4K=4 provides little additional gain. Removing Embedding Enhancement degrades the same four metrics from (0.86,2.19,3.61,11.8)(0.86,2.19,3.61,11.8) to (0.92,2.30,3.85,12.6)(0.92,2.30,3.85,12.6). Changing bb from the default 0.50.5 to 0.10.1 or 11 gives (0.87,2.21,3.71,12.0)(0.87,2.21,3.71,12.0) or (0.89,2.23,3.75,12.2)(0.89,2.23,3.75,12.2), indicating limited sensitivity to this scale factor.

    Conditional-RDD also improves several existing RAFT-like decoders under C+T training. Relative to RAFT, GMA, KPA-Flow, and SKFlow, the corresponding FlowDiffuser variants achieve Sintel Clean/Final and KITTI EPE/F1-all of (1.29,2.63,4.51,16.2)(1.29,2.63,4.51,16.2), (1.26,2.59,4.32,15.6)(1.26,2.59,4.32,15.6), (1.24,2.45,3.97,14.7)(1.24,2.45,3.97,14.7), and (1.18,2.39,3.91,14.2)(1.18,2.39,3.91,14.2), respectively. The reported average gains are 6.4%/8.7%6.4\%/8.7\%, 4.3%/8.3%4.3\%/8.3\%, 5.8%/9.3%5.8\%/9.3\%, and 3.1%/8.4%3.1\%/8.4\% on Sintel/KITTI. Finally, changing the training data from C+T to AF+T improves FlowDiffuser from (0.89,2.38,3.84,12.7)(0.89,2.38,3.84,12.7) to (0.86,2.19,3.61,11.8)(0.86,2.19,3.61,11.8) on Sintel Clean, Sintel Final, KITTI EPE, and KITTI F1-all.

Coverage note — The complete rows for every competing method and the qualitative prediction and variance-map visualizations were omitted because the benchmark knowl preserves the paper's central numerical claims, while the ablation knowl preserves the component-level evidence.

References

  1. 1.Hyemin Ahn, Esteve Valls Mascaro, and Dongheui Lee. Can we use diffusion probabilistic models for 3d motion prediction? Preprint arXiv:2302.14503, 2023. 3
  2. 2.Min Bai, Wenjie Luo, Kaustav Kundu, and Raquel Urtasun. Exploiting semantic information and deep matching for optical flow. In ECCV, 2016. 2
  3. 3.Emmanuel Asiedu Brempong, Simon Kornblith, Ting Chen, Niki Parmar, Matthias Minderer, and Mohammad Norouzi. Denoising pretraining for semantic segmentation. In CVPR, 2022. 3
  4. 4.Daniel Butler, Jonas Wulff, Garrett Stanley, and Michael Black. A naturalistic open source movie for optical flow evaluation. In ECCV, 2012. 6
  5. 5.Shoufa Chen, Peize Sun, Yibing Song, and Ping Luo. Diffusiondet: Diffusion model for object detection. Preprint arXiv:2211.09788, 2022. 3, 4, 8
  6. 6.Ting Chen, Lala Li, Saurabh Saxena, Geoffrey Hinton, and David J Fleet. A generalist framework for panoptic segmentation of images and videos. Preprint arXiv:2210.06366, 2022. 4, 8
  7. 7.Jiaxin Cheng, Xiao Liang, Xingjian Shi, Tong He, Tianjun Xiao, and Mu Li. Layoutdiffuse: Adapting foundational diffusion models for layout-to-image generation. Preprint arXiv:2302.08908, 2023. 2, 3
  8. 8.Changxing Deng, Ao Luo, Haibin Huang, Shaodan Ma, Jiangyu Liu, and Shuaicheng Liu. Explicit motion disentangling for efficient optical flow estimation. In ICCV, 2023. 1, 6
  9. 9.Prafulla Dhariwal and Alexander Nichol. Diffusion models beat gans on image synthesis. 2021. 2, 3
  10. 10.Qiaole Dong, Chenjie Cao, and Yanwei Fu. Rethinking optical flow from geometric matching consistent perspective. In CVPR, 2023. 6, 7
  11. 11.A. Dosovitskiy, P. Fischer, Eddy Ilg, Philip Hausser, Caner ¨Hazirbas, V. Golkov, P. V. D. Smagt, D. Cremers, and T. Brox. Flownet: Learning optical flow with convolutional networks. In ICCV, 2015. 6
  12. 12.Alexey Dosovitskiy, Philipp Fischer, Eddy Ilg, Philip Hausser, Caner Hazirbas, Vladimir Golkov, Patrick van der Smagt, Daniel Cremers, and Thomas Brox. Flownet: Learning optical flow with convolutional networks. In ICCV, 2015. 1, 2
  13. 13.Zhangxuan Gu, Haoxing Chen, Zhuoer Xu, Jun Lan, Changhua Meng, and Weiqiang Wang. Diffusioninst: Diffusion model for instance segmentation. Preprint arXiv:2212.02773, 2022. 3
  14. 14.Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. In NeurIPS, 2020. 3, 4, 5, 8
  15. 15.Zhaoyang Huang, Xiaoyu Shi, Chao Zhang, Qiang Wang, Ka Chun Cheung, Hongwei Qin, Jifeng Dai, and Hongsheng Li. Flowformer: A transformer architecture for optical flow. In ECCV, 2022. 1, 5, 6, 7
  16. 16.Tak-Wai Hui, Xiaoou Tang, and Chen Change Loy. Liteflownet: A lightweight convolutional neural network for optical flow estimation. In CVPR, 2018. 2
  17. 17.Junhwa Hur and S. Roth. Iterative residual refinement for joint optical flow and occlusion estimation. In CVPR, 2019. 1, 2
  18. 18.Jisoo Jeong, Jamie Menjay Lin, Fatih Porikli, and Nojun Kwak. Imposing consistency for optical flow estimation. In CVPR, 2022. 6
  19. 19.Shihao Jiang, Dylan Campbell, Yao Lu, Hongdong Li, and Richard Hartley. Learning to estimate hidden motions with global motion aggregation. In ICCV, 2021. 1, 2, 3, 4, 5, 6, 7, 8
  20. 20.D. Kondermann, Rahul Nair, Katrin Honauer, Karsten Krispin, Jonas Andrulis, Alexander Brock, Burkhard Gussefeld, Mohsen Rahimimoghaddam, Sabine Hofmann, ¨C. Brenner, and B. Jähne. The hci benchmark suite: Stereo and flow ground truth with uncertainties for urban autonomous driving. In CVPRW, 2016. 6
  21. 21.Bo Li, Xiaolin Wei, Fengwei Chen, and Bin Liu. 3d colored shape reconstruction from a single rgb image through diffusion. Preprint arXiv:2302.05573, 2023. 3
  22. 22.Haipeng Li, Hai Jiang, Ao Luo, Ping Tan, Haoqiang Fan, Bing Zeng, and Shuaicheng Liu. Dmhomo: Learning homography with diffusion models. ACM Transactions on Graphics, 2024. 3
  23. 23.Zhengqi Li and Noah Snavely. Megadepth: Learning single-view depth prediction from internet photos. In CVPP, 2018. 7
  24. 24.Yawen Lu, Qifan Wang, Siqi Ma, Tong Geng, Yingjie Victor Chen, Huaijin Chen, and Dongfang Liu. Transflow: Transformer as flow learner. In CVPR, 2023. 6, 7
  25. 25.Ao Luo, Fan Yang, Xin Li, and Shuaicheng Liu. Learning optical flow with kernel patch attention. In CVPR, 2022. 1, 2, 3, 4, 5, 6, 7, 8
  26. 26.Ao Luo, Fan Yang, Kunming Luo, Xin Li, Haoqiang Fan, and Shuaicheng Liu. Learning optical flow with adaptive graph reasoning. In AAAI, 2022. 2
  27. 27.Ao Luo, Fan Yang, Xin Li, Lang Nie, Chunyu Lin, Haoqiang Fan, and Shuaicheng Liu. Gaflow: Incorporating gaussian attention into optical flow. In ICCV, 2023. 6
  28. 28.N. Mayer, Eddy Ilg, Philip Hausser, P. Fischer, D. Cremers, ¨A. Dosovitskiy, and T. Brox. A large dataset to train convolutional networks for disparity, optical flow, and scene flow estimation. In CVPR, 2016. 6
  29. 29.Luke Melas-Kyriazi, Christian Rupprecht, and Andrea Vedaldi. pcˆ2: Projection-conditioned point cloud diffusion for single-image 3d reconstruction. Preprint arXiv:2302.10668, 2023. 3
  30. 30.Moritz Menze and Andreas Geiger. Object scene flow for autonomous vehicles. In CVPR, 2015. 6
  31. 31.Saurabh Saxena, Charles Herrmann, Junhwa Hur, Abhishek Kar, Mohammad Norouzi, Deqing Sun, and David J Fleet. The surprising effectiveness of diffusion models for optical flow and monocular depth estimation. Preprint arXiv:2306.01923, 2023. 2, 3, 4, 6, 7
  32. 32.Xiaoyu Shi, Zhaoyang Huang, Dasong Li, Manyuan Zhang, Ka Chun Cheung, Simon See, Hongwei Qin, Jifeng Dai, and Hongsheng Li. Flowformer++: Masked cost volume autoencoding for pretraining optical flow estimation. In CVPR, 2023. 5, 6, 7
  33. 33.Jiaming Song, Chenlin Meng, and Stefano Ermon. Denoising diffusion implicit models. In ICLR, 2021. 3, 4, 5
  34. 34.Yang Song and Stefano Ermon. Generative modeling by estimating gradients of the data distribution. 2019. 2, 3
  35. 35.Xiuchao Sui, Shaohua Li, Xue Geng, Yan Wu, Xinxing Xu, Yong Liu, Rick Goh, and Hongyuan Zhu. Craft: Cross-attentional flow transformer for robust optical flow. In CVPR, 2022. 6
  36. 36.Deqing Sun, Xiaodong Yang, Ming-Yu Liu, and Jan Kautz. Pwc-net: Cnns for optical flow using pyramid, warping, and cost volume. In CVPR, 2018. 1, 2
  37. 37.Deqing Sun, Daniel Vlasic, Charles Herrmann, Varun Jampani, Michael Krainin, Huiwen Chang, Ramin Zabih, William T Freeman, and Ce Liu. Autoflow: Learning a better training set for optical flow. In CVPR, 2021. 6, 7
  38. 38.Deqing Sun, Charles Herrmann, Fitsum Reda, Michael Rubinstein, David J Fleet, and William T Freeman. Disentangling architecture and training for optical flow. In ECCV, 2022. 6, 7
  39. 39.Shangkun Sun, Yuanqi Chen, Yu Zhu, Guodong Guo, and Ge Li. Skflow: Learning optical flow with super kernels. In NeurIPS, 2022. 3, 4, 5, 6, 7, 8
  40. 40.Haoru Tan, Sitong Wu, and Jimin Pi. Semantic diffusion network for semantic segmentation. Preprint arXiv:2302.02057, 2023. 3
  41. 41.Zachary Teed and Jun Deng. Raft: Recurrent all-pairs field transforms for optical flow. In ECCV, 2020. 1, 2, 3, 5, 6, 7, 8
  42. 42.Guy Tevet, Sigal Raab, Brian Gordon, Yoni Shafir, Daniel Cohen-or, and Amit Haim Bermano. Human motion diffusion model. In ICLR, 2022. 5
  43. 43.Xiaodong Wang, Chenfei Wu, Shengming Yin, Minheng Ni, Jianfeng Wang, Linjie Li, Zhengyuan Yang, Fan Yang, Lijuan Wang, Zicheng Liu, et al. Learning 3d photography videos via self-supervised diffusion on single images. Preprint arXiv:2302.10781, 2023. 2, 3
  44. 44.Philippe Weinzaepfel, Thomas Lucas, Vincent Leroy, Yohann Cabon, Vaibhav Arora, Romain Bregier, Gabriela ´Csurka, Leonid Antsfeld, Boris Chidlovskii, and Jer´ome ˆRevaud. Croco v2: Improved cross-view completion pretraining for stereo matching and optical flow. In ICCV, 2023. 7
  45. 45.Haofei Xu, Jing Zhang, Jianfei Cai, Hamid Rezatofighi, and Dacheng Tao. Gmflow: Learning optical flow via global matching. In CVPR, 2022. 1, 6
  46. 46.Haofei Xu, Jing Zhang, Jianfei Cai, Hamid Rezatofighi, Fisher Yu, Dacheng Tao, and Andreas Geiger. Unifying flow, stereo and depth estimation. TPAMI, 2023. 6, 7
  47. 47.Ning Xu, Linjie Yang, Yuchen Fan, Dingcheng Yue, Yuchen Liang, Jianchao Yang, and Thomas Huang. Youtube-vos: A large-scale video object segmentation benchmark. Preprint arXiv:1809.03327, 2018. 7
  48. 48.Feihu Zhang, Oliver J Woodford, Victor Adrian Prisacariu, and Philip HS Torr. Separable flow: Learning motion cost volumes for optical flow estimation. In ICCV, 2021. 6
  49. 49.Chengqian Zhao, Cheng Feng, Dengwang Li, and Shuo Li. Of-msrn: optical flow-auxiliary multi-task regression network for direct quantitative measurement, segmentation and motion estimation. In AAAI, 2020. 1, 2
  50. 50.Shengyu Zhao, Yilun Sheng, Yue Dong, E. Chang, and Yan Xu. Maskflownet: Asymmetric feature matching with learnable occlusion mask. In CVPR, 2020. 1, 2
  51. 51.Shiyu Zhao, Long Zhao, Zhixing Zhang, Enyu Zhou, and Dimitris Metaxas. Global matching with overlapping attention for optical flow estimation. In CVPR, 2022. 1, 6

Citation

MLA
Luo, A., et al. “FlowDiffuser: Advancing Optical Flow Estimation with Diffusion Models”. 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024, pp. 19167–76, https://doi.org/10.1109/CVPR52733.2024.01813.
APA
Luo, A., Li, X., Yang, F., Liu, J., Fan, H., & Liu, S. (2024). FlowDiffuser: Advancing Optical Flow Estimation with Diffusion Models. 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 19167–19176. https://doi.org/10.1109/CVPR52733.2024.01813
Chicago
Luo, A., X. Li, F. Yang, J. Liu, H. Fan, and S. Liu. 2024. “FlowDiffuser: Advancing Optical Flow Estimation with Diffusion Models”. 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 19167–76. https://doi.org/10.1109/CVPR52733.2024.01813.
Harvard
Luo, A. et al. (2024) “FlowDiffuser: Advancing Optical Flow Estimation with Diffusion Models”, 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp. 19167–19176. Available at: https://doi.org/10.1109/CVPR52733.2024.01813.
Vancouver
1. Luo A, Li X, Yang F, Liu J, Fan H, Liu S (2024) FlowDiffuser: Advancing Optical Flow Estimation with Diffusion Models. In: 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp 19167–19176

BibTeX

@inproceedings{Luo_2024, title={FlowDiffuser: Advancing Optical Flow Estimation with Diffusion Models}, url={http://dx.doi.org/10.1109/CVPR52733.2024.01813}, DOI={10.1109/cvpr52733.2024.01813}, booktitle={2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)}, publisher={IEEE}, author={Luo, Ao and Li, Xin and Yang, Fan and Liu, Jiangyu and Fan, Haoqiang and Liu, Shuaicheng}, year={2024}, month=June, pages={19167–19176} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE