FlowDiffuser: Advancing Optical Flow Estimation with Diffusion Models
Ao LuoXin LiFan YangJiangyu LiuHaoqiang FanShuaicheng Liu
Reformulates optical flow estimation as a conditional generative task by introducing a diffusion model with a specialized recurrent denoising decoder that boosts cross-dataset generalization on benchmarks like Sintel and KITTI.
Optical flow estimation calculates pixel displacement between consecutive video frames, serving as a critical foundation for modern computer vision applications such as autonomous driving and video processing. Prevailing systems predominantly frame this as a direct regression problem, which maps image pairs straight to motion vectors. However, this established paradigm exhibits serious vulnerabilities when applied to complex real-world conditions, including severe motion blur, occlusions, and illumination shifts. Standard generative diffusion models offer a potential alternative by modeling motion distributions directly, but they introduce prohibitive computational bottlenecks and lack task-specific designs for optical flow.
The article introduces FlowDiffuser, an optical flow framework that reformulates optical flow estimation as a conditional generative diffusion task. The objective is to demonstrate that progressively removing noise from an initial flow field using a specialized recurrent decoder significantly improves estimation accuracy and model generalization while remaining computationally efficient.
The researchers designed an architecture centered on a Conditional Recurrent Denoising Decoder that integrates a Hidden State Denoising strategy. Rather than using conventional UNet diffusion decoders that restart each step from scratch, FlowDiffuser connects intermediate latent states within a recurrent update loop. The framework was evaluated across standard synthetic and real-world benchmark datasets, including Sintel and KITTI-2015, under two-frame evaluation protocols against leading competitive baselines.
The experimental findings show that FlowDiffuser achieves leading performance across standard benchmarks, securing an average rank of 1.1 across evaluated metrics. In generalization tests on the KITTI dataset, it reduced the outlier error rate (F1-all) to 11.8% and achieved an end-point error of 3.61. In online testing on Sintel, it surpassed prominent models such as MatchFlow and EMD-Flow by 13.2% and 20.9% in error reduction. Furthermore, integrating the conditional denoising decoder into established base architectures, such as RAFT and GMA, delivered performance improvements of 4.3% to 8.7% with minimal parameter overhead. Ablation experiments also confirmed that the framework achieves optimal accuracy with only three denoising steps, avoiding the high latency typical of standard diffusion networks.
These results indicate that generative formulation provides superior motion modeling under challenging conditions where standard regression fails. By combining task-specific recurrent decoding with diffusion modeling, teams can deploy models that are both robust and computationally practical for latency-sensitive applications. Practitioners can adopt this denoising decoder as a modular upgrade to boost existing optical flow pipelines without overhauling underlying feature extraction architectures.
Organizations developing motion estimation systems should evaluate the open-source FlowDiffuser framework for integration into production pipelines. Future efforts should explore scaling the architecture to multi-frame video inputs and optimizing hardware deployment for real-time edge processing. Although the approach demonstrates strong benchmark performance, users should note that the evaluation is constrained to standard two-frame datasets and synthetic pre-training setups, warranting pilot testing on target domain-specific sensor data before deployment in safety-critical environments.
- Paper: RAFT: Recurrent All-Pairs Field Transforms for Optical Flow, Zachary Teed et al. (2020). RAFT establishes the foundational recurrent refinement and correlation volume architecture that FlowDiffuser directly adapts into a conditional recurrent denoising decoder.
- Paper: Denoising Diffusion Probabilistic Models, Jonathan Ho et al. (2020). This seminal work introduces denoising diffusion probabilistic models, establishing the core noise-to-sample generation process that FlowDiffuser reformulates for optical flow estimation.
- Paper: GMFlow: Learning Optical Flow via Global Matching, Haofei Xu et al. (2022). GMFlow provides important context on modern transformer-based global matching frameworks that serve as comparative baselines and backbones for flow estimation networks.
- Paper: FlowNet 2.0: Evolution of Optical Flow Estimation with Deep Networks, Eddy Ilg et al. (2016). FlowNet 2.0 provides crucial background on deep learning regression paradigms and stacked iterative architectures for optical flow that FlowDiffuser seeks to replace with generative diffusion.
- Paper: PWC-Net: CNNs for Optical Flow Using Pyramid, Warping, and Cost Volume, Deqing Sun et al. (2018). PWC-Net outlines the standard pyramid, warping, and cost volume design patterns utilized across deep regression-based flow estimators.
- Paper: Denoising Diffusion Implicit Models, Jiaming Song et al. (2021). DDIM introduces deterministic, non-Markovian accelerated sampling for diffusion models, which is essential for efficient multi-step reverse denoising.
- Paper: Elucidating the Design Space of Diffusion-Based Generative Models, Tero Karras et al. (2022). This paper elucidates the concrete design choices for noise schedules, network preconditioning, and sampling algorithms in diffusion models.
- Paper: A Large Dataset to Train Convolutional Networks for Disparity, Optical Flow, and Scene Flow Estimation, Nikolaus Mayer et al. (2016). This paper introduces the synthetic FlyingThings3D dataset that serves as the standard pre-training foundation for FlowDiffuser and modern optical flow estimators.
- Paper: A Naturalistic Open Source Movie for Optical Flow Evaluation, Daniel J. Butler et al. (2012). This benchmark paper establishes the MPI-Sintel evaluation dataset used to test generalization and accuracy in FlowDiffuser.
- Paper: MemFlow: Optical Flow Estimation and Prediction with Memory, Qiaole Dong et al. (2024). MemFlow extends recurrent flow estimation by incorporating dynamic memory buffers to aggregate long-range temporal context across consecutive video frames.
- Paper: History-Guided Video Diffusion, Kiwhan Song et al. (2025). History-Guided Video Diffusion advances generative temporal modeling by applying per-frame independent noise and history guidance to ensure multi-frame consistency.
- Paper: Score-Based Diffusion Models in Function Space, Jae Hyun Lim 0001 et al. (2025). This work generalizes score-based generative modeling to infinite-dimensional function spaces and continuous dynamic operators like fluid flows.
- Paper: Adjoint Matching: Fine-tuning Flow and Diffusion Generative Models with Memoryless Stochastic Optimal Control, Carles Domingo-Enrich et al. (2025). Adjoint Matching provides a stochastic optimal control framework for fine-tuning continuous flow matching and diffusion models toward downstream objectives.
