Flow-Guided Sparse Transformer for Video Deblurring

Jing LinYuanhao CaiXiaowan HuHaoqian WangYouliang YanXueyi ZouHenghui DingYulun ZhangRadu TimofteLuc Van Gool

article2022ICML80 citations

Proposes a flow-guided sparse Transformer that uses optical flow offsets to accurately sample sharp non-local patches across neighboring frames, overcoming the blur artifacts and computational overhead of traditional video restoration models.

Listen

Video recorded on hand-held cameras and mobile devices frequently suffers from motion blur caused by rapid movement and camera shake. Removing blur is essential for downstream applications like autonomous driving, automated tracking, and video stabilization. Traditional restoration models and standard convolutional neural networks struggle to capture distant context across frames. Meanwhile, standard Transformer architectures either incur prohibitive computing costs or restrict their analysis to small windows that miss critical details during rapid motion.

The article designs and evaluates a video restoration framework called the Flow-Guided Sparse Transformer (FGST). Its main objective is to demonstrate that integrating estimated motion cues directly into a sparse self-attention mechanism and using recurrent connections enables higher-quality video deblurring with lower computational overhead.

To evaluate this framework, the authors conducted comprehensive experiments across two established video restoration benchmarks, the DVD and GOPRO datasets, comprising thousands of blurry and sharp image pairs. They also tested real-world footage captured with hand-held devices. The approach pairs an optical flow estimator with a window-based sparse attention module, which selectively samples relevant patches from adjacent frames without altering original image textures through traditional image warping. The model also incorporates a recurrent embedding module to transfer historical frame context sequentially.

The findings show that FGST outperforms previous state-of-the-art models on standard quality metrics. On the DVD benchmark, it achieved a peak signal-to-noise ratio of 33.36 dB, outperforming the previous leading method by 0.56 dB. On the GOPRO benchmark, it achieved 32.90 dB, surpassing existing models by up to 1.23 dB. In terms of efficiency, FGST required approximately 40% fewer parameters, about 63% fewer floating-point operations, and ran more than twice as fast during inference compared to leading alternatives. Ablation studies confirmed that combining flow-guided window sampling with recurrent embeddings provided a 1.72 dB boost over the baseline model.

These results indicate that video restoration systems can achieve superior visual clarity, cleaner edges, and fewer artifacts without escalating computational budgets or hardware costs. Unlike prior methods that degrade textures by pre-warping frames, sampling directly guided by motion preserves critical image fidelity. The modular design also allows performance to scale upward when paired with more accurate motion estimation tools.

Organizations developing computer vision pipelines should consider adopting flow-guided sparse Transformer architectures for video enhancement tasks. When implementing the framework, engineering teams should use a 3x3 window configuration to balance receptive field coverage and processing efficiency, and select advanced optical flow estimators to maximize restoration quality.

A limitation of the study is that quantitative evaluations rely primarily on synthetic and controlled benchmark datasets, whereas real-world footage lacks objective ground-truth references for numerical benchmarking. Additionally, performance depends on the accuracy of the underlying motion estimator. Confidence in the reported performance gains remains high given consistent improvements across multiple benchmarks and architectural configurations.

Cover for Flow-Guided Sparse Transformer for Video Deblurring

Abstract

Exploiting similar and sharper scene patches in spatio-temporal neighborhoods is critical for video deblurring. However, CNN-based methods show limitations in capturing long-range dependencies and modeling non-local self-similarity. In this paper, we propose a novel framework, Flow-Guided Sparse Transformer (FGST), for video deblurring. In FGST, we customize a self-attention module, Flow-Guided Sparse Window-based Multi-head Self-Attention (FGSW-MSA). For each query element on the blurry reference frame, FGSW-MSA enjoys the guidance of the estimated optical flow to globally sample spatially sparse yet highly related key elements corresponding to the same scene patch in neighboring frames. Besides, we present a Recurrent Embedding (RE) mechanism to transfer information from past frames and strengthen long-range temporal dependencies. Comprehensive experiments demonstrate that our proposed FGST outperforms state-of-the-art (SOTA) methods on both DVD and GOPRO datasets and yields visually pleasant results in real video deblurring. https://github.com/linjing7/VR-Baseline

Citation

MLA
Lin, J., et al. “Flow-Guided Sparse Transformer for Video Deblurring”. International Conference on Machine Learning, vol. 162, 2022, pp. 13334–43, https://proceedings.mlr.press/v162/lin22a.html.
APA
Lin, J., Cai, Y., Hu, X., Wang, H., Yan, Y., Zou, X., Ding, H., Zhang, Y., Timofte, R., & Gool, L. V. (2022). Flow-Guided Sparse Transformer for Video Deblurring. International Conference on Machine Learning, 162, 13334–13343. https://proceedings.mlr.press/v162/lin22a.html
Chicago
Lin, J., Y. Cai, X. Hu, et al. 2022. “Flow-Guided Sparse Transformer for Video Deblurring”. International Conference on Machine Learning 162: 13334–43. https://proceedings.mlr.press/v162/lin22a.html.
Harvard
Lin, J. et al. (2022) “Flow-Guided Sparse Transformer for Video Deblurring”, International Conference on Machine Learning. PMLR, pp. 13334–13343. Available at: https://proceedings.mlr.press/v162/lin22a.html.
Vancouver
1. Lin J, Cai Y, Hu X, Wang H, Yan Y, Zou X, Ding H, Zhang Y, Timofte R, Gool LV (2022) Flow-Guided Sparse Transformer for Video Deblurring. In: International Conference on Machine Learning. PMLR, pp 13334–13343

BibTeX

@InProceedings{pmlr-v162-lin22a,
  title = 	 {Flow-Guided Sparse Transformer for Video Deblurring},
  author =       {Lin, Jing and Cai, Yuanhao and Hu, Xiaowan and Wang, Haoqian and Yan, Youliang and Zou, Xueyi and Ding, Henghui and Zhang, Yulun and Timofte, Radu and Van Gool, Luc},
  booktitle = 	 {Proceedings of the 39th International Conference on Machine Learning},
  pages = 	 {13334--13343},
  year = 	 {2022},
  editor = 	 {Chaudhuri, Kamalika and Jegelka, Stefanie and Song, Le and Szepesvari, Csaba and Niu, Gang and Sabato, Sivan},
  volume = 	 {162},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {17--23 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v162/lin22a/lin22a.pdf},
  url = 	 {https://proceedings.mlr.press/v162/lin22a.html},
  abstract = 	 {Exploiting similar and sharper scene patches in spatio-temporal neighborhoods is critical for video deblurring. However, CNN-based methods show limitations in capturing long-range dependencies and modeling non-local self-similarity. In this paper, we propose a novel framework, Flow-Guided Sparse Transformer (FGST), for video deblurring. In FGST, we customize a self-attention module, Flow-Guided Sparse Window-based Multi-head Self-Attention (FGSW-MSA). For each $query$ element on the blurry reference frame, FGSW-MSA enjoys the guidance of the estimated optical flow to globally sample spatially sparse yet highly related $key$ elements corresponding to the same scene patch in neighboring frames. Besides, we present a Recurrent Embedding (RE) mechanism to transfer information from past frames and strengthen long-range temporal dependencies. Comprehensive experiments demonstrate that our proposed FGST outperforms state-of-the-art (SOTA) methods on both DVD and GOPRO datasets and yields visually pleasant results in real video deblurring. https://github.com/linjing7/VR-Baseline}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/