Flow-Guided Sparse Transformer for Video Deblurring
Jing LinYuanhao CaiXiaowan HuHaoqian WangYouliang YanXueyi ZouHenghui DingYulun ZhangRadu TimofteLuc Van Gool
Proposes a flow-guided sparse Transformer that uses optical flow offsets to accurately sample sharp non-local patches across neighboring frames, overcoming the blur artifacts and computational overhead of traditional video restoration models.
Video recorded on hand-held cameras and mobile devices frequently suffers from motion blur caused by rapid movement and camera shake. Removing blur is essential for downstream applications like autonomous driving, automated tracking, and video stabilization. Traditional restoration models and standard convolutional neural networks struggle to capture distant context across frames. Meanwhile, standard Transformer architectures either incur prohibitive computing costs or restrict their analysis to small windows that miss critical details during rapid motion.
The article designs and evaluates a video restoration framework called the Flow-Guided Sparse Transformer (FGST). Its main objective is to demonstrate that integrating estimated motion cues directly into a sparse self-attention mechanism and using recurrent connections enables higher-quality video deblurring with lower computational overhead.
To evaluate this framework, the authors conducted comprehensive experiments across two established video restoration benchmarks, the DVD and GOPRO datasets, comprising thousands of blurry and sharp image pairs. They also tested real-world footage captured with hand-held devices. The approach pairs an optical flow estimator with a window-based sparse attention module, which selectively samples relevant patches from adjacent frames without altering original image textures through traditional image warping. The model also incorporates a recurrent embedding module to transfer historical frame context sequentially.
The findings show that FGST outperforms previous state-of-the-art models on standard quality metrics. On the DVD benchmark, it achieved a peak signal-to-noise ratio of 33.36 dB, outperforming the previous leading method by 0.56 dB. On the GOPRO benchmark, it achieved 32.90 dB, surpassing existing models by up to 1.23 dB. In terms of efficiency, FGST required approximately 40% fewer parameters, about 63% fewer floating-point operations, and ran more than twice as fast during inference compared to leading alternatives. Ablation studies confirmed that combining flow-guided window sampling with recurrent embeddings provided a 1.72 dB boost over the baseline model.
These results indicate that video restoration systems can achieve superior visual clarity, cleaner edges, and fewer artifacts without escalating computational budgets or hardware costs. Unlike prior methods that degrade textures by pre-warping frames, sampling directly guided by motion preserves critical image fidelity. The modular design also allows performance to scale upward when paired with more accurate motion estimation tools.
Organizations developing computer vision pipelines should consider adopting flow-guided sparse Transformer architectures for video enhancement tasks. When implementing the framework, engineering teams should use a 3x3 window configuration to balance receptive field coverage and processing efficiency, and select advanced optical flow estimators to maximize restoration quality.
A limitation of the study is that quantitative evaluations rely primarily on synthetic and controlled benchmark datasets, whereas real-world footage lacks objective ground-truth references for numerical benchmarking. Additionally, performance depends on the accuracy of the underlying motion estimator. Confidence in the reported performance gains remains high given consistent improvements across multiple benchmarks and architectural configurations.
- Paper: EDVR: Video Restoration With Enhanced Deformable Convolutional Networks, Xintao Wang et al. (2019). EDVR establishes the foundational multi-frame alignment and spatio-temporal feature fusion principles for video restoration that video deblurring transformers aim to overcome.
- Paper: RAFT: Recurrent All-Pairs Field Transforms for Optical Flow, Zachary Teed et al. (2020). RAFT provides the modern deep optical flow estimation methodology crucial for optical flow-guided sampling in video deblurring.
- Paper: SwinIR: Image Restoration Using Swin Transformer, Jingyun Liang et al. (2021). SwinIR demonstrates how window-based self-attention can be effectively adapted for low-level image restoration tasks, motivating window-based sparse transformer attention.
- Paper: Uformer: A General U-Shaped Transformer for Image Restoration, Zhendong Wang et al. (2021). Uformer introduces locally-enhanced window self-attention for image deblurring and restoration, providing the architecture baseline for transformer-based deblurring.
- Paper: Scale-Recurrent Network for Deep Image Deblurring, Xin Tao et al. (2018). SRN-DeblurNet pioneers deep recurrent mechanisms for propagating structural deblurring information across frames and scales.
- Paper: Video Enhancement with Task-Oriented Flow, Tianfan Xue et al. (2017). TOFlow introduces task-oriented optical flow estimation to guide low-level video restoration and processing.
- Paper: Deep Multi-scale Convolutional Neural Network for Dynamic Scene Deblurring, Seungjun Nah et al. (2017). Nah et al. establish the standard multi-scale dynamic scene deblurring formulation and GOPRO benchmark used to evaluate modern video deblurring models.
- Paper: Non-local Neural Networks, Xiaolong Wang et al. (2018). Non-local Neural Networks formalizes non-local self-similarity and spatio-temporal long-range dependency modeling in video architectures.
- Paper: Restormer: Efficient Transformer for High-Resolution Image Restoration, Syed Waqas Zamir et al. (2022). Restormer extends efficient transformer modeling to high-resolution restoration tasks by introducing transposed self-attention across channels.
- Paper: Cross Aggregation Transformer for Image Restoration, Zheng Chen et al. (2022). Cross Aggregation Transformer expands on window-based attention in image restoration by utilizing rectangular and axial cross-attention to capture long-range contextual structures.
- Paper: Fourier Priors-Guided Diffusion for Zero-Shot Joint Low-Light Enhancement and Deblurring, Xiaoqian Lv et al. (2024). FourierDiff advances the deblurring domain beyond deterministic transformers by incorporating frequency priors into generative diffusion models for low-light deblurring.
