Recurrent Video Restoration Transformer with Guided Deformable Attention
Jingyun LiangYuchen FanXiaoyu XiangRakesh RanjanEddy IlgSimon GreenJiezhang CaoKai ZhangRadu TimofteLuc Van Gool
Proposes a hybrid video restoration transformer that combines clip-level parallel processing with global recurrence and guided deformable attention, achieving state-of-the-art super-resolution, deblurring, and denoising performance while maintaining low memory consumption and runtime.
Restoring high-quality video from degraded, blurry, or noisy footage is critical for modern technologies such as live streaming, surveillance systems, and archival media restoration. Existing restoration methods face a fundamental operational trade-off: parallel transformer-based systems achieve high restoration quality by processing all frames at once but demand excessive memory and computing power, whereas sequential recurrent methods use fewer resources but suffer from information loss, noise buildup, and slower linear processing.
The article demonstrates a hybrid deep learning framework, named the Recurrent Video Restoration Transformer (RVRT), designed to achieve an optimal balance among restoration quality, model size, and operational efficiency. The approach divides a video into small multi-frame clips, processing frames within each clip in parallel while linking consecutive clips through a globally recurrent framework that features a novel guided deformable attention mechanism to align video clips smoothly in a single step.
The authors evaluated the framework across eight standard benchmark datasets covering three major restoration tasks: video super-resolution, deblurring, and denoising. The experimental results show that the proposed system establishes state-of-the-art performance across all tasks. Compared to standard parallel transformer models, the framework reduces parameter size and testing memory usage by over 50% while cutting runtime by at least 25% to up to 85% in deblurring and denoising tasks. Furthermore, it consistently outperforms existing recurrent networks in restoration accuracy, improving peak signal-to-noise ratio by 0.2 to 0.5 decibels on key super-resolution benchmarks and mitigating frame error propagation.
These findings indicate that organizations deploying automated video restoration can achieve top-tier visual clarity without incurring prohibitive infrastructure costs or latency penalties. The architecture significantly lowers memory consumption and processing times, making high-quality video enhancement practical for cost-sensitive and near-real-time production pipelines.
Decision-makers should consider piloting the open-source framework for compute-constrained video workflows such as media upscaling and streaming pipelines. Future technical development should focus on building end-to-end video-level motion estimation to eliminate computational overhead in larger clip configurations, while deployment teams should remain mindful of ethical and privacy risks when clarifying real-world surveillance footage.
- Paper: EDVR: Video Restoration With Enhanced Deformable Convolutional Networks, Xintao Wang et al. (2019). This seminal work establishes deformable convolutional alignment and temporal fusion for video restoration, providing foundational architectural principles that RVRT adapts into guided deformable attention.
- Paper: SwinIR: Image Restoration Using Swin Transformer, Jingyun Liang et al. (2021). This paper establishes the application of window-based vision transformers to low-level vision restoration tasks, which directly informs RVRT's clip-level parallel transformer design.
- Paper: Video Enhancement with Task-Oriented Flow, Tianfan Xue et al. (2017). It introduces task-oriented motion estimation alongside the benchmark Vimeo-90K dataset used to evaluate and train recurrent video restoration models like RVRT.
- Paper: Scale-Recurrent Network for Deep Image Deblurring, Xin Tao et al. (2018). It demonstrates recurrent state propagation in deep restoration networks, offering foundational context for combining recurrent connections with multi-frame processing.
- Paper: Uformer: A General U-Shaped Transformer for Image Restoration, Zhendong Wang et al. (2021). It introduces locally enhanced window-based transformer architectures for image restoration, addressing the computational bottlenecks of standard self-attention.
- Paper: Pre-Trained Image Processing Transformer, Hanting Chen et al. (2020). It lays the groundwork for applying large-scale transformer architectures to low-level vision and image restoration benchmarks.
- Paper: Flow-Guided Sparse Transformer for Video Deblurring, Jing Lin et al. (2022). This work advances motion-guided attention in video restoration by integrating sparse flow-guided self-attention into a recurrent video deblurring framework.
- Paper: Learning Trajectory-Aware Transformer for Video Super-Resolution, Chengxu Liu et al. (2022). This paper explores an alternative trajectory-aware attention paradigm across video sequences, expanding on long-range temporal modeling for video super-resolution.
- Paper: Dual-Domain Attention for Image Deblurring, Yuning Cui et al. (2023). It builds on efficient attention designs for restoration by extending spatial attention into the frequency domain for deblurring.
