BasicVSR++: Improving Video Super-Resolution with Enhanced Propagation and Alignment
Kelvin C. K. ChanShangchen ZhouXiangyu XuChen Change Loy
Proposes second-order grid propagation and flow-guided deformable alignment to dramatically improve video super-resolution performance while maintaining high computational efficiency.
Video super-resolution enhances low-resolution video into clear, high-resolution footage by combining complementary details across moving, misaligned frames. While recurrent neural networks offer compact and efficient architectures for this task, existing designs struggle to effectively propagate information over long sequences and accurately align misaligned features, especially in occluded or complex regions.
The article demonstrates an enhanced recurrent architecture, named BasicVSR++, designed to significantly improve video super-resolution accuracy and detail preservation without increasing computational costs. It evaluates whether refining feature propagation across time and stabilizing feature alignment can surpass existing high-capacity and transformer-based methods.
To achieve this, the authors introduced two core redesigns: second-order grid propagation, which alternates forward and backward information flow while connecting frames across multiple time steps, and flow-guided deformable alignment, which uses estimated optical motion as a base guide to stably learn flexible alignment offsets. The approach was evaluated using standard benchmark video datasets (REDS and Vimeo-90K) under various degradation settings, and was compared against 17 leading restoration models as well as competition standards.
The evaluations yielded several major findings. BasicVSR++ set a new state of the art across all standard benchmarks, outperforming large-capacity models like EDVR by up to 1.3 decibels in peak signal-to-noise ratio while requiring 65% fewer parameters. Compared to its direct predecessor, BasicVSR, the redesigned model achieved a significant 0.82-decibel gain with nearly identical model size and runtime. The architecture also outperformed transformer-based models while operating 18 times faster and using 78% fewer parameters. In visual comparisons, the method successfully recovered fine textures and details in occluded regions that caused blurring or temporal flickering in other approaches. Furthermore, the architecture won three champions and one first runner-up in the NTIRE 2021 video restoration challenge for compressed video enhancement.
These results demonstrate that architectural refinements in temporal propagation and alignment provide far greater performance gains than simply scaling up model size. For organizations handling video processing, streaming, or enhancement, this method reduces computational overhead, lowers hardware costs, and speeds up processing pipelines while delivering superior visual consistency.
Organizations developing or deploying video restoration systems should adopt second-order grid propagation and flow-guided alignment as foundational design principles for super-resolution, compressed video enhancement, and related restoration tasks. However, practitioners should note that the model experiences quality degradation when processing severely degraded real-world footage. Additional research and pilot testing on extreme real-world degradations are recommended before deploying the system in unconstrained, heavily corrupted environments.
- Paper: EDVR: Video Restoration With Enhanced Deformable Convolutional Networks, Xintao Wang et al. (2019). EDVR introduced multi-scale deformable convolutions for temporal frame alignment in video restoration, serving as the direct foundational alignment paradigm that BasicVSR++ redesigns into flow-guided deformable alignment.
- Paper: Video Enhancement with Task-Oriented Flow, Tianfan Xue et al. (2017). TOFlow established task-oriented optical flow estimation and introduced the Vimeo-90K benchmark dataset heavily relied upon and evaluated in BasicVSR++.
- Paper: Real-Time Single Image and Video Super-Resolution Using an Efficient Sub-Pixel Convolutional Neural Network, Wenzhe Shi et al. (2016). ESPCN introduced sub-pixel convolution for efficient low-resolution feature extraction and upscaling, providing a core architectural component used in modern deep super-resolution networks.
- Paper: Image Super-Resolution Using Deep Convolutional Networks, Chao Dong et al. (2014). SRCNN established the fundamental paradigm of end-to-end convolutional neural networks for super-resolution upon which recurrent deep restoration architectures are built.
- Paper: Learning Trajectory-Aware Transformer for Video Super-Resolution, Chengxu Liu et al. (2022). TTVSR benchmarks against and extends the recurrent video super-resolution paradigm by using trajectory-aware transformers to overcome the remaining long-range degradation limits of recurrent models like BasicVSR++.
- Paper: Recurrent Video Restoration Transformer with Guided Deformable Attention, Jingyun Liang et al. (2022). RVRT builds on recurrent video restoration and deformable alignment concepts by integrating guided deformable attention within a hybrid recurrent-transformer architecture.
- Paper: Flow-Guided Sparse Transformer for Video Deblurring, Jing Lin et al. (2022). FGST leverages optical flow guidance and recurrent embedding to drive sparse self-attention for video restoration, extending the flow-guided temporal alignment principles explored in BasicVSR++.
