BasicVSR++: Improving Video Super-Resolution with Enhanced Propagation and Alignment

Kelvin C. K. ChanShangchen ZhouXiangyu XuChen Change Loy

article2022CVPR681 citations3 champions and 1 runner-up in NTIRE 2021 Video Restoration and Enhancement Challenge

Proposes second-order grid propagation and flow-guided deformable alignment to dramatically improve video super-resolution performance while maintaining high computational efficiency.

Listen

Video super-resolution enhances low-resolution video into clear, high-resolution footage by combining complementary details across moving, misaligned frames. While recurrent neural networks offer compact and efficient architectures for this task, existing designs struggle to effectively propagate information over long sequences and accurately align misaligned features, especially in occluded or complex regions.

The article demonstrates an enhanced recurrent architecture, named BasicVSR++, designed to significantly improve video super-resolution accuracy and detail preservation without increasing computational costs. It evaluates whether refining feature propagation across time and stabilizing feature alignment can surpass existing high-capacity and transformer-based methods.

To achieve this, the authors introduced two core redesigns: second-order grid propagation, which alternates forward and backward information flow while connecting frames across multiple time steps, and flow-guided deformable alignment, which uses estimated optical motion as a base guide to stably learn flexible alignment offsets. The approach was evaluated using standard benchmark video datasets (REDS and Vimeo-90K) under various degradation settings, and was compared against 17 leading restoration models as well as competition standards.

The evaluations yielded several major findings. BasicVSR++ set a new state of the art across all standard benchmarks, outperforming large-capacity models like EDVR by up to 1.3 decibels in peak signal-to-noise ratio while requiring 65% fewer parameters. Compared to its direct predecessor, BasicVSR, the redesigned model achieved a significant 0.82-decibel gain with nearly identical model size and runtime. The architecture also outperformed transformer-based models while operating 18 times faster and using 78% fewer parameters. In visual comparisons, the method successfully recovered fine textures and details in occluded regions that caused blurring or temporal flickering in other approaches. Furthermore, the architecture won three champions and one first runner-up in the NTIRE 2021 video restoration challenge for compressed video enhancement.

These results demonstrate that architectural refinements in temporal propagation and alignment provide far greater performance gains than simply scaling up model size. For organizations handling video processing, streaming, or enhancement, this method reduces computational overhead, lowers hardware costs, and speeds up processing pipelines while delivering superior visual consistency.

Organizations developing or deploying video restoration systems should adopt second-order grid propagation and flow-guided alignment as foundational design principles for super-resolution, compressed video enhancement, and related restoration tasks. However, practitioners should note that the model experiences quality degradation when processing severely degraded real-world footage. Additional research and pilot testing on extreme real-world degradations are recommended before deploying the system in unconstrained, heavily corrupted environments.

arXiv: 2104.13371
Cover for BasicVSR++: Improving Video Super-Resolution with Enhanced Propagation and Alignment

Abstract

A recurrent structure is a popular framework choice for the task of video super-resolution. The state-of-the-art method BasicVSR adopts bidirectional propagation with feature alignment to effectively exploit information from the entire input video. In this study, we redesign BasicVSR by proposing second-order grid propagation and flow-guided deformable alignment. We show that by empowering the recurrent framework with enhanced propagation and alignment, one can exploit spatiotemporal information across misaligned video frames more effectively. The new components lead to an improved performance under a similar computational constraint. In particular, our model BasicVSR++ surpasses BasicVSR by a significant 0.82 dB in PSNR with similar number of parameters. BasicVSR++ is generalizable to other video restoration tasks, and obtains three champions and one first runner-up in NTIRE 2021 video restoration challenge.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Methodology
  • 3.1 Second-Order Grid Propagation
  • 3.2 Flow-Guided Deformable Alignment
  • 4 Experiments
  • 4.1 Comparisons with State-of-the-Art Methods
  • 5 Ablation Studies
  • 6 NTIRE 2021 Challenge Results
  • 7 Conclusion
  • References
  • A Network Architecture
  • B Experimental Settings
  • C Qualitative Comparisons

Knowls

  1. Knowl 1 — Second-Order Grid Propagation in Video Super-Resolution

    model/method

    Second-order grid propagation is a recurrent temporal modeling scheme for video super-resolution (VSR) that addresses two key shortcomings of conventional unidirectional and bidirectional recurrent frameworks: the limitation of single-pass propagation and the restriction of first-order Markov temporal connections.

    1. Grid Propagation: Instead of passing features across the temporal sequence once, intermediate features are propagated backward and forward in time across multiple alternating propagation iterations (branches). At each propagation iteration jj, frame features are updated using both the temporally aligned features from neighboring timesteps in iteration jj and the corresponding refined feature from the preceding iteration j−1j-1. This bidirectional grid structure enables the network to repeatedly revisit and refine features across the full sequence.

    2. Second-Order Connections: Conventional recurrent VSR models compute the current state at timestep ii strictly from timestep i−1i-1 (first-order Markov chain). Second-order grid propagation relaxes this by simultaneously aggregating aligned features from both the (i−1)(i-1)-th and (i−2)(i-2)-th timesteps during forward propagation (and (i+1)(i+1)-th and (i+2)(i+2)-th timesteps during backward propagation). This allows direct information flow across wider spatiotemporal baselines, improving feature retrieval in fine-textured and occluded regions.

  2. Knowl 2 — Flow-Guided Deformable Alignment

    model/method

    Flow-guided deformable alignment (FGDA) synergizes optical flow and deformable convolution (DCN) to provide flexible and stable cross-frame feature alignment in video super-resolution.

    Standard deformable alignment directly predicts spatial sampling offsets from feature maps via convolution, which frequently suffers from training instability and offset overflow (divergence). Conversely, standard flow-based alignment is stable but constrained to single-point spatial interpolation, leading to blurry features in complex motion regions.

    FGDA resolves this trade-off:

    1. Optical flow si→i−1s_{i \to i-1} from frame ii to frame i−1i-1 is first estimated using a dedicated flow network.
    2. The candidate feature fi−1f_{i-1} is pre-aligned via bilinear spatial warping W\mathcal{W} using si→i−1s_{i \to i-1}, producing fˉi−1=W(fi−1,si→i−1)\bar{f}_{i-1} = \mathcal{W}(f_{i-1}, s_{i \to i-1}).
    3. The base optical flow serves as a direct spatial anchor. Convolutional layers predict only residual offsets Δo\Delta o and modulation masks mm conditioned on the concatenation of current LR features gig_i and the pre-aligned feature fˉi−1\bar{f}_{i-1}.
    4. The final DCN sampling offset is formed by oi→i−1=si→i−1+Δoo_{i \to i-1} = s_{i \to i-1} + \Delta o.
    5. Deformable convolution is then executed directly on the original unwarped feature fi−1f_{i-1} using oi→i−1o_{i \to i-1} and modulation mask mi→i−1m_{i \to i-1}.

    For second-order alignment, FGDA simultaneously processes warped features fˉi−1\bar{f}_{i-1} and ar{f}_{i-2} jointly by concatenating them with gig_i to predict multi-frame offsets and masks in a single forward pass, preventing redundant computations and exploiting complementary multi-frame cues.

  3. Knowl 3 — Offset and Mask Formulation for Flow-Guided Deformable Alignment

    equation

    In flow-guided deformable alignment, let gi∈RC×H×Wg_i \in \mathbb{R}^{C \times H \times W} denote the feature map extracted from the low-resolution frame xix_i at timestep ii, and let fi−p∈RC×H×Wf_{i-p} \in \mathbb{R}^{C \times H \times W} denote the latent feature at timestep i−pi-p for temporal offsets p∈{1,2}p \in \{1, 2\}. Let si→i−p∈R2×H×Ws_{i \to i-p} \in \mathbb{R}^{2 \times H \times W} denote the estimated optical flow from frame ii to frame i−pi-p.

    Pre-alignment of candidate features is performed via spatial warping W\mathcal{W}: fˉi−p=W(fi−p,si→i−p),p∈{1,2}\bar{f}_{i-p} = \mathcal{W}(f_{i-p}, s_{i \to i-p}), \quad p \in \{1, 2\}

    For second-order alignment, the DCN offsets oi→i−po_{i \to i-p} and modulation masks mi→i−pm_{i \to i-p} for each candidate frame are computed jointly from the concatenated pre-aligned features and current frame features: oi→i−p=si→i−p+Co(c(gi,fˉi−1,fˉi−2))o_{i \to i-p} = s_{i \to i-p} + \mathcal{C}^o(c(g_i, \bar{f}_{i-1}, \bar{f}_{i-2})) mi→i−p=σ(Cm(c(gi,fˉi−1,fˉi−2)))m_{i \to i-p} = \sigma\left(\mathcal{C}^m(c(g_i, \bar{f}_{i-1}, \bar{f}_{i-2}))\right) where c(⋅)c(\cdot) denotes concatenation along the channel dimension, Co\mathcal{C}^o and Cm\mathcal{C}^m denote stacks of convolutional layers predicting offset residues and modulation features respectively, and σ(⋅)\sigma(\cdot) denotes the sigmoid activation function.

    The composite offsets oio_i and masks mim_i are concatenated across temporal orders, and deformable convolution D\mathcal{D} is applied to the concatenated unwarped features: oi=c(oi→i−1,oi→i−2)o_i = c(o_{i \to i-1}, o_{i \to i-2}) mi=c(mi→i−1,mi→i−2)m_i = c(m_{i \to i-1}, m_{i \to i-2}) f^i=D(c(fi−1,fi−2);oi,mi)\hat{f}_i = \mathcal{D}(c(f_{i-1}, f_{i-2}); o_i, m_i)

  4. Knowl 4 — Recurrent State Update in Second-Order Grid Propagation

    equation

    Let xix_i be the ii-th input frame of a video sequence (i∈{1,…,T}i \in \{1, \dots, T\}), gig_i be the shallow feature map extracted from xix_i via residual blocks, and fijf_i^j be the feature map at timestep ii in the jj-th propagation branch (j∈{1,…,N}j \in \{1, \dots, N\}), with fi0=gif_i^0 = g_i.

    For a forward propagation branch jj, the recurrent feature alignment and state update are defined by: f^ij=A(gi,fi−1j,fi−2j,si→i−1,si→i−2)\hat{f}_i^j = \mathcal{A}\left(g_i, f_{i-1}^j, f_{i-2}^j, s_{i \to i-1}, s_{i \to i-2}\right) fij=f^ij+R(c(fij−1,f^ij))f_i^j = \hat{f}_i^j + \mathcal{R}\left(c(f_i^{j-1}, \hat{f}_i^j)\right) where:

    • si→i−1s_{i \to i-1} and si→i−2s_{i \to i-2} are optical flow fields between the respective frames.
    • A\mathcal{A} denotes the second-order flow-guided deformable alignment operation.
    • f^ij\hat{f}_i^j is the aligned spatiotemporal candidate feature.
    • c(⋅)c(\cdot) denotes concatenation along the channel dimension.
    • R\mathcal{R} denotes a stack of residual blocks refining the combined state from the previous grid pass fij−1f_i^{j-1} and aligned state f^ij\hat{f}_i^j.
    • Boundary conditions are defined as s0→−1=s0→−2=s1→−1=f−1j=f−2j=0s_{0 \to -1} = s_{0 \to -2} = s_{1 \to -1} = f_{-1}^j = f_{-2}^j = 0.

    Backward propagation branches are defined symmetrically over timesteps i+1i+1 and i+2i+2 with inverted optical flows.

  5. Knowl 5 — Quantitative Evaluation of BasicVSR++ across VSR Benchmarks

    data/table

    BasicVSR++ was evaluated against 17 video super-resolution methods on four standard benchmark datasets: REDS4, Vimeo-90K-T, Vid4, and UDM10 under 4×4\times super-resolution with Bicubic (BI) and Blur Downsampling (BD) degradation models. Evaluation metrics are PSNR (dB) and SSIM. All PSNR/SSIM scores are calculated on the Y-channel, except REDS4 which is evaluated on RGB channels. Runtime is measured on low-resolution inputs of size 180×320180 \times 320.

    Method Params (M) Runtime (ms) REDS4 (BI) Vimeo-90K-T (BI) Vid4 (BI) UDM10 (BD) Vimeo-90K-T (BD) Vid4 (BD)
    Bicubic - - 26.14/0.7292 31.32/0.8684 23.78/0.6347 28.47/0.8253 31.30/0.8687 21.80/0.5246
    TOFlow - - 27.98/0.7990 33.08/0.9054 25.89/0.7651 36.26/0.9438 34.62/0.9212 -
    FRVSR 5.1 137 - - - 37.09/0.9522 35.64/0.9319 26.69/0.8103
    DUF 5.8 974 28.63/0.8251 - - 38.48/0.9605 36.87/0.9447 27.38/0.8329
    RBPN 12.2 1507 30.09/0.8590 37.07/0.9435 27.12/0.8180 38.66/0.9596 37.20/0.9458 -
    EDVR-M 3.3 118 30.53/0.8699 37.09/0.9446 27.10/0.8186 39.40/0.9663 37.33/0.9484 27.45/0.8406
    EDVR 20.6 378 31.09/0.8800 37.61/0.9489 27.35/0.8264 39.89/0.9686 37.81/0.9523 27.85/0.8503
    PFNL 3.0 295 29.63/0.8502 36.14/0.9363 26.73/0.8029 38.74/0.9627 - 27.16/0.8355
    MuCAN - - 30.88/0.8750 37.32/0.9465 - - - -
    TGA 5.8 - - - - - 37.59/0.9516 27.63/0.8423
    RLSP 4.2 49 - - - 38.48/0.9606 36.49/0.9403 27.48/0.8388
    RSDN 6.2 94 - - - 39.35/0.9653 37.23/0.9471 27.92/0.8505
    RRN 3.4 45 - - - 38.96/0.9644 - 27.69/0.8488
    BasicVSR 6.3 63 31.42/0.8909 37.18/0.9450 27.24/0.8251 39.96/0.9694 37.53/0.9498 27.96/0.8553
    IconVSR 8.7 70 31.67/0.8948 37.47/0.9476 27.39/0.8279 40.03/0.9694 37.84/0.9524 28.04/0.8570
    VSR-Tran 32.6 4312 31.06/0.8815 37.71/0.9494 27.36/0.8258 - - -
    BasicVSR++ 7.3 77 32.39/0.9069 37.79/0.9500 27.79/0.8400 40.72/0.9722 38.21/0.9550 29.04/0.8753

    BasicVSR++ outperforms EDVR by up to 1.30 dB1.30\text{ dB} on REDS4 and 1.19 dB1.19\text{ dB} on Vid4 (BD) while using 65%65\% fewer parameters (7.3 M7.3\text{ M} vs. 20.6 M20.6\text{ M}). Compared to its immediate predecessor BasicVSR, BasicVSR++ achieves a gain of 0.97 dB0.97\text{ dB} on REDS4 and 1.08 dB1.08\text{ dB} on Vid4 (BD).

  6. Knowl 6 — Ablation Study of BasicVSR++ Core Components

    data/table

    An ablation study evaluated the progressive contribution of Flow-Guided Deformable Alignment (FGDA), Second-Order Propagation, and Grid Propagation against a baseline recurrent bidirectional network. Models were trained and evaluated on the REDS4 dataset for 4×4\times super-resolution under Bicubic degradation.

    Component (A) (B) (C) BasicVSR++
    Flow-Guided Deform. Align. ✓ ✓ ✓
    Second-Order Propagation ✓ ✓
    Grid Propagation ✓
    PSNR (dB) 31.48 31.94 32.08 32.39
    • Adding Flow-Guided Deformable Alignment improves PSNR from 31.48 dB31.48\text{ dB} to 31.94 dB31.94\text{ dB} (+0.46 dB+0.46\text{ dB}).
    • Introducing Second-Order temporal propagation increases performance to 32.08 dB32.08\text{ dB} (+0.14 dB+0.14\text{ dB}).
    • Incorporating Grid Propagation across iterations increases performance to 32.39 dB32.39\text{ dB} (+0.31 dB+0.31\text{ dB}).
    • The cumulative improvement of the full BasicVSR++ model over the baseline is 0.91 dB0.91\text{ dB} in PSNR.
  7. Knowl 7 — Ablation of Optical Flow Guidance Mechanisms in Deformable Alignment

    data/table

    To analyze how optical flow should be integrated with deformable alignment, three design variants were compared on the REDS4 dataset:

    1. Without Optical Flow (w/o Flow): Standard deformable alignment where offsets are predicted directly from feature maps.
    2. Offset-Fidelity Loss: Direct offset prediction supervised by optical flow via an auxiliary loss function during training.
    3. Flow-Guided Deformable Alignment (Ours): Optical flow is directly used as a base offset and added to learned residual offsets during both training and inference.
    Alignment Design w/o Flow Offset-Fidelity Loss Ours (FGDA)
    PSNR (dB) 27.44 30.22 32.39

    Without optical flow guidance, offset training collapses due to instability, resulting in a degraded PSNR of 27.44 dB27.44\text{ dB}. Using offset-fidelity loss stabilizes training (30.22 dB30.22\text{ dB}), but directly embedding optical flow as base offsets inside the architecture achieves 32.39 dB32.39\text{ dB}, outperforming the loss-based supervision by 2.17 dB2.17\text{ dB}.

  8. Knowl 8 — Performance and Efficiency of Light-BasicVSR++

    data/table

    A lightweight variant, Light-BasicVSR++, was constructed to match the parameter count and computational runtime of BasicVSR and IconVSR on the REDS4 benchmark (RGB channel PSNR, runtime measured on 180×320180 \times 320 low-resolution inputs for 4×4\times upsampling).

    Metric BasicVSR IconVSR Light-BasicVSR++
    Params (M) 6.3 8.7 6.4
    Runtime (ms) 63 70 69
    PSNR (dB) 31.42 31.67 32.24

    Light-BasicVSR++ achieves 32.24 dB32.24\text{ dB} PSNR with 6.4 M6.4\text{ M} parameters and 69 ms69\text{ ms} runtime, surpassing BasicVSR by 0.82 dB0.82\text{ dB} with a nearly identical parameter count (+0.1 M+0.1\text{ M}) and runtime (+6 ms+6\text{ ms}), and outperforming IconVSR by 0.57 dB0.57\text{ dB} while requiring 26%26\% fewer parameters (6.4 M6.4\text{ M} vs. 8.7 M8.7\text{ M}).

  9. Knowl 9 — Generalization to Compressed Video Enhancement

    data/table

    The BasicVSR++ architecture generalizes directly to video restoration tasks beyond non-blind super-resolution. On the NTIRE 2021 Quality Enhancement of Compressed Video Challenge benchmark, BasicVSR++ was evaluated against top-ranking challenge methods.

    Metric BasicVSR++ Method 1 Method 2 Method 3
    PSNR (dB) 30.37 29.95 29.69 29.64
    SSIM 0.9484 0.9468 0.9423 0.9405

    BasicVSR++ achieves 30.37 dB30.37\text{ dB} PSNR and 0.94840.9484 SSIM, outperforming the winning competition methods (achieving margins of +0.42 dB+0.42\text{ dB}, +0.68 dB+0.68\text{ dB}, and +0.73 dB+0.73\text{ dB} in PSNR over Methods 1, 2, and 3, respectively).

  10. Knowl 10 — Scaling Limits and Real-World Degradation Bottlenecks in BasicVSR++

    limitation

    Two operational limitations are identified for BasicVSR++:

    1. Diminishing Returns on Higher-Order and Grid Iteration Scaling: While transitioning from first-order to second-order propagation improves PSNR by 0.14 dB0.14\text{ dB} and transitioning from 1 to 2 grid iterations improves PSNR by 0.31 dB0.31\text{ dB}, scaling beyond second-order connections (e.g., third-order) or beyond two grid propagation iterations yields marginal gains (approximately 0.05 dB0.05\text{ dB} in PSNR), because the nearest two frames already contain the vast majority of complementary temporal information.
    2. Severe Real-World Degradations: When applied to real-world blind VSR using second-order degradation modeling and aggressive augmentation, BasicVSR++ effectively removes unknown artifacts and enhances character legibility. However, restoration performance deteriorates when input video frames suffer from severe, compound in-the-wild degradations.

Coverage note — None was omitted; all key contributions including network architectures, exact offset and state equations, comprehensive quantitative benchmarks, ablations on alignment and propagation schemes, efficiency models, downstream challenge results, and model limitations are included.

References

  1. 1.Jose Caballero, Christian Ledig, Aitken Andrew, Acosta Alejandro, Johannes Totz, Zehan Wang, and Wenzhe Shi. Real-time video super-resolution with spatio-temporal networks and motion compensation. In CVPR, 2017. 5
  2. 2.Jiezhang Cao, Yawei Li, Kai Zhang, and Luc Van Gool. Video super-resolution transformer. arXiv preprint arXiv:2106.06847, 2021. 5
  3. 3.Kelvin C.K. Chan, Xintao Wang, Ke Yu, Chao Dong, and Chen Change Loy. BasicVSR: The search for essential components in video super-resolution and beyond. In CVPR, 2021. 1, 2, 3, 4, 5, 8
  4. 4.Kelvin C.K. Chan, Xintao Wang, Ke Yu, Chao Dong, and Chen Change Loy. Understanding deformable alignment in video super-resolution. In AAAI, 2021. 2, 4, 7, 8
  5. 5.Kelvin C.K. Chan, Shangchen Zhou, Xiangyu Xu, and Chen Change Loy. Investigating tradeoffs in real-world video super-resolution. In CVPR, 2022. 8
  6. 6.Jifeng Dai, Haozhi Qi, Yuwen Xiong, Yi Li, Guodong Zhang, Han Hu, and Yichen Wei. Deformable convolutional networks. In ICCV, 2017. 2, 4
  7. 7.Damien Fourure, Remi Emonet, Elisa Fromont, Damien Muselet, Alain Tremeau, and Christian Wolf. Residual conv-deconv grid network for semantic segmentation. In BMVC, 2017. 2
  8. 8.Dario Fuoli, Shuhang Gu, and Radu Timofte. Efficient video super-resolution through recurrent latent space propagation. In ICCVW, 2019. 1, 2, 5
  9. 9.Muhammad Haris, Greg Shakhnarovich, and Norimichi Ukita. Recurrent back-projection network for video super-resolution. In CVPR, 2019. 1, 4, 5
  10. 10.Yan Huang, Wei Wang, and Liang Wang. Bidirectional recurrent convolutional networks for multi-frame super-resolution. In NIPS, 2015. 1, 2
  11. 11.Yan Huang, Wei Wang, and Liang Wang. Video super-resolution via bidirectional recurrent convolutional networks. TPAMI, 2018. 1, 2
  12. 12.Takashi Isobe, Xu Jia, Shuhang Gu, Songjiang Li, Shengjin Wang, and Qi Tian. Video super-resolution with recurrent structure-detail network. In ECCV, 2020. 1, 2, 5
  13. 13.Takashi Isobe, Songjiang Li, Xu Jia, Shanxin Yuan, Gregory Slabaugh, Chunjing Xu, Ya-Li Li, Shengjin Wang, and Qi Tian. Video super-resolution with temporal group attention. In CVPR, 2020. 5
  14. 14.Takashi Isobe, Fang Zhu, and Shengjin Wang. Revisiting temporal modeling for video super-resolution. In BMVC, 2020. 1, 2, 5
  15. 15.Younghyun Jo, Seoung Wug Oh, Jaeyeon Kang, and Seon Joo Kim. Deep video super-resolution network using dynamic upsampling filters without explicit motion compensation. In CVPR, 2018. 5
  16. 16.Nan Rosemary Ke, Anirudh Goyal, Olexa Bilaniuk, Jonathan Binas, Michael C Mozer, Chris Pal, and Yoshua Bengio. Sparse attentive backtracking: Temporal creditassignment through reminding. In NIPS, 2018. 2
  17. 17.Wenbo Li, Xin Tao, Taian Guo, Lu Qi, Jiangbo Lu, and Jiaya Jia. MuCAN: Multi-correspondence aggregation network for video super-resolution. In ECCV, 2020. 5
  18. 18.Tsungnan Lin, Bill G Horne, Peter Tino, and C Lee Giles. Learning long-term dependencies in NARX recurrent neural networks. IEEE Transactions on Neural Networks, 1996. 2
  19. 19.Ce Liu and Deqing Sun. On bayesian adaptive video super resolution. TPAMI, 2014. 4, 5
  20. 20.Contributors MMEditing. MMEditing: OpenMMLab Image and Video Editing Toolbox, 3 2022. 5
  21. 21.Seungjun Nah, Sungyong Baik, Seokil Hong, Gyeongsik Moon, Sanghyun Son, Radu Timofte, and Kyoung Mu Lee. NTIRE 2019 challenge on video deblurring and super-resolution: Dataset and study. In CVPRW, 2019. 4, 5
  22. 22.Seungjun Nah, Sanghyun Son, and Kyoung Mu Lee. Recurrent neural networks with intra-frame iterations for video deblurring. In CVPR, 2019. 2
  23. 23.Simon Niklaus and Feng Liu. Softmax splatting for video frame interpolation. In CVPR, 2020. 2
  24. 24.Mehdi S M Sajjadi, Raviteja Vemulapalli, and Matthew Brown. Frame-recurrent video super-resolution. In CVPR, 2018. 1, 2, 5
  25. 25.Rohollah Soltani and Hui Jiang. Higher order recurrent neural networks. arXiv preprint arXiv:1605.00064, 2016. 2
  26. 26.Ke Sun, Yang Zhao, Borui Jiang, Tianheng Cheng, Bin Xiao, Dong Liu, Yadong Mu, Xinggang Wang, Wenyu Liu, and Jingdong Wang. High-resolution representations for labeling pixels and regions. arXiv preprint arXiv:1904.04514, 2019. 2
  27. 27.Xin Tao, Hongyun Gao, Renjie Liao, Jue Wang, and Jiaya Jia. Detail-revealing deep video super-resolution. In CVPR, 2017. 5
  28. 28.Yapeng Tian, Yulun Zhang, Yun Fu, and Chenliang Xu. TDAN: Temporally deformable alignment network for video super-resolution. In CVPR, 2018. 1, 2, 4
  29. 29.Hua Wang, Dewei Su, Longcun Jin, and Chuangchuang Liu. Deformable non-local network for video super-resolution. IEEE Access, 2019. 2, 4
  30. 30.Jingdong Wang, Ke Sun, Tianheng Cheng, Borui Jiang, Chaorui Deng, Yang Zhao, Dong Liu, Yadong Mu, Mingkui Tan, Xinggang Wang, Wenyu Liu, and Bin Xiao. Deep high-resolution representation learning for visual recognition. TPAMI, 2020. 2
  31. 31.Xintao Wang, Kelvin C.K. Chan, Ke Yu, Chao Dong, and Chen Change Loy. EDVR: Video restoration with enhanced deformable convolutional networks. In CVPRW, 2019. 1, 2, 4, 5, 6
  32. 32.Xintao Wang, Liangbin Xie, Chao Dong, and Ying Shan. Real-ESRGAN: Training real-world blind super-resolution with pure synthetic data. In ICCVW, 2021. 8
  33. 33.Xiaoyu Xiang, Yapeng Tian, Yulun Zhang, Yun Fu, Jan P Allebach, and Chenliang Xu. Zooming Slow-Mo: Fast and accurate one-stage space-time video super-resolution. In CVPR, 2020. 2
  34. 34.Xiangyu Xu, Muchen Li, Wenxiu Sun, and Ming-Hsuan Yang. Learning spatial and spatio-temporal pixel aggregations for image and video denoising. TIP, 2020. 4
  35. 35.Tianfan Xue, Baian Chen, Jiajun Wu, Donglai Wei, and William T Freeman. Video enhancement with task-oriented flow. IJCV, 2019. 1, 4, 5, 6
  36. 36.Ren Yang, Radu Timofte, Jing Liu, Yi Xu, Xinjian Zhang, Minyi Zhao, Shuigeng Zhou, Kelvin CK Chan, Shangchen Zhou, Xiangyu Xu, et al. NTIRE 2021 challenge on quality enhancement of compressed video: Methods and results. In CVPRW, 2021. 8
  37. 37.Peng Yi, Zhongyuan Wang, Kui Jiang, Junjun Jiang, and Jiayi Ma. Progressive fusion video super-resolution network via exploiting non-local spatio-temporal correlations. In ICCV, 2019. 4, 5
  38. 38.Shangchen Zhou, Jiawei Zhang, Jinshan Pan, Haozhe Xie, Wangmeng Zuo, and Jimmy Ren. Spatio-temporal filter adaptive network for video deblurring. In ICCV, 2019. 2
  39. 39.Xizhou Zhu, Han Hu, Stephen Lin, and Jifeng Dai. Deformable ConvNets v2: More deformable, better results. In CVPR, 2019. 2, 4
  40. 40.Juntang Zhuang, Junlin Yang, Lin Gu, and Nicha Dvornek. ShelfNet for fast semantic segmentation. In ICCVW, 2019. 2

Citation

MLA
Chan, K. C. K., et al. “BasicVSR++: Improving Video Super-Resolution with Enhanced Propagation and Alignment”. arXiv, 2021, http://arxiv.org/abs/2104.13371v1.
APA
Chan, K. C. K., Zhou, S., Xu, X., & Loy, C. C. (2021). BasicVSR++: Improving Video Super-Resolution with Enhanced Propagation and Alignment. arXiv. http://arxiv.org/abs/2104.13371v1
Chicago
Chan, K. C. K., S. Zhou, X. Xu, and C. C. Loy. 2021. “BasicVSR++: Improving Video Super-Resolution with Enhanced Propagation and Alignment”. arXiv. http://arxiv.org/abs/2104.13371v1.
Harvard
Chan, K.C.K. et al. (2021) “BasicVSR++: Improving Video Super-Resolution with Enhanced Propagation and Alignment”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2104.13371v1.
Vancouver
1. Chan KCK, Zhou S, Xu X, Loy CC (2021) BasicVSR++: Improving Video Super-Resolution with Enhanced Propagation and Alignment. arXiv

BibTeX

@article{chan2021basicvsr,
  title = {BasicVSR++: Improving Video Super-Resolution with Enhanced Propagation and Alignment},
  author = {Chan, Kelvin C. K. and Zhou, Shangchen and Xu, Xiangyu and Loy, Chen Change},
  year = {2021},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2104.13371v1},
  eprint = {2104.13371}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE