EDVR: Video Restoration With Enhanced Deformable Convolutional Networks

Xintao WangKelvin C. K. ChanKe YuChao DongChen Change Loy

article2019CVPR1,321 citationsNTIRE 2019 Video Restoration and Enhancement Challenge Champion (all four tracks)

Presents EDVR, a video restoration framework that couples pyramid deformable convolution alignment with spatiotemporal attention fusion to resolve large motions and blur, winning all four tracks of the NTIRE19 competition.

Listen

Video restoration tasks, such as enhancing resolution and removing blur, are critical for modern visual systems. However, real-world videos frequently contain severe motion blur, occlusion, and large displacements between frames. Existing approaches struggle because standard motion estimation methods fail under complex motion, and typical frame fusion methods treat all neighboring image regions equally, regardless of blur or alignment quality.

The article evaluates and demonstrates a unified video restoration framework named EDVR (Enhanced Deformable Video Restoration). The main objective is to establish an effective system capable of accurately aligning neighboring frames with large motions and dynamically fusing the most informative visual features across multiple video restoration benchmarks.

The approach introduces two central components: a multi-scale alignment module that uses coarse-to-fine deformable convolutions to align features without requiring separate optical flow estimation, and an attention-based fusion module that dynamically assigns weights to neighboring frames across both space and time. To handle severe blur, a preliminary deblurring module prepares features before alignment, and an optional second-stage network refines the outputs. The system was evaluated across multiple standard benchmarks, including the realistic REDS and Vimeo-90K datasets, across four distinct competitive video restoration tracks.

The findings demonstrate substantial improvements over existing methods. In the NTIRE 2019 competition, the framework won first place across all four tracks, including clean and blurry video super-resolution as well as clean and compressed video deblurring. On the REDS benchmark, the framework achieved an average video super-resolution peak signal-to-noise ratio of 31.09 dB compared to 28.63 dB for the prior leading method, and a deblurring ratio of 34.80 dB compared to 26.98 dB for existing methods. Ablation experiments confirmed that the multi-scale alignment module improved image quality by 0.4 dB without adding computational complexity, while temporal and spatial attention contributed an additional 0.14 dB gain. Furthermore, adding a second refinement stage provided an extra 0.5 dB improvement on highly challenging inputs.

These results show that handling video frame alignment implicitly at the feature level, combined with selective attention during fusion, significantly reduces visual artifacts and improves reconstruction quality over traditional motion compensation. For operational deployments, adopting this unified architecture can lower pipeline complexity by replacing separate optical flow and deblurring steps with an end-to-end learning model.

Organizations seeking to implement advanced video enhancement should adopt feature-level deformable alignment and attention mechanisms as a baseline. For high-demand scenarios with severe blurring, teams should implement the two-stage cascade, weighing the added reconstruction fidelity against the extra computational cost.

A key limitation identified in the article is dataset bias: models trained on one video distribution experienced a performance drop of 0.5 to 1.5 dB when tested on mismatched datasets. Consequently, decision-makers should have high confidence in the framework's algorithmic architecture, but should ensure that models are fine-tuned on representative target data before wide deployment.

arXiv: 1905.02716
Cover for EDVR: Video Restoration With Enhanced Deformable Convolutional Networks

Abstract

Video restoration tasks, including super-resolution, deblurring, etc, are drawing increasing attention in the computer vision community. A challenging benchmark named REDS is released in the NTIRE19 Challenge. This new benchmark challenges existing methods from two aspects: (1) how to align multiple frames given large motions, and (2) how to effectively fuse different frames with diverse motion and blur. In this work, we propose a novel Video Restoration framework with Enhanced Deformable networks, termed EDVR, to address these challenges. First, to handle large motions, we devise a Pyramid, Cascading and Deformable (PCD) alignment module, in which frame alignment is done at the feature level using deformable convolutions in a coarse-to-fine manner. Second, we propose a Temporal and Spatial Attention (TSA) fusion module, in which attention is applied both temporally and spatially, so as to emphasize important features for subsequent restoration. Thanks to these modules, our EDVR wins the champions and outperforms the second place by a large margin in all four tracks in the NTIRE19 video restoration and enhancement challenges. EDVR also demonstrates superior performance to state-of-the-art published methods on video super-resolution and deblurring. The code is available at this https URL.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Methodology
  • 3.1 Overview
  • 3.2 Alignment with Pyramid, Cascading and Deformable Convolution
  • 3.3 Fusion with Temporal and Spatial Attention
  • 3.4 Two-Stage Restoration
  • 4 Experiments
  • 4.1 Training Datasets and Details
  • 4.2 Comparisons with State-of-the-art Methods
  • 4.3 Ablation Studies
  • 4.4 Evaluation on REDS Dataset
  • 5 Conclusion
  • References

Knowls

  1. Knowl 1 — EDVR Framework for Video Restoration

    model/method

    The Enhanced Deformable Video Restoration (EDVR) framework is a deep learning architecture designed for multiple video restoration tasks, such as video super-resolution (4×4\times) and video deblurring.

    Given 2N+12N+1 consecutive low-quality input frames I[t−N:t+N]I_{[t-N:t+N]}, the framework reconstructs a high-quality central reference frame O^t\hat{O}_t targeting ground truth OtO_t through four main stages:

    1. Feature Extraction and Pre-Deblurring: Shallow convolutional layers extract feature representations from each frame. For inputs containing motion blur, a PreDeblur module precedes alignment to refine blurry features and improve alignment accuracy. For high-resolution tasks like video deblurring, inputs are first downsampled via strided convolutions to execute subsequent processing in a lower-resolution feature space.
    2. Feature Alignment: Neighboring frame features Ft+iF_{t+i} (i∈[−N,+N],i≠0i \in [-N, +N], i \neq 0) are aligned to the reference frame feature FtF_t using a Pyramid, Cascading and Deformable (PCD) alignment module at the feature level.
    3. Spatio-Temporal Fusion: Aligned features Ft+iaF^a_{t+i} are dynamically aggregated across temporal and spatial dimensions via a Temporal and Spatial Attention (TSA) fusion module.
    4. Reconstruction and Upsampling: The fused feature representation is processed by a deep reconstruction network comprising cascaded residual blocks (40 residual blocks with 128 feature channels in the primary stage). Finally, spatial upsampling (e.g., sub-pixel convolution) produces the residual image, which is added to a directly upsampled reference image.
  2. Knowl 2 — Pyramid, Cascading and Deformable (PCD) Alignment Module

    model/method

    The PCD alignment module aligns neighboring frame features Ft+iF_{t+i} to the reference frame feature FtF_t at the feature level without relying on explicit optical flow estimation or image warping.

    For a deformable convolution kernel with KK sample locations, fixed coordinates pkp_k, and weights wkw_k, the aligned feature Ft+iaF^a_{t+i} at location p0p_0 is computed via modulated deformable convolution:

    Ft+ia(p0)=∑k=1Kwk⋅Ft+i(p0+pk+Δpk)⋅ΔmkF^a_{t+i}(p_0) = \sum_{k=1}^K w_k \cdot F_{t+i}(p_0 + p_k + \Delta p_k) \cdot \Delta m_k

    where Δpk\Delta p_k represents the learned spatial offset and Δmk∈[0,1]\Delta m_k \in [0, 1] represents the modulation scalar.

    To handle large and complex motions, PCD employs an LL-level feature pyramid (L=3L=3) generated via strided convolutions (downsampling factor of 2 per level) without increasing channel counts:

    1. Pyramidal Estimation: Offsets ΔPt+il\Delta P^l_{t+i} and aligned features (Ft+ia)l(F^a_{t+i})^l at pyramid level ll are estimated using concatenated features alongside the 2×2\times upsampled offsets and features from the coarser level l+1l+1:

    ΔPt+il=f([Ft+il,Ftl],(ΔPt+il+1)↑2)\Delta P^l_{t+i} = f\left( [F^l_{t+i}, F^l_t], (\Delta P^{l+1}_{t+i})^{\uparrow 2} \right)

    (Ft+ia)l=g(DConv(Ft+il,ΔPt+il),((Ft+ia)l+1)↑2)(F^a_{t+i})^l = g\left( \text{DConv}(F^l_{t+i}, \Delta P^l_{t+i}), ((F^a_{t+i})^{l+1})^{\uparrow 2} \right)

    where ff and gg are sub-networks of convolution layers, (⋅)↑2(\cdot)^{\uparrow 2} denotes 2×2\times bilinear upsampling, and DConv\text{DConv} is the deformable convolution.

    1. Cascading Refinement: Following the finest level (l=1l=1) of pyramidal alignment, an additional deformable convolution layer is cascaded to refine the coarsely aligned features to sub-pixel accuracy.

    The entire PCD alignment module is trained end-to-end within the video restoration pipeline without explicit optical flow supervision or pretraining.

  3. Knowl 3 — Temporal and Spatial Attention (TSA) Fusion Module

    model/method

    The Temporal and Spatial Attention (TSA) fusion module dynamically aggregates multi-frame aligned features by assigning pixel-level weights across temporal frames and spatial channels.

    1. Temporal Attention: Computes frame-to-frame similarity between each aligned neighboring feature Ft+iaF^a_{t+i} (i∈[−N,+N]i \in [-N, +N]) and the reference feature FtaF^a_t in an embedded space via convolution layers θ\theta and ϕ\phi:

    h(Ft+ia,Fta)=sigmoid(θ(Ft+ia)Tϕ(Fta))h(F^a_{t+i}, F^a_t) = \text{sigmoid}\left( \theta(F^a_{t+i})^T \phi(F^a_t) \right)

    The resulting spatial-specific temporal attention map h(Ft+ia,Fta)h(F^a_{t+i}, F^a_t) scales the input aligned features element-wise before passing them into a fusion convolution layer:

    F~t+ia=Ft+ia⊙h(Ft+ia,Fta)\tilde{F}^a_{t+i} = F^a_{t+i} \odot h(F^a_{t+i}, F^a_t)

    Ffusion=Conv([F~t−Na,…,F~ta,…,F~t+Na])F_{\text{fusion}} = \text{Conv}\left( [\tilde{F}^a_{t-N}, \dots, \tilde{F}^a_t, \dots, \tilde{F}^a_{t+N}] \right)

    where ⊙\odot is element-wise multiplication and [… ][\dots] denotes channel concatenation.

    1. Spatial Attention: Spatial attention masks are computed from FfusionF_{\text{fusion}} using a multi-scale pyramid structure to expand the receptive field. The fused features are modulated by these masks via element-wise multiplication and residual addition to exploit cross-channel and spatial context.
  4. Knowl 4 — Two-Stage Video Restoration Strategy

    model/method

    To enhance restoration under severe motion blur or distortions, a two-stage restoration architecture cascades a secondary, shallower EDVR network to refine the output frames generated by the first-stage model.

    • Stage 1 (Primary Model): Standard EDVR architecture with 40 residual blocks in the reconstruction network, producing an initial restored frame sequence O^t(1)\hat{O}_t^{(1)}.
    • Stage 2 (Refinement Model): A shallower EDVR network (20 residual blocks in the reconstruction network) takes the restored frames from Stage 1 as input and performs secondary alignment, attention fusion, and reconstruction to produce final output O^t(2)\hat{O}_t^{(2)}.

    This two-stage cascading mitigates the alignment burden on individual frames, removes residual severe blur, and reduces temporal inconsistency across the output video sequence.

  5. Knowl 5 — Training Setup and Charbonnier Loss for EDVR

    experimental setup

    EDVR is optimized using the Charbonnier loss function computed between the reconstructed frame O^t\hat{O}_t and ground truth frame OtO_t:

    L=∥O^t−Ot∥2+ε2\mathcal{L} = \sqrt{\|\hat{O}_t - O_t\|^2 + \varepsilon^2}

    with constant ε=1×10−3\varepsilon = 1 \times 10^{-3}.

    Training Configuration:

    • Input Dimensions: 5 consecutive frames (N=2N=2). Mini-batch size of 32 RGB patches of size 64×6464 \times 64 for Video Super-Resolution and 256×256256 \times 256 for Video Deblurring.
    • Optimization: Adam optimizer with β1=0.9\beta_1 = 0.9, β2=0.999\beta_2 = 0.999, and an initial learning rate of 4×10−44 \times 10^{-4}.
    • Hardware: 8 NVIDIA Titan Xp GPUs.
    • Data Augmentation: Random horizontal flips and 90∘90^\circ rotations.
    • Layer Initialization: Deeper network variants are initialized using parameter weights transferred from shallower models to accelerate convergence.
  6. Knowl 6 — Ablation Study on PCD Alignment and TSA Fusion

    data/table

    An ablation study evaluated on REDS demonstrates the isolated and combined performance gains of the PCD alignment and TSA fusion modules. Experiments use a lightweight model configuration (10 residual blocks in the reconstruction module, 64 feature channels, evaluated on 1280×7201280 \times 720 resolution):

    Model Model 1 Model 2 Model 3 Model 4
    PCD Alignment ×\times (1 DConv) ×\times (4 DConv) ✓\checkmark ✓\checkmark
    TSA Fusion ×\times ×\times ×\times ✓\checkmark
    PSNR (dB) 29.78 29.98 30.39 30.53
    FLOPs 640.2G 932.9G 939.3G 936.5G
    • Replacing a standard 4-layer deformable convolution design (Model 2, TDAN style) with the hierarchical PCD module (Model 3) improves PSNR by +0.41 dB+0.41\text{ dB} at nearly identical computational complexity (939.3G vs 932.9G FLOPs).
    • Incorporating the TSA fusion module (Model 4) adds an additional +0.14 dB+0.14\text{ dB} improvement without increasing FLOPs.
  7. Knowl 7 — Benchmark Evaluation for 4x Video Super-Resolution

    data/table

    Performance comparison of EDVR against state-of-the-art single-image super-resolution (RCAN) and video super-resolution methods across Vid4, Vimeo-90K-T, and REDS4 datasets (4×4\times upscaling):

    Method Vid4 (Y / RGB) Vimeo-90K-T (Y / RGB) REDS4 (RGB)
    Bicubic 23.78 / 22.37 31.32 / 29.79 26.14 / 0.7292
    RCAN 25.46 / 24.02 35.35 / 33.61 28.78 / 0.8200
    TOFlow 25.89 / 24.41 34.83 / 33.08 27.98 / 0.7990
    DUF 27.33 / 25.79 36.37 / 34.33 28.63 / 0.8251
    RBPN 27.12 / – 37.07 / – – / –
    EDVR (Ours) 27.35 / 25.83 37.61 / 35.79 31.09 / 0.8800

    Note: Vid4 and Vimeo-90K-T report PSNR (dB) on luminance (Y) and RGB channels. REDS4 reports PSNR (dB) / SSIM on RGB channels. EDVR outperforms flow-based methods (TOFlow) and implicit motion compensation models (DUF, RBPN), outperforming DUF on REDS4 by +2.46 dB+2.46\text{ dB}.

  8. Knowl 8 — Benchmark Evaluation for Video Deblurring on REDS4

    data/table

    Quantitative comparison of EDVR against state-of-the-art image and video deblurring approaches evaluated on the clean video deblurring track of REDS4 (RGB channels, PSNR [dB] / SSIM):

    Method Clip 000 Clip 011 Clip 015 Clip 020 Average
    DeblurGAN 26.57 / 0.8597 22.37 / 0.6637 26.48 / 0.8258 20.93 / 0.6436 24.09 / 0.7482
    DeepDeblur 29.13 / 0.9024 24.28 / 0.7648 28.58 / 0.8822 22.66 / 0.6493 26.16 / 0.8249
    SRN-Deblur 28.95 / 0.8734 25.48 / 0.7595 29.26 / 0.8706 24.21 / 0.7528 26.98 / 0.8141
    DBN 30.03 / 0.9015 24.28 / 0.7331 29.40 / 0.8878 22.51 / 0.7039 26.55 / 0.8066
    EDVR (Ours) 36.66 / 0.9743 34.33 / 0.9393 36.09 / 0.9542 32.12 / 0.9269 34.80 / 0.9487

    EDVR exceeds the previous state of the art (SRN-Deblur) by +7.82 dB+7.82\text{ dB} PSNR on average across complex dynamic motion blur scenes.

  9. Knowl 9 — NTIRE 2019 Video Restoration Challenge Results and Component Gains

    data/table

    EDVR achieved first place in all four tracks of the NTIRE 2019 Video Restoration and Enhancement Challenge. The top validation scores (PSNR [dB] / SSIM) on the competition test sets are:

    Track EDVR (Ours) 2nd Method 3rd Method 4th Method
    Video SR Clean 31.79 / 0.8962 31.13 / 0.8811 31.00 / 0.8822 30.97 / 0.8804
    Video SR Blur 30.17 / 0.8647 – 27.71 / 0.8067 28.92 / 0.8333
    Deblur Clean 36.96 / 0.9657 35.71 / 0.9522 34.09 / 0.9361 33.71 / 0.9363
    Deblur Compression 31.69 / 0.8783 29.78 / 0.8285 29.63 / 0.8261 29.19 / 0.8190

    On the REDS4 dataset, the progressive contributions of two-stage refinement (-S2) and test-time self-ensemble (+, averaging across 4 rotated/flipped transformations) are:

    • SR Clean: Single EDVR: 31.09/0.880031.09 / 0.8800 →\rightarrow EDVR+: 31.23/0.881831.23 / 0.8818 →\rightarrow EDVR-S2: 31.54/0.888831.54 / 0.8888 →\rightarrow EDVR-S2+: 31.56/0.889131.56 / 0.8891.
    • Deblur Clean: Single EDVR: 34.80/0.948734.80 / 0.9487 →\rightarrow EDVR+: 35.27/0.952635.27 / 0.9526 →\rightarrow EDVR-S2: 36.37/0.963236.37 / 0.9632 →\rightarrow EDVR-S2+: 36.49/0.963936.49 / 0.9639.
  10. Knowl 10 — Cross-Dataset Bias in Video Super-Resolution

    empirical result

    Cross-dataset evaluation shows that video super-resolution networks suffer noticeable performance drops when the training and testing data distributions diverge:

    Training Dataset REDS4 Test Vid4 Test Vimeo-90K-T Test
    REDS (5 frames) 31.09 / 0.8800 25.37 / 0.7956 34.33 / 0.9246
    Vimeo-90K (7 frames) 30.49 / 0.8700 25.83 / 0.8077 35.79 / 0.9374
    Difference (Δ\Delta) +0.60 / +0.0100 -0.46 / -0.0121 -1.46 / -0.0128

    Training on REDS improves REDS4 performance by +0.60 dB+0.60\text{ dB} over training on Vimeo-90K, whereas evaluation on Vimeo-90K-T drops by 1.46 dB1.46\text{ dB} when trained on REDS compared to native Vimeo-90K training, indicating significant distribution shifts across video restoration benchmarks.

Coverage note — None was omitted; all contributed models (EDVR framework, PCD alignment module, TSA fusion module, Two-stage restoration), optimization details, challenge results, benchmark evaluations, ablation studies, and dataset bias analyses are covered.

References

  1. 1.Gedas Bertasius, Lorenzo Torresani, and Jianbo Shi. Object detection in video with spatiotemporal sampling networks. In ECCV, 2018. 2
  2. 2.Jose Caballero, Christian Ledig, Aitken Andrew, Acosta Alejandro, Johannes Totz, Zehan Wang, and Wenzhe Shi. Real-time video super-resolution with spatio-temporal networks and motion compensation. In CVPR, 2017. 1, 2, 5, 6
  3. 3.Jifeng Dai, Haozhi Qi, Yuwen Xiong, Yi Li, Guodong Zhang, Han Hu, and Yichen Wei. Deformable convolutional networks. In ICCV, 2017. 2, 3
  4. 4.Qiqin Dai, Seunghwan Yoo, Armin Kappeler, and Aggelos K Katsaggelos. Dictionary-based multiple frame video super-resolution. In ICIP, 2015. 1
  5. 5.Chao Dong, Chen Change Loy, Kaiming He, and Xiaoou Tang. Learning a deep convolutional network for image super-resolution. In ECCV, 2014. 1, 2
  6. 6.Muhammad Haris, Greg Shakhnarovich, and Norimichi Ukita. Recurrent back-projection network for video super-resolution. In CVPR, 2019. 1, 5, 6
  7. 7.Tak-Wai Hui, Xiaoou Tang, and Chen Change Loy. Liteflownet: A lightweight convolutional neural network for optical flow estimation. In CVPR, 2018. 2, 3
  8. 8.Tak-Wai Hui, Xiaoou Tang, and Chen Change Loy. A lightweight optical flow cnn–revisiting data fidelity and regularization. arXiv preprint arXiv:1903.07414, 2019. 3
  9. 9.Eddy Ilg, Nikolaus Mayer, Tonmoy Saikia, Margret Keuper, Alexey Dosovitskiy, and Thomas Brox. FlowNet 2.0: Evolution of optical flow estimation with deep networks. In CVPR, 2017. 2, 3
  10. 10.Younghyun Jo, Seoung Wug Oh, Jaeyeon Kang, and Seon Joo Kim. Deep video super-resolution network using dynamic upsampling filters without explicit motion compensation. In CVPR, 2018. 1, 2, 5, 6
  11. 11.Armin Kappeler, Seunghwan Yoo, Qiqin Dai, and Aggelos K Katsaggelos. Video super-resolution with convolutional neural networks. IEEE Transactions on Computational Imaging, 2016. 1
  12. 12.Tae Hyun Kim, Seungjun Nah, and Kyoung Mu Lee. Dynamic video deblurring using a locally adaptive blur model. TPAMI, 40(10):2374–2387, 2018. 2
  13. 13.Tae Hyun Kim, Mehdi S M Sajjadi, Michael Hirsch, and Bernhard Schölkopf. Spatio-temporal transformer network for video restoration. In ECCV, 2018. 1
  14. 14.Diederik Kingma and Jimmy Ba. Adam: A method for stochastic optimization. In ICLR, 2015. 5
  15. 15.Orest Kupyn, Volodymyr Budzan, Mykola Mykhailych, Dmytro Mishkin, and Jiri Matas. Deblurgan: Blind motion deblurring using conditional adversarial networks. arXiv preprint arXiv:1711.07064, 2017. 1
  16. 16.Orest Kupyn, Volodymyr Budzan, Mykola Mykhailych, Dmytro Mishkin, and Jiří Matas. Deblurgan: Blind motion deblurring using conditional adversarial networks. In CVPR, 2018. 5, 6
  17. 17.Wei-Sheng Lai, Jia-Bin Huang, Narendra Ahuja, and Ming-Hsuan Yang. Deep laplacian pyramid networks for fast and accurate super-resolution. In CVPR, 2017. 5
  18. 18.Christian Ledig, Lucas Theis, Ferenc Huszár, Jose Caballero, Andrew Cunningham, Alejandro Acosta, Andrew P Aitken, Alykhan Tejani, Johannes Totz, Zehan Wang, et al. Photo-realistic single image super-resolution using a generative adversarial network. In CVPR, 2017. 1, 2
  19. 19.Renjie Liao, Xin Tao, Ruiyu Li, Ziyang Ma, and Jiaya Jia. Video super-resolution via deep draft-ensemble learning. In CVPR, 2015. 1, 5, 6
  20. 20.Bee Lim, Sanghyun Son, Heewon Kim, Seungjun Nah, and Kyoung Mu Lee. Enhanced deep residual networks for single image super-resolution. In CVPRW, 2017. 1, 2, 5, 8
  21. 21.Ce Liu and Deqing Sun. On bayesian adaptive video super resolution. TPAMI, 36(2):346–360, 2014. 5, 6, 7
  22. 22.Ding Liu, Zhaowen Wang, Yuchen Fan, Xianming Liu, Zhangyang Wang, Shiyu Chang, and Thomas Huang. Robust video super-resolution with learned temporal dynamics. In ICCV, 2017. 1, 2
  23. 23.Ding Liu, Bihan Wen, Yuchen Fan, Chen Change Loy, and Thomas S Huang. Non-local recurrent network for image restoration. In NIPS, 2018. 2
  24. 24.Ziyang Ma, Renjie Liao, Xin Tao, Li Xu, Jiaya Jia, and Enhua Wu. Handling motion blur in multi-frame super-resolution. In CVPR, 2015. 2
  25. 25.Pavlo Molchanov, Stephen Tyree, Tero Karras, Timo Aila, and Jan Kautz. Pruning convolutional neural networks for resource efficient inference. arXiv preprint arXiv:1611.06440, 2016. 7
  26. 26.Seungjun Nah, Sungyong Baik, Seokil Hong, Gyeongsik Moon, Sanghyun Son, Radu Timofte, and Kyoung Mu Lee. Ntire 2019 challenges on video deblurring and super-resolution: Dataset and study. In CVPRW, June 2019. 1, 5
  27. 27.Seungjun Nah, Tae Hyun Kim, and Kyoung Mu Lee. Deep multi-scale convolutional neural network for dynamic scene deblurring. In CVPR, 2017. 1, 5, 6
  28. 28.Seungjun Nah, Radu Timofte, Sungyong Baik, Seokil Hong, Gyeongsik Moon, Sanghyun Son, Kyoung Mu Lee, Xintao Wang, Kelvin C.K. Chan, Ke Yu, Chao Dong, Chen Change Loy, et al. Ntire 2019 challenge on video deblurring: methods and results. In CVPRW, 2019. 2, 8
  29. 29.Seungjun Nah, Radu Timofte, Shuhang Gu, Sungyong Baik, Seokil Hong, Gyeongsik Moon, Sanghyun Son, Kyoung Mu Lee, Xintao Wang, Kelvin C.K. Chan, Ke Yu, Chao Dong, Chen Change Loy, et al. Ntire 2019 challenge on video super-resolution: Methods and results. In CVPRW, 2019. 2, 8
  30. 30.Liyuan Pan, Yuchao Dai, Miaomiao Liu, and Fatih Porikli. Simultaneous stereo video deblurring and scene flow estimation. In CVPR, 2017. 2
  31. 31.Anurag Ranjan and Black J. Michael. Optical flow estimation using a spatial pyramid network. arXiv preprint arXiv:1611.00850, 2016. 3
  32. 32.Mehdi S M Sajjadi, Raviteja Vemulapalli, and Matthew Brown. Frame-recurrent video super-resolution. In CVPR, 2018. 1, 2, 5, 6
  33. 33.Oded Shahar, Alon Faktor, and Michal Irani. Space-time super-resolution from a single video. In CVPR, 2011. 1
  34. 34.Shuochen Su, Mauricio Delbracio, Jue Wang, Guillermo Sapiro, Wolfgang Heidrich, and Oliver Wang. Deep video deblurring for hand-held cameras. In CVPR, 2017. 2, 5, 6
  35. 35.Deqing Sun, Xiaodong Yang, Ming-Yu Liu, and Jan Kautz. PWC-Net: CNNs for optical flow using pyramid, warping, and cost volume. In CVPR, 2018. 3, 7
  36. 36.Hiroyuki Takeda, Peyman Milanfar, Matan Protter, and Michael Elad. Super-resolution without explicit subpixel motion estimation. TIP, 18(9):1958–1975, 2009. 1
  37. 37.Xin Tao, Hongyun Gao, Renjie Liao, Jue Wang, and Jiaya Jia. Detail-revealing deep video super-resolution. In CVPR, 2017. 1, 2, 5, 6
  38. 38.Xin Tao, Hongyun Gao, Xiaoyong Shen, Jue Wang, and Jiaya Jia. Scale-recurrent network for deep image deblurring. In CVPR, 2018. 1
  39. 39.Xin Tao, Hongyun Gao, Xiaoyong Shen, Jue Wang, and Jiaya Jia. Scale-recurrent network for deep image deblurring. In CVPR, 2018. 5, 6
  40. 40.Yapeng Tian, Yulun Zhang, Yun Fu, and Chenliang Xu. TDAN: Temporally deformable alignment network for video super-resolution. arXiv preprint arXiv:1812.02898, 2018. 1, 2, 3, 4, 7, 8
  41. 41.Radu Timofte, Eirikur Agustsson, Luc Van Gool, Ming-Hsuan Yang, Lei Zhang, Bee Lim, Sanghyun Son, Heewon Kim, Seungjun Nah, Kyoung Mu Lee, Xintao Wang, Yapeng Tian, Ke Yu, Yulun Zhang, Wu Shixiang, Chao Dong, Yu Qiao, Chen Change Loy, et al. Ntire 2017 challenge on single image super-resolution: Methods and results. In CVPRW, 2017. 1, 2
  42. 42.Radu Timofte, Rasmus Rothe, and Luc Van Gool. Seven ways to improve example-based single image super resolution. In CVPR, 2016. 8
  43. 43.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In NIPS, 2017. 2
  44. 44.Xiaolong Wang, Ross Girshick, Abhinav Gupta, and Kaiming He. Non-local neural networks. In CVPR, 2018. 2
  45. 45.Xintao Wang, Ke Yu, Chao Dong, and Chen Change Loy. Recovering realistic texture in image super-resolution by deep spatial feature transform. In CVPR, 2018. 1, 4
  46. 46.Xintao Wang, Ke Yu, Shixiang Wu, Jinjin Gu, Yihao Liu, Chao Dong, Yu Qiao, and Chen Change Loy. Esrgan: Enhanced super-resolution generative adversarial networks. In ECCVW, 2018. 3
  47. 47.Sanghyun Woo, Jongchan Park, Joon-Young Lee, and In So Kweon. Cbam: Convolutional block attention module. In ECCV, 2018. 2
  48. 48.Tianfan Xue, Baian Chen, Jiajun Wu, Donglai Wei, and William T Freeman. Video enhancement with task-oriented flow. arXiv preprint arXiv:1711.09078, 2017. 1, 2, 4, 5, 6, 7
  49. 49.Ke Yu, Chao Dong, Liang Lin, and Chen Change Loy. Crafting a toolchain for image restoration by deep reinforcement learning. In CVPR, 2018. 2
  50. 50.Ke Yu, Xintao Wang, Chao Dong, Xiaoou Tang, and Chen Change Loy. Path-restore: Learning network path selection for image restoration. arXiv preprint arXiv:1904.10343, 2019. 2
  51. 51.Kaihao Zhang, Wenhan Luo, Yiran Zhong, Lin Ma, Wei Liu, and Hongdong Li. Adversarial spatio-temporal learning for video deblurring. TIP, 28(1):291–301, 2019. 2
  52. 52.Yulun Zhang, Kunpeng Li, Kai Li, Lichen Wang, Bineng Zhong, and Yun Fu. Image super-resolution using very deep residual channel attention networks. In ECCV, 2018. 1, 2, 3, 5, 6
  53. 53.Yue Zhao, Yuanjun Xiong, and Dahua Lin. Trajectory convolution for action recognition. In NIPS, 2018. 2
  54. 54.Xizhou Zhu, Han Hu, Stephen Lin, and Jifeng Dai. Deformable convnets v2: More deformable, better results. arXiv preprint arXiv:1811.11168, 2018. 3

Citation

MLA
Wang, X., et al. “EDVR: Video Restoration with Enhanced Deformable Convolutional Networks”. arXiv, 2019, http://arxiv.org/abs/1905.02716v1.
APA
Wang, X., Chan, K. C. K., Yu, K., Dong, C., & Loy, C. C. (2019). EDVR: Video Restoration with Enhanced Deformable Convolutional Networks. arXiv. http://arxiv.org/abs/1905.02716v1
Chicago
Wang, X., K. C. K. Chan, K. Yu, C. Dong, and C. C. Loy. 2019. “EDVR: Video Restoration with Enhanced Deformable Convolutional Networks”. arXiv. http://arxiv.org/abs/1905.02716v1.
Harvard
Wang, X. et al. (2019) “EDVR: Video Restoration with Enhanced Deformable Convolutional Networks”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1905.02716v1.
Vancouver
1. Wang X, Chan KCK, Yu K, Dong C, Loy CC (2019) EDVR: Video Restoration with Enhanced Deformable Convolutional Networks. arXiv

BibTeX

@article{wang2019edvr,
  title = {EDVR: Video Restoration with Enhanced Deformable Convolutional Networks},
  author = {Wang, Xintao and Chan, Kelvin C. K. and Yu, Ke and Dong, Chao and Loy, Chen Change},
  year = {2019},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1905.02716v1},
  eprint = {1905.02716}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE