Recurrent Video Restoration Transformer with Guided Deformable Attention

Jingyun LiangYuchen FanXiaoyu XiangRakesh RanjanEddy IlgSimon GreenJiezhang CaoKai ZhangRadu TimofteLuc Van Gool

article2022NeurIPS316 citations

Proposes a hybrid video restoration transformer that combines clip-level parallel processing with global recurrence and guided deformable attention, achieving state-of-the-art super-resolution, deblurring, and denoising performance while maintaining low memory consumption and runtime.

Listen

Restoring high-quality video from degraded, blurry, or noisy footage is critical for modern technologies such as live streaming, surveillance systems, and archival media restoration. Existing restoration methods face a fundamental operational trade-off: parallel transformer-based systems achieve high restoration quality by processing all frames at once but demand excessive memory and computing power, whereas sequential recurrent methods use fewer resources but suffer from information loss, noise buildup, and slower linear processing.

The article demonstrates a hybrid deep learning framework, named the Recurrent Video Restoration Transformer (RVRT), designed to achieve an optimal balance among restoration quality, model size, and operational efficiency. The approach divides a video into small multi-frame clips, processing frames within each clip in parallel while linking consecutive clips through a globally recurrent framework that features a novel guided deformable attention mechanism to align video clips smoothly in a single step.

The authors evaluated the framework across eight standard benchmark datasets covering three major restoration tasks: video super-resolution, deblurring, and denoising. The experimental results show that the proposed system establishes state-of-the-art performance across all tasks. Compared to standard parallel transformer models, the framework reduces parameter size and testing memory usage by over 50% while cutting runtime by at least 25% to up to 85% in deblurring and denoising tasks. Furthermore, it consistently outperforms existing recurrent networks in restoration accuracy, improving peak signal-to-noise ratio by 0.2 to 0.5 decibels on key super-resolution benchmarks and mitigating frame error propagation.

These findings indicate that organizations deploying automated video restoration can achieve top-tier visual clarity without incurring prohibitive infrastructure costs or latency penalties. The architecture significantly lowers memory consumption and processing times, making high-quality video enhancement practical for cost-sensitive and near-real-time production pipelines.

Decision-makers should consider piloting the open-source framework for compute-constrained video workflows such as media upscaling and streaming pipelines. Future technical development should focus on building end-to-end video-level motion estimation to eliminate computational overhead in larger clip configurations, while deployment teams should remain mindful of ethical and privacy risks when clarifying real-world surveillance footage.

Cover for Recurrent Video Restoration Transformer with Guided Deformable Attention

Abstract

Video restoration aims at restoring multiple high-quality frames from multiple low-quality frames. Existing video restoration methods generally fall into two extreme cases, i.e., they either restore all frames in parallel or restore the video frame by frame in a recurrent way, which would result in different merits and drawbacks. Typically, the former has the advantage of temporal information fusion. However, it suffers from large model size and intensive memory consumption; the latter has a relatively small model size as it shares parameters across frames; however, it lacks long-range dependency modeling ability and parallelizability. In this paper, we attempt to integrate the advantages of the two cases by proposing a recurrent video restoration transformer, namely RVRT. RVRT processes local neighboring frames in parallel within a globally recurrent framework which can achieve a good trade-off between model size, effectiveness, and efficiency. Specifically, RVRT divides the video into multiple clips and uses the previously inferred clip feature to estimate the subsequent clip feature. Within each clip, different frame features are jointly updated with implicit feature aggregation. Across different clips, the guided deformable attention is designed for clip-to-clip alignment, which predicts multiple relevant locations from the whole inferred clip and aggregates their features by the attention mechanism. Extensive experiments on video super-resolution, deblurring, and denoising show that the proposed RVRT achieves state-of-the-art performance on benchmark datasets with balanced model size, testing memory and runtime. The codes are available at https://github.com/JingyunLiang/RVRT.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Video Restoration
  • 2.2 Vision Transformer
  • 3 Methodology
  • 3.1 Overall Architecture
  • 3.2 Recurrent Feature Refinement
  • 3.3 Guided Deformable Attention for Video Alignment
  • 4 Experiments
  • 4.1 Experimental Setup
  • 4.2 Ablation Study
  • 4.3 Video Super-Resolution
  • 4.4 Video Deblurring
  • 6 Limitations and Societal Impacts
  • Acknowledgments and Disclosure of Funding
  • References
  • Checklist

Knowls

  1. Knowl 1 — RVRT Framework for Recurrent Video Restoration

    model/method

    Recurrent Video Restoration Transformer (RVRT) integrates recurrent architectures and transformer architectures by dividing a video into fixed-length clips of NN frames and processing frames within each clip in parallel while propagating features across clips recurrently.

    Given a degraded low-quality (LQ) video sequence ILQimesRT×H×W×CI^{LQ} imes \mathbb{R}^{T \times H \times W \times C} (where TT, HH, WW, and CC are video length, height, width, and channels, respectively), RVRT aims to reconstruct the high-quality (HQ) video IHQ∈RT×sH×sW×CI^{HQ} \in \mathbb{R}^{T \times sH \times sW \times C} with scale factor ss. RVRT consists of three stages:

    1. Shallow Feature Extraction: An initial convolution layer extracts shallow features. For deblurring and denoising (s=1s = 1), two strided convolutional layers downsample the features. Several Residual Swin Transformer Blocks (RSTBs) extract the shallow feature representation F0F^0.
    2. Recurrent Feature Refinement (RFR): The video features at layer ii, Fi∈RT×H×W×CF^i \in \mathbb{R}^{T \times H \times W \times C}, are reshaped into T/NT/N clip features F1i,F2i,…,FT/Ni∈RN×H×W×CF^i_1, F^i_2, \dots, F^i_{T/N} \in \mathbb{R}^{N \times H \times W \times C}, each comprising NN neighboring frame features Ft,1i,…,Ft,Ni∈RH×W×CF^i_{t,1}, \dots, F^i_{t,N} \in \mathbb{R}^{H \times W \times C}. The (t−1)(t-1)-th clip feature Ft−1iF^i_{t-1} is aligned to the tt-th clip via Guided Deformable Attention (GDA): F^t−1i=GDA(Ft−1i;Ot−1→ti,Ft−1i−1,Fti−1)\widehat{F}^i_{t-1} = \text{GDA}(F^i_{t-1}; O^i_{t-1 \to t}, F^{i-1}_{t-1}, F^{i-1}_t) where Ot−1→tiO^i_{t-1 \to t} is the optical flow between the clips, and F^t−1i\widehat{F}^i_{t-1} is the aligned clip feature. The current clip feature is updated as: Fti=RFR(Ft0,Ft1,…,Fti−1,F^t−1i)F^i_t = \text{RFR}(F^0_t, F^1_t, \dots, F^{i-1}_t, \widehat{F}^i_{t-1}) where RFR(⋅)\text{RFR}(\cdot) includes a convolutional fusion layer and Modified Residual Swin Transformer Blocks (MRSTBs). To propagate temporal context bidirectionally, the sequence order is reversed for all even-indexed refinement modules across LL stacked refinement stages.
    3. HQ Frame Reconstruction: Several RSTBs process the refined features, and a sub-pixel convolution (pixel shuffle) layer reconstructs the output HQ video IRHQI^{RHQ}.

    The whole network is trained using the Charbonnier loss: L=∥IRHQ−IHQ∥2+ϵ2\mathcal{L} = \sqrt{\|I^{RHQ} - I^{HQ}\|^2 + \epsilon^2} where ϵ=10−3\epsilon = 10^{-3}.

  2. Knowl 2 — Guided Deformable Attention for Clip-to-Clip Alignment

    model/method

    Guided Deformable Attention (GDA) performs one-stage video-to-video alignment between the (t−1)(t-1)-th supporting clip feature Ft−1iF^i_{t-1} and the tt-th reference clip feature FtiF^i_t across NN frames in each clip.

    For every frame pair (n′,n)(n', n) with 1≤n′,n≤N1 \le n', n \le N, where n′n' indexes the supporting frame in clip t−1t-1 and nn indexes the reference frame in clip tt:

    1. Pre-alignment via Optical Flow Warping: The supporting clip features are warped using optical flow Ot−1→ti,(1:N)O^{i,(1:N)}_{t-1 \to t}: Fˉt−1i,(1:N)=W(Ft−1i,Ot−1→ti,(1:N))\bar{F}^{i,(1:N)}_{t-1} = \mathcal{W}(F^i_{t-1}, O^{i,(1:N)}_{t-1 \to t}) where W\mathcal{W} is the bilinear backward warping operator.

    2. Offset Prediction: Optical flow offsets ot−1→ti,(1:N)o^{i,(1:N)}_{t-1 \to t} are estimated by a small convolutional neural network from the channel concatenation of reference features, warped supporting features, and optical flows: ot−1→ti,(1:N)=CNN(Concat(Fti−1,Fˉt−1i,(1:N),Ot−1→ti,(1:N)))o^{i,(1:N)}_{t-1 \to t} = \text{CNN}(\text{Concat}(F^{i-1}_t, \bar{F}^{i,(1:N)}_{t-1}, O^{i,(1:N)}_{t-1 \to t})) For each frame pair, MM sampling offsets are predicted (NMNM offsets total per reference frame).

    3. Projection and Deformable Sampling: Linear projection matrices PQ,PK,PV∈RC×CP_Q, P_K, P_V \in \mathbb{R}^{C \times C} project features prior to spatial sampling. For reference frame nn: Q=Ft,ni−1PQ∈R1×CQ = F^{i-1}_{t,n} P_Q \in \mathbb{R}^{1 \times C} K=Sampling(Ft−1i−1PK,Ot−1→ti,(n)+ot−1→ti,(n))∈RNM×CK = \text{Sampling}(F^{i-1}_{t-1} P_K, O^{i,(n)}_{t-1 \to t} + o^{i,(n)}_{t-1 \to t}) \in \mathbb{R}^{NM \times C} V=Sampling(Ft−1iPV,Ot−1→ti,(n)+ot−1→ti,(n))∈RNM×CV = \text{Sampling}(F^i_{t-1} P_V, O^{i,(n)}_{t-1 \to t} + o^{i,(n)}_{t-1 \to t}) \in \mathbb{R}^{NM \times C} where Sampling(⋅,⋅)\text{Sampling}(\cdot, \cdot) performs bilinear interpolation at non-integer sampling coordinates defined by the sum of optical flow and predicted offsets.

    4. Attention Aggregation and Channel Interaction: The aligned feature F^t−1i,(n)\widehat{F}^{i,(n)}_{t-1} is computed via scaled dot-product attention followed by a residual Multi-Layer Perceptron (MLP) with GELU activation (hidden dimension RCRC, where RR is the expansion ratio): F^t−1i,(n)=SoftMax(QKTC)V\widehat{F}^{i,(n)}_{t-1} = \text{SoftMax}\left(\frac{Q K^T}{\sqrt{C}}\right) V F^t−1i=F^t−1i+MLP(F^t−1i)\widehat{F}^i_{t-1} = \widehat{F}^i_{t-1} + \text{MLP}(\widehat{F}^i_{t-1})

    GDA supports multi-group deformable sampling along channels and multi-head attention within each deformable group.

  3. Knowl 3 — Layer-by-Layer Optical Flow and Offset Refinement in GDA

    equation

    In Guided Deformable Attention (GDA), initial optical flows Ot−1→t1,(1:N)O^{1,(1:N)}_{t-1 \to t} are computed from the low-quality input video using SPyNet. Across stacked recurrent refinement layers i=1,…,L−1i = 1, \dots, L-1, the optical flows are progressively refined by averaging the MM predicted offsets for each frame pair (n′,n)(n', n):

    Ot−1→t,n′i+1,(n)=Ot−1→t,n′i,(n)+1M∑m=1M{ot−1→t,n′i,(n)}mO^{i+1,(n)}_{t-1 \to t, n'} = O^{i,(n)}_{t-1 \to t, n'} + \frac{1}{M} \sum_{m=1}^{M} \{o^{i,(n)}_{t-1 \to t, n'}\}_m

    where:

    • Ot−1→t,n′i,(n)∈RH×W×2O^{i,(n)}_{t-1 \to t, n'} \in \mathbb{R}^{H \times W \times 2} is the optical flow from frame n′n' in clip t−1t-1 to frame nn in clip tt at refinement layer ii.
    • {ot−1→t,n′i,(n)}m∈RH×W×2\{o^{i,(n)}_{t-1 \to t, n'}\}_m \in \mathbb{R}^{H \times W \times 2} is the mm-th predicted offset (1≤m≤M1 \le m \le M) for the frame pair (n′,n)(n', n) at layer ii.
    • MM is the number of candidate offset points predicted per frame.
  4. Knowl 4 — Modified Residual Swin Transformer Blocks (MRSTB) with 3D Attention

    model/method

    To perform joint spatio-temporal feature aggregation within an NN-frame clip, the Recurrent Feature Refinement (RFR) module uses Modified Residual Swin Transformer Blocks (MRSTBs).

    In an MRSTB, the standard 2D spatial self-attention window of size h×wh \times w from Swin Transformer is extended to a 3D spatio-temporal window of size N×h×wN \times h \times w. Consequently, every token in every frame within a clip computes self-attention across both its own spatial neighborhood and all other N−1N-1 frames in the clip simultaneously. This performs implicit intra-clip feature alignment and fusion in parallel without sequential recurrent propagation within the clip.

  5. Knowl 5 — RVRT Experimental and Architecture Hyperparameters

    experimental setup

    The standard RVRT architecture and training configuration are specified as follows:

    • Network Architecture: 1 RSTB (2 Swin Transformer layers) for shallow feature extraction; L=4L = 4 Recurrent Feature Refinement (RFR) modules with clip size N=2N = 2, each containing 2 MRSTBs (2 modified Swin Transformer layers per MRSTB); 1 RSTB (2 layers) for reconstruction. Spatial attention window size is 8×88 \times 8 with 6 attention heads in RSTB and MRSTB. Channel dimension is C=144C = 144 for video super-resolution and C=192C = 192 for video deblurring and denoising.
    • GDA Configuration: 12 deformable groups, 12 attention heads, and M=9M = 9 candidate sampling offsets per frame. The query projection expands channels to 2C2C.
    • Training Details: Optimized using the Adam optimizer with batch size 8 for 600,000 iterations. Learning rate starts at 4×10−44 \times 10^{-4} and decays with a Cosine Annealing schedule. SPyNet is initialized with pretrained weights, frozen for the first 30,000 iterations, and trained with its learning rate reduced by 75%. Training patch size is 256×256256 \times 256 HQ patches. Sequence lengths during training are 30 frames for REDS, 14 frames for Vimeo-90K, and 16 frames for DVD, GoPro, and DAVIS.
  6. Knowl 6 — Quantitative Video Super-Resolution Benchmark Results

    data/table

    RVRT was evaluated for 4×4\times video super-resolution under bicubic (BI) and blur-downsampling (BD) degradations on REDS4, Vimeo-90K-T, Vid4, and UDM10 datasets. Metrics reported are RGB-channel PSNR/SSIM for REDS4 and Y-channel PSNR/SSIM for the others.

    Method REDS4 (BI) Vimeo-90K-T (BI) Vid4 (BI) Vimeo-90K-T (BD) Vid4 (BD)
    Bicubic 26.14/0.7292 31.32/0.8684 23.78/0.6347 31.30/0.8687 21.80/0.5246
    TOFlow 27.98/0.7990 33.08/0.9054 25.89/0.7651 34.62/0.9212 25.85/0.7659
    DUF 28.63/0.8251 - 27.33/0.8319 36.87/0.9447 27.38/0.8329
    EDVR 31.09/0.8800 37.61/0.9489 27.35/0.8264 37.81/0.9523 27.85/0.8503
    BasicVSR 31.42/0.8909 37.18/0.9450 27.24/0.8251 37.53/0.9498 27.96/0.8553
    IconVSR 31.67/0.8948 37.47/0.9476 27.39/0.8279 37.84/0.9524 28.04/0.8570
    VRT 32.19/0.9006 38.20/0.9530 27.93/0.8425 38.72/0.9584 29.42/0.8795
    BasicVSR++ 32.39/0.9069 37.79/0.9500 27.79/0.8400 38.21/0.9550 29.04/0.8753
    RVRT (ours) 32.75/0.9113 38.15/0.9527 27.99/0.8462 38.59/0.9576 29.54/0.8810

    RVRT achieves state-of-the-art PSNR and SSIM on REDS4 and Vid4 for both degradations, outperforming the recurrent BasicVSR++ by 0.36 dB0.36\text{ dB} on REDS4 (BI) and 0.50 dB0.50\text{ dB} on Vid4 (BD), and outperforming the transformer model VRT on REDS4 (+0.56 dB+0.56\text{ dB}) and Vid4 (+0.06 dB+0.06\text{ dB} BI, +0.12 dB+0.12\text{ dB} BD).

  7. Knowl 7 — Model Size, Memory, and Runtime Comparison in Video Super-Resolution

    data/table

    Efficiency and restoration quality were benchmarked for 4×4\times video super-resolution with an input resolution of 320×180320 \times 180 on REDS4:

    Method #Param (M) Memory (MB) Runtime (ms) PSNR (dB)
    BasicVSR++ 7.3 223 77 32.39
    BasicVSR++ + RSTB 9.3 1021 201 32.61
    EDVR 20.6 3535 378 31.09
    VSRT 32.6 27487 328 31.19
    VRT 35.6 2149 243 32.19
    RVRT (ours) 10.8 1056 183 32.75

    Compared to parallel transformer/sliding-window models (EDVR, VSRT, VRT), RVRT reduces parameters by 47%–70%47\%\text{--}70\%, decreases testing memory by over 50%50\%, and lowers runtime by at least 25%25\%, while delivering higher PSNR (32.75 dB32.75\text{ dB}). Replacing CNN blocks in BasicVSR++ with RSTB blocks results in comparable memory (1021 MB1021\text{ MB}) and higher runtime (201 ms201\text{ ms}) than RVRT (183 ms183\text{ ms}) with lower PSNR (32.61 dB32.61\text{ dB} vs 32.75 dB32.75\text{ dB}).

  8. Knowl 8 — Quantitative Video Deblurring and Denoising Performance

    data/table

    RVRT was evaluated on video deblurring benchmarks (DVD and GoPro) and video denoising benchmarks (DAVIS and Set8 across noise standard deviations σ∈{10,20,30,40,50}\sigma \in \{10, 20, 30, 40, 50\} in RGB PSNR/SSIM).

    Video Deblurring Results:

    Dataset / Metric EDVR GSTA FGST ESTRNN VRT RVRT (ours)
    DVD PSNR (dB) 31.82 32.53 33.36 31.07 34.24 34.30
    DVD SSIM 0.9160 0.9468 0.9500 0.9023 0.9651 0.9655
    GoPro PSNR (dB) 31.54 32.10 32.90 31.07 34.81 34.92
    GoPro SSIM 0.9260 0.9600 0.9610 0.9023 0.9724 0.9738

    On 1280×7201280 \times 720 inputs, RVRT deblurring requires 13.6M13.6\text{M} parameters and 0.3 s0.3\text{ s} runtime compared to VRT's 18.3M18.3\text{M} parameters and 2.2 s2.2\text{ s}.

    Video Denoising Results (PSNR in dB):

    Dataset σ\sigma DVDNet FastDVDNet PaCNet VRT RVRT (ours)
    DAVIS 10 38.13 38.71 39.97 40.82 40.57
    DAVIS 20 35.70 35.77 36.82 38.15 38.05
    DAVIS 30 34.08 34.04 34.79 36.52 36.57
    DAVIS 40 32.86 32.82 33.34 35.32 35.47
    DAVIS 50 31.85 31.86 32.20 34.36 34.57
    Set8 10 36.08 36.44 37.06 37.88 37.53
    Set8 20 33.49 33.43 33.94 35.02 34.83
    Set8 30 31.79 31.68 32.05 33.35 33.30
    Set8 40 30.55 30.46 30.70 32.15 32.21
    Set8 50 29.56 29.53 29.66 31.22 31.33

    For video denoising on 1280×7201280 \times 720 inputs, RVRT uses 12.8M12.8\text{M} parameters and 0.2 s0.2\text{ s} runtime compared to VRT (18.4M18.4\text{M} parameters, 1.5 s1.5\text{ s}).

  9. Knowl 9 — Ablation Studies on Clip Length, Alignment Modules, and Temporal Robustness

    empirical result

    Ablations on REDS for video super-resolution demonstrated:

    1. Clip Length (NN): Increasing clip length from N=1N=1 (fully recurrent frame-by-frame) to N=2N=2 increases PSNR from 31.98 dB31.98\text{ dB} to 32.10 dB32.10\text{ dB}. At N=3N=3, PSNR reaches 32.07 dB32.07\text{ dB} due to larger within-clip motion, but reaches 32.21 dB32.21\text{ dB} when all pair optical flows are explicitly computed.
    2. Alignment Strategy: Comparing alignment approaches within the same backbone:
      • Optical flow warping: 28.88 dB28.88\text{ dB}
      • Temporal Mutual Self-Attention (TMSA): 30.45 dB30.45\text{ dB}
      • Deformable Convolution (DCN): 31.93 dB31.93\text{ dB}
      • Frame-to-frame GDA: 32.00 dB32.00\text{ dB}
      • Clip-to-clip GDA: 32.10 dB32.10\text{ dB}
    3. GDA Components: Removing optical flow guidance drops performance by 1.11 dB1.11\text{ dB} (30.99 dB30.99\text{ dB} vs 32.10 dB32.10\text{ dB}). Disabling layer-wise optical flow updates drops PSNR to 31.83 dB31.83\text{ dB}. Removing the channel MLP reduces PSNR to 32.03 dB32.03\text{ dB}.
    4. Temporal Error Propagation (Hack Experiment): Zeroing out all pixels of the 50-th frame in a 100-frame video sequence revealed that N=2N=2 experiences a smaller PSNR drop across all frames than N=1N=1, demonstrating that clip-level recurrent processing reduces noise amplification while propagating long-range context across more neighboring frames.
  10. Knowl 10 — Computational and Application Limitations of RVRT

    limitation

    RVRT exhibits the following limitations:

    1. Quadratic Pre-Alignment Complexity: The optical flow pre-alignment step between consecutive clips scales quadratically with clip length NN, as N×NN \times N frame pairs must be aligned between adjacent NN-frame clips.
    2. Sensitivity to Motion at Larger Clip Lengths: Without dense inter-frame optical flow estimation, performance saturates when increasing clip length beyond N=2N=2 due to large intra-clip displacements and flow inaccuracy.
    3. Societal and Application Risks: Like other high-capacity video restoration systems, hallucinated high-frequency details may produce inaccurate reconstructions if deployed in sensitive applications such as medical diagnostic video or forensics.

Coverage note — None was omitted; all contributed models (RVRT, MRSTB, GDA), equations, architectural configurations, experimental tables (VSR, deblurring, denoising, runtime/memory comparisons), ablation studies, and stated limitations are fully represented.

References

  1. 1.Pablo Arias and Jean-Michel Morel. Video denoising via empirical bayesian estimation of space-time patches. Journal of Mathematical Imaging and Vision, 60(1):70–93, 2018.
  2. 2.Jose Caballero, Christian Ledig, Andrew Aitken, Alejandro Acosta, Johannes Totz, Zehan Wang, and Wenzhe Shi. Real-time video super-resolution with spatio-temporal networks and motion compensation. In IEEE Conference on Computer Vision and Pattern Recognition, pages 4778–4787, 2017.
  3. 3.Jose Caballero, Christian Ledig, Andrew Aitken, Alejandro Acosta, Johannes Totz, Zehan Wang, and Wenzhe Shi. Real-time video super-resolution with spatio-temporal networks and motion compensation. In IEEE Conference on Computer Vision and Pattern Recognition, pages 4778–4787, 2017.
  4. 4.Jiezhang Cao, Yawei Li, Kai Zhang, and Luc Van Gool. Video super-resolution transformer. arXiv preprint arXiv:2106.06847, 2021.
  5. 5.Jiezhang Cao, Jingyun Liang, Kai Zhang, Wenguan Wang, Qin Wang, Yulun Zhang, Hao Tang, and Luc Van Gool. Towards interpretable video super-resolution via alternating optimization. arXiv preprint arXiv:2207.10765, 2022.
  6. 6.Jiezhang Cao, Qin Wang, Jingyun Liang, Yulun Zhang, Kai Zhang, and Luc Van Gool. Practical real video denoising with realistic degradation model. arXiv preprint arXiv:2208.11803, 2022.
  7. 7.Mingdeng Cao, Yanbo Fan, Yong Zhang, Jue Wang, and Yujiu Yang. Vdtr: Video deblurring with transformer. arXiv preprint arXiv:2204.08023, 2022.
  8. 8.Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. End-to-end object detection with transformers. In European Conference on Computer Vision, pages 213–229, 2020.
  9. 9.Kelvin CK Chan, Xintao Wang, Ke Yu, Chao Dong, and Chen Change Loy. Basicvsr: The search for essential components in video super-resolution and beyond. In IEEE Conference on Computer Vision and Pattern Recognition, pages 4947–4956, 2021.
  10. 10.Kelvin CK Chan, Xintao Wang, Ke Yu, Chao Dong, and Chen Change Loy. Understanding deformable alignment in video super-resolution. In AAAI Conference on Artificial Intelligence, pages 973–981, 2021.
  11. 11.Kelvin CK Chan, Shangchen Zhou, Xiangyu Xu, and Chen Change Loy. Basicvsr++: Improving video super-resolution with enhanced propagation and alignment. arXiv preprint arXiv:2104.13371, 2021.
  12. 12.Pierre Charbonnier, Laure Blanc-Feraud, Gilles Aubert, and Michel Barlaud. Two deterministic half-quadratic regularization algorithms for computed imaging. In International Conference on Image Processing, pages 168–172, 1994.
  13. 13.Hanting Chen, Yunhe Wang, Tianyu Guo, Chang Xu, Yiping Deng, Zhenhua Liu, Siwei Ma, Chunjing Xu, Chao Xu, and Wen Gao. Pre-trained image processing transformer. In IEEE Conference on Computer Vision and Pattern Recognition, pages 12299–12310, 2021.
  14. 14.Mengyu Chu, You Xie, Jonas Mayer, Laura Leal-Taixé, and Nils Thuerey. Learning temporal coherence via self-supervision for gan-based video generation. ACM Transactions on Graphics (TOG), 39(4):75–1, 2020.
  15. 15.Jifeng Dai, Haozhi Qi, Yuwen Xiong, Yi Li, Guodong Zhang, Han Hu, and Yichen Wei. Deformable convolutional networks. In IEEE International Conference on Computer Vision, pages 764–773, 2017.
  16. 16.Chao Dong, Chen Change Loy, Kaiming He, and Xiaoou Tang. Learning a deep convolutional network for image super-resolution. In European Conference on Computer Vision, pages 184–199, 2014.
  17. 17.Xiaoyi Dong, Jianmin Bao, Dongdong Chen, Weiming Zhang, Nenghai Yu, Lu Yuan, Dong Chen, and Baining Guo. Cswin transformer: A general vision transformer backbone with cross-shaped windows. arXiv preprint arXiv:2107.00652, 2021.
  18. 18.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929, 2020.
  19. 19.Alexey Dosovitskiy, Philipp Fischer, Eddy Ilg, Philip Hausser, Caner Hazirbas, Vladimir Golkov, Patrick Van Der Smagt, Daniel Cremers, and Thomas Brox. Flownet: Learning optical flow with convolutional networks. In IEEE International Conference on Computer Vision, pages 2758–2766, 2015.
  20. 20.Dario Fuoli, Martin Danelljan, Radu Timofte, and Luc Van Gool. Fast online video super-resolution with deformable attention pyramid. arXiv preprint arXiv:2202.01731, 2022.
  21. 21.Dario Fuoli, Shuhang Gu, and Radu Timofte. Efficient video super-resolution through recurrent latent space propagation. In IEEE International Conference on Computer Vision Workshop, pages 3476–3485, 2019.
  22. 22.Zhicheng Geng, Luming Liang, Tianyu Ding, and Ilya Zharkov. Rstt: Real-time spatial temporal transformer for space-time video super-resolution. arXiv preprint arXiv:2203.14186, 2022.
  23. 23.Muhammad Haris, Gregory Shakhnarovich, and Norimichi Ukita. Recurrent back-projection network for video super-resolution. In IEEE Conference on Computer Vision and Pattern Recognition, pages 3897–3906, 2019.
  24. 24.Yan Huang, Wei Wang, and Liang Wang. Bidirectional recurrent convolutional networks for multi-frame super-resolution. Advances in Neural Information Processing Systems, 28:235–243, 2015.
  25. 25.Yan Huang, Wei Wang, and Liang Wang. Video super-resolution via bidirectional recurrent convolutional networks. IEEE Transactions on Pattern Analysis and Machine Intelligence, 40(4):1015–1028, 2017.
  26. 26.Takashi Isobe, Xu Jia, Shuhang Gu, Songjiang Li, Shengjin Wang, and Qi Tian. Video super-resolution with recurrent structure-detail network. In European Conference on Computer Vision, pages 645–660, 2020.
  27. 27.Takashi Isobe, Songjiang Li, Xu Jia, Shanxin Yuan, Gregory Slabaugh, Chunjing Xu, Ya-Li Li, Shengjin Wang, and Qi Tian. Video super-resolution with temporal group attention. In IEEE Conference on Computer Vision and Pattern Recognition, pages 8008–8017, 2020.
  28. 28.Takashi Isobe, Fang Zhu, Xu Jia, and Shengjin Wang. Revisiting temporal modeling for video super-resolution. arXiv preprint arXiv:2008.05765, 2020.
  29. 29.Younghyun Jo, Seoung Wug Oh, Jaeyeon Kang, and Seon Joo Kim. Deep video super-resolution network using dynamic upsampling filters without explicit motion compensation. In IEEE Conference on Computer Vision and Pattern Recognition, pages 3224–3232, 2018.
  30. 30.Armin Kappeler, Seunghwan Yoo, Qiqin Dai, and Aggelos K Katsaggelos. Video super-resolution with convolutional neural networks. IEEE Transactions on Computational Imaging, 2(2):109–122, 2016.
  31. 31.Anna Khoreva, Anna Rohrbach, and Bernt Schiele. Video object segmentation with language referring expressions. In Asian Conference on Computer Vision, pages 123–141, 2018.
  32. 32.Tae Hyun Kim, Mehdi SM Sajjadi, Michael Hirsch, and Bernhard Scholkopf. Spatio-temporal transformer network for video restoration. In European Conference on Computer Vision, pages 106–122, 2018.
  33. 33.Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014.
  34. 34.Jie Lei, Liwei Wang, Yelong Shen, Dong Yu, Tamara L Berg, and Mohit Bansal. Mart: Memory-augmented recurrent transformer for coherent video paragraph captioning. arXiv preprint arXiv:2005.05402, 2020.
  35. 35.Dongxu Li, Chenchen Xu, Kaihao Zhang, Xin Yu, Yiran Zhong, Wenqi Ren, Hanna Suominen, and Hongdong Li. Arvo: Learning all-range volumetric correspondence for video deblurring. In IEEE Conference on Computer Vision and Pattern Recognition, pages 7721–7731, 2021.
  36. 36.Wenbo Li, Xin Tao, Taian Guo, Lu Qi, Jiangbo Lu, and Jiaya Jia. Mucan: Multi-correspondence aggregation network for video super-resolution. In European Conference on Computer Vision, pages 335–351, 2020.
  37. 37.Yawei Li, Kai Zhang, Jiezhang Cao, Radu Timofte, and Luc Van Gool. Localvit: Bringing locality to vision transformers. arXiv preprint arXiv:2104.05707, 2021.
  38. 38.Jingyun Liang, Jiezhang Cao, Yuchen Fan, Kai Zhang, Rakesh Ranjan, Yawei Li, Radu Timofte, and Luc Van Gool. Vrt: A video restoration transformer. arXiv preprint arXiv:2201.12288, 2022.
  39. 39.Jingyun Liang, Jiezhang Cao, Guolei Sun, Kai Zhang, Luc Van Gool, and Radu Timofte. SwinIR: Image restoration using swin transformer. In IEEE Conference on International Conference on Computer Vision Workshops, 2021.
  40. 40.Jingyun Liang, Andreas Lugmayr, Kai Zhang, Martin Danelljan, Luc Van Gool, and Radu Timofte. Hierarchical conditional flow: A unified framework for image super-resolution and image rescaling. In IEEE Conference on International Conference on Computer Vision, 2021.
  41. 41.Jingyun Liang, Guolei Sun, Kai Zhang, Luc Van Gool, and Radu Timofte. Mutual affine network for spatially variant kernel estimation in blind image super-resolution. In IEEE Conference on International Conference on Computer Vision, 2021.
  42. 42.Jingyun Liang, Kai Zhang, Shuhang Gu, Luc Van Gool, and Radu Timofte. Flow-based kernel prior with application to blind super-resolution. In IEEE Conference on Computer Vision and Pattern Recognition, pages 10601–10610, 2021.
  43. 43.Renjie Liao, Xin Tao, Ruiyu Li, Ziyang Ma, and Jiaya Jia. Video super-resolution via deep draft-ensemble learning. In IEEE International Conference on Computer Vision, pages 531–539, 2015.
  44. 44.Jing Lin, Yuanhao Cai, Xiaowan Hu, Haoqian Wang, Youliang Yan, Xueyi Zou, Henghui Ding, Yulun Zhang, Radu Timofte, and Luc Van Gool. Flow-guided sparse transformer for video deblurring. arXiv preprint arXiv:2201.01893, 2022.
  45. 45.Jiayi Lin, Yan Huang, and Liang Wang. Fdan: Flow-guided deformable alignment network for video super-resolution. arXiv preprint arXiv:2105.05640, 2021.
  46. 46.Ce Liu and Deqing Sun. On bayesian adaptive video super resolution. IEEE Transactions on Pattern Analysis and Machine Intelligence, 36(2):346–360, 2013.
  47. 47.Chengxu Liu, Huan Yang, Jianlong Fu, and Xueming Qian. Learning trajectory-aware transformer for video super-resolution. arXiv preprint arXiv:2204.04216, 2022.
  48. 48.Ding Liu, Zhaowen Wang, Yuchen Fan, Xianming Liu, Zhangyang Wang, Shiyu Chang, and Thomas Huang. Robust video super-resolution with learned temporal dynamics. In IEEE International Conference on Computer Vision, pages 2507–2515, 2017.
  49. 49.Hongying Liu, Zhubo Ruan, Peng Zhao, Chao Dong, Fanhua Shang, Yuanyuan Liu, Linlin Yang, and Radu Timofte. Video super-resolution based on deep learning: a comprehensive survey. Artificial Intelligence Review, pages 1–55, 2022.
  50. 50.Li Liu, Wanli Ouyang, Xiaogang Wang, Paul Fieguth, Jie Chen, Xinwang Liu, and Matti Pietikäinen. Deep learning for generic object detection: A survey. International Journal of Computer Vision, 128(2):261–318, 2020.
  51. 51.Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. arXiv preprint arXiv:2103.14030, 2021.
  52. 52.Ilya Loshchilov and Frank Hutter. Sgdr: Stochastic gradient descent with warm restarts. arXiv preprint arXiv:1608.03983, 2016.
  53. 53.Seungjun Nah, Sungyong Baik, Seokil Hong, Gyeongsik Moon, Sanghyun Son, Radu Timofte, and Kyoung Mu Lee. Ntire 2019 challenge on video deblurring and super-resolution: Dataset and study. In IEEE Conference on Computer Vision and Pattern Recognition Workshops, pages 1996–2005, 2019.
  54. 54.Seungjun Nah, Tae Hyun Kim, and Kyoung Mu Lee. Deep multi-scale convolutional neural network for dynamic scene deblurring. In IEEE Conference on Computer Vision and Pattern Recognition, pages 3883–3891, 2017.
  55. 55.Seungjun Nah, Sanghyun Son, and Kyoung Mu Lee. Recurrent neural networks with intra-frame iterations for video deblurring. In IEEE Conference on Computer Vision and Pattern Recognition, pages 8102–8111, 2019.
  56. 56.Simon Niklaus. A reimplementation of SPyNet using PyTorch. https://github.com/sniklaus/pytorch-spynet, 2018.
  57. 57.Jinshan Pan, Haoran Bai, and Jinhui Tang. Cascaded deep video deblurring using temporal sharpness prior. In IEEE Conference on Computer Vision and Pattern Recognition, pages 3043–3051, 2020.
  58. 58.Anurag Ranjan and Michael J Black. Optical flow estimation using a spatial pyramid network. In IEEE Conference on Computer Vision and Pattern Recognition, pages 4161–4170, 2017.
  59. 59.Mehdi SM Sajjadi, Raviteja Vemulapalli, and Matthew Brown. Frame-recurrent video super-resolution. In IEEE Conference on Computer Vision and Pattern Recognition, pages 6626–6634, 2018.
  60. 60.Dev Yashpal Sheth, Sreyas Mohan, Joshua L Vincent, Ramon Manzorro, Peter A Crozier, Mitesh M Khapra, Eero P Simoncelli, and Carlos Fernandez-Granda. Unsupervised deep video denoising. In IEEE International Conference on Computer Vision, pages 1759–1768, 2021.
  61. 61.Wenzhe Shi, Jose Caballero, Ferenc Huszár, Johannes Totz, Andrew P Aitken, Rob Bishop, Daniel Rueckert, and Zehan Wang. Real-time single image and video super-resolution using an efficient sub-pixel convolutional neural network. In IEEE Conference on Computer Vision and Pattern Recognition, pages 1874–1883, 2016.
  62. 62.Hyeongseok Son, Junyong Lee, Jonghyeop Lee, Sunghyun Cho, and Seungyong Lee. Recurrent video deblurring with blur-invariant motion estimation and pixel volumes. ACM Transactions on Graphics, 40(5):1–18, 2021.
  63. 63.Shuochen Su, Mauricio Delbracio, Jue Wang, Guillermo Sapiro, Wolfgang Heidrich, and Oliver Wang. Deep video deblurring for hand-held cameras. In IEEE Conference on Computer Vision and Pattern Recognition, pages 1279–1288, 2017.
  64. 64.Maitreya Suin and AN Rajagopalan. Gated spatio-temporal attention-guided video deblurring. In IEEE Conference on Computer Vision and Pattern Recognition, pages 7802–7811, 2021.
  65. 65.Deqing Sun, Xiaodong Yang, Ming-Yu Liu, and Jan Kautz. Pwc-net: Cnns for optical flow using pyramid, warping, and cost volume. In IEEE Conference on Computer Vision and Pattern Recognition, pages 8934–8943, 2018.
  66. 66.Guolei Sun, Yun Liu, Thomas Probst, Danda Pani Paudel, Nikola Popovic, and Luc Van Gool. Boosting crowd counting with transformers. arXiv preprint arXiv:2105.10926, 2021.
  67. 67.Lei Sun, Christos Sakaridis, Jingyun Liang, Qi Jiang, Kailun Yang, Peng Sun, Yaozu Ye, Kaiwei Wang, and Luc Van Gool. Mefnet: Multi-scale event fusion network for motion deblurring. arXiv preprint arXiv:2112.00167, 2021.
  68. 68.Xin Tao, Hongyun Gao, Renjie Liao, Jue Wang, and Jiaya Jia. Detail-revealing deep video super-resolution. In IEEE International Conference on Computer Vision, pages 4472–4480, 2017.
  69. 69.Xin Tao, Hongyun Gao, Xiaoyong Shen, Jue Wang, and Jiaya Jia. Scale-recurrent network for deep image deblurring. In IEEE Conference on Computer Vision and Pattern Recognition, pages 8174–8182, 2018.
  70. 70.Matias Tassano, Julie Delon, and Thomas Veit. Dvdnet: A fast network for deep video denoising. In IEEE International Conference on Image Processing, pages 1805–1809, 2019.
  71. 71.Matias Tassano, Julie Delon, and Thomas Veit. Fastdvdnet: Towards real-time deep video denoising without flow estimation. In IEEE Conference on Computer Vision and Pattern Recognition, pages 1354–1363, 2020.
  72. 72.Yapeng Tian, Yulun Zhang, Yun Fu, and Chenliang Xu. Tdan: Temporally-deformable alignment network for video super-resolution. In IEEE Conference on Computer Vision and Pattern Recognition, pages 3360–3369, 2020.
  73. 73.Zhengzhong Tu, Hossein Talebi, Han Zhang, Feng Yang, Peyman Milanfar, Alan Bovik, and Yinxiao Li. Maxim: Multi-axis mlp for image processing. arXiv preprint arXiv:2201.02973, 2022.
  74. 74.Zhengzhong Tu, Hossein Talebi, Han Zhang, Feng Yang, Peyman Milanfar, Alan Bovik, and Yinxiao Li. Maxvit: Multi-axis vision transformer. arXiv preprint arXiv:2204.01697, 2022.
  75. 75.Gregory Vaksman, Michael Elad, and Peyman Milanfar. Patch craft: Video denoising by deep modeling and patch matching. In IEEE International Conference on Computer Vision, pages 1759–1768, 2021.
  76. 76.Ashish Vaswani, Prajit Ramachandran, Aravind Srinivas, Niki Parmar, Blake Hechtman, and Jonathon Shlens. Scaling local self-attention for parameter efficient visual backbones. arXiv preprint arXiv:2103.12731, 2021.
  77. 77.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. arXiv preprint arXiv:1706.03762, 2017.
  78. 78.Ziyu Wan, Bo Zhang, Dongdong Chen, and Jing Liao. Bringing old films back to life. arXiv preprint arXiv:2203.17276, 2022.
  79. 79.Longguang Wang, Yulan Guo, Li Liu, Zaiping Lin, Xinpu Deng, and Wei An. Deep video super-resolution using hr optical flow estimation. IEEE Transactions on Image Processing, 29:4323–4336, 2020.
  80. 80.Xintao Wang, Kelvin CK Chan, Ke Yu, Chao Dong, and Chen Change Loy. Edvr: Video restoration with enhanced deformable convolutional networks. In IEEE Conference on Computer Vision and Pattern Recognition Workshops, pages 1954–1963, 2019.
  81. 81.Zhendong Wang, Xiaodong Cun, Jianmin Bao, and Jianzhuang Liu. Uformer: A general u-shaped transformer for image restoration. arXiv preprint arXiv:2106.03106, 2021.
  82. 82.Zhiwei Wang, Yao Ma, Zitao Liu, and Jiliang Tang. R-transformer: Recurrent neural network enhanced transformer. arXiv preprint arXiv:1907.05572, 2019.
  83. 83.Bichen Wu, Chenfeng Xu, Xiaoliang Dai, Alvin Wan, Peizhao Zhang, Zhicheng Yan, Masayoshi Tomizuka, Joseph Gonzalez, Kurt Keutzer, and Peter Vajda. Visual transformers: Token-based image representation and processing for computer vision. arXiv preprint arXiv:2006.03677, 2020.
  84. 84.Zhuofan Xia, Xuran Pan, Shiji Song, Li Erran Li, and Gao Huang. Vision transformer with deformable attention. arXiv preprint arXiv:2201.00520, 2022.
  85. 85.Xiaoyu Xiang, Yapeng Tian, Yulun Zhang, Yun Fu, Jan P Allebach, and Chenliang Xu. Zooming slow-mo: Fast and accurate one-stage space-time video super-resolution. In IEEE Conference on Computer Vision and Pattern Recognition, pages 3370–3379, 2020.
  86. 86.Xinguang Xiang, Hao Wei, and Jinshan Pan. Deep video deblurring using sharpness features from exemplars. IEEE Transactions on Image Processing, 29:8976–8987, 2020.
  87. 87.Tianfan Xue, Baian Chen, Jiajun Wu, Donglai Wei, and William T Freeman. Video enhancement with task-oriented flow. International Journal of Computer Vision, 127(8):1106–1125, 2019.
  88. 88.Peng Yi, Zhongyuan Wang, Kui Jiang, Junjun Jiang, Tao Lu, Xin Tian, and Jiayi Ma. Omniscient video super-resolution. In IEEE International Conference on Computer Vision, pages 4429–4438, 2021.
  89. 89.Peng Yi, Zhongyuan Wang, Kui Jiang, Junjun Jiang, and Jiayi Ma. Progressive fusion video super-resolution network via exploiting non-local spatio-temporal correlations. In IEEE International Conference on Computer Vision, pages 3106–3115, 2019.
  90. 90.Wulian Yun, Mengshi Qi, Chuanming Wang, Huiyuan Fu, and Huadong Ma. Coarse-to-fine video denoising with dual-stage spatial-channel transformer. arXiv preprint arXiv:2205.00214, 2022.
  91. 91.Syed Waqas Zamir, Aditya Arora, Salman Khan, Munawar Hayat, Fahad Shahbaz Khan, Ming-Hsuan Yang, and Ling Shao. Multi-stage progressive image restoration. In IEEE Conference on Computer Vision and Pattern Recognition, pages 14821–14831, 2021.
  92. 92.Kai Zhang, Yawei Li, Jingyun Liang, Jiezhang Cao, Yulun Zhang, Hao Tang, Radu Timofte, and Luc Van Gool. Practical blind denoising via swin-conv-unet and data synthesis. arXiv preprint arXiv:2203.13278, 2022.
  93. 93.Kai Zhang, Jingyun Liang, Luc Van Gool, and Radu Timofte. Designing a practical degradation model for deep blind image super-resolution. In IEEE Conference on International Conference on Computer Vision, 2021.
  94. 94.Kai Zhang, Wangmeng Zuo, Yunjin Chen, Deyu Meng, and Lei Zhang. Beyond a gaussian denoiser: Residual learning of deep cnn for image denoising. IEEE Transactions on Image Processing, 26(7):3142–3155, 2017.
  95. 95.Kai Zhang, Wangmeng Zuo, and Lei Zhang. Learning a single convolutional super-resolution network for multiple degradations. In IEEE Conference on Computer Vision and Pattern Recognition, pages 3262–3271, 2018.
  96. 96.Yulun Zhang, Kunpeng Li, Kai Li, Lichen Wang, Bineng Zhong, and Yun Fu. Image super-resolution using very deep residual channel attention networks. In European Conference on Computer Vision, pages 286–301, 2018.
  97. 97.Yinjie Zhang, Yuanxing Zhang, Yi Wu, Yu Tao, Kaigui Bian, Pan Zhou, Lingyang Song, and Hu Tuo. Improving quality of experience by adaptive video streaming with super-resolution. In IEEE Conference on Computer Communications, pages 1957–1966, 2020.
  98. 98.Zhihang Zhong, Ye Gao, Yinqiang Zheng, and Bo Zheng. Efficient spatio-temporal recurrent neural network for video deblurring. In European Conference on Computer Vision, pages 191–207, 2020.
  99. 99.Shangchen Zhou, Jiawei Zhang, Jinshan Pan, Haozhe Xie, Wangmeng Zuo, and Jimmy Ren. Spatio-temporal filter adaptive network for video deblurring. In IEEE International Conference on Computer Vision, pages 2482–2491, 2019.
  100. 100.Xizhou Zhu, Han Hu, Stephen Lin, and Jifeng Dai. Deformable convnets v2: More deformable, better results. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 9308–9316, 2019.
  101. 101.Xizhou Zhu, Weijie Su, Lewei Lu, Bin Li, Xiaogang Wang, and Jifeng Dai. Deformable detr: Deformable transformers for end-to-end object detection. arXiv preprint arXiv:2010.04159, 2020.

Citation

MLA
Liang, J., et al. “Recurrent Video Restoration Transformer with Guided Deformable Attention”. Advances in Neural Information Processing Systems, vol. 35, 2022, pp. 378–93, https://proceedings.neurips.cc/paper_files/paper/2022/file/02687e7b22abc64e651be8da74ec610e-Paper-Conference.pdf.
APA
Liang, J., Fan, Y., Xiang, X., Ranjan, R., Ilg, E., Green, S., Cao, J., Zhang, K., Timofte, R., & Gool, L. V. (2022). Recurrent Video Restoration Transformer with Guided Deformable Attention. Advances in Neural Information Processing Systems, 35, 378–393. https://proceedings.neurips.cc/paper_files/paper/2022/file/02687e7b22abc64e651be8da74ec610e-Paper-Conference.pdf
Chicago
Liang, J., Y. Fan, X. Xiang, et al. 2022. “Recurrent Video Restoration Transformer with Guided Deformable Attention”. Advances in Neural Information Processing Systems 35: 378–93. https://proceedings.neurips.cc/paper_files/paper/2022/file/02687e7b22abc64e651be8da74ec610e-Paper-Conference.pdf.
Harvard
Liang, J. et al. (2022) “Recurrent Video Restoration Transformer with Guided Deformable Attention”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 378–393. Available at: https://proceedings.neurips.cc/paper_files/paper/2022/file/02687e7b22abc64e651be8da74ec610e-Paper-Conference.pdf.
Vancouver
1. Liang J, Fan Y, Xiang X, Ranjan R, Ilg E, Green S, Cao J, Zhang K, Timofte R, Gool LV (2022) Recurrent Video Restoration Transformer with Guided Deformable Attention. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 378–393

BibTeX

@inproceedings{liang2022recurrent,
  title = {Recurrent Video Restoration Transformer with Guided Deformable Attention},
  author = {Liang, Jingyun and Fan, Yuchen and Xiang, Xiaoyu and Ranjan, Rakesh and Ilg, Eddy and Green, Simon and Cao, Jiezhang and Zhang, Kai and Timofte, Radu and Gool, Luc V},
  year = {2022},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {35},
  pages = {378-393},
  url = {https://proceedings.neurips.cc/paper_files/paper/2022/file/02687e7b22abc64e651be8da74ec610e-Paper-Conference.pdf}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors