Upscale-A-Video: Temporal-Consistent Diffusion Model for Real-World Video Super-Resolution

Shangchen ZhouPeiqing YangJianyi WangYihang LuoChen Change Loy

article2024CVPR137 citations

Proposes a text-guided latent diffusion framework for real-world video super-resolution that integrates local temporal layers into the U-Net and VAE decoder alongside training-free recurrent latent propagation to produce temporally consistent, high-quality video sequences.

Listen

Real-world video super-resolution requires upscaling low-quality footage plagued by complex, unknown degradations such as blur, compression artifacts, noise, and downsampling. Traditional methods based on convolutional neural networks struggle to reconstruct realistic fine details and often produce over-smoothed results. While diffusion models have emerged as a powerful tool for generating highly detailed, photo-realistic imagery, adapting them to video restoration introduces significant challenges: the inherent randomness of diffusion denoising causes temporal instability, severe frame-to-frame flickering, and low-level texture inconsistencies across video sequences.

The article introduces and evaluates Upscale-A-Video, a text-guided latent diffusion framework designed specifically for real-world video super-resolution. The primary objective is to adapt pretrained image diffusion priors to video upscaling while ensuring robust temporal consistency both within short video clips and across entire long sequences.

To achieve this without the prohibitive cost of training large diffusion models from scratch, the authors adapt a pretrained image upscaler using a hybrid local-global strategy. At the local level, they freeze the core spatial layers and train lightweight temporal layers—specifically 3D convolutions and temporal attention—within the denoiser, followed by fine-tuning the latent decoder using temporal blocks and spatial feature transform layers to preserve low-level textures and color fidelity. At the global level, they introduce a training-free, optical-flow-guided recurrent propagation module that bidirectionally aligns and aggregates latent representations across video segments during inference. The framework was trained on roughly 335,000 video clips from public datasets and an additional curated set of 37,000 high-definition YouTube videos, then evaluated across six synthetic, real-world, and artificial intelligence-generated benchmarks.

The evaluation yielded several key findings in order of importance. First, Upscale-A-Video established top-tier performance across diverse benchmarks, achieving the highest peak signal-to-noise ratio on all four synthetic datasets (reaching up to 30.79 dB on UDM10) and outperforming existing methods on perceptual quality and non-reference metrics for real-world and AI-generated videos. Second, the framework substantially improved temporal stability, outperforming baseline diffusion models and achieving flow warping error rates that beat or matched specialized convolutional video networks (e.g., reducing warping error on YouHQ40 from 2.398 to 0.737 in ablation testing). Third, fine-tuning the latent decoder and applying the training-free propagation module proved critical, as omitting them led to severe visual flickering and measurable drops in consistency. Fourth, incorporating text prompts and adjustable input noise levels enabled controllable generation, allowing users to intentionally balance faithful low-level restoration against generative fine detail synthesis.

These findings demonstrate that organizations can successfully leverage large pretrained image generative models for video enhancement tasks without retraining foundational networks from scratch, thereby avoiding massive compute expenditures. The framework enables high-fidelity video restoration for challenging media archives, user-generated mobile content, and AI-generated video workflows where conventional enhancers fail. Furthermore, the ability to steer details via text prompts and tune restoration-generation trade-offs provides operational flexibility depending on whether exact reconstruction fidelity or heightened aesthetic quality is desired.

For practical implementation, organizations upgrading video enhancement pipelines should consider adopting this local-global latent propagation framework. Operators should tune noise parameters to the specific degradation severity of their input material, using lower noise levels for faithful restoration and higher levels when aggressive detail synthesis is needed. While the results are backed by strong empirical metrics across multiple standard benchmarks, potential limitations remain around computational memory constraints during inference and the reliance on optical flow accuracy in scenes with extreme motion or heavy occlusions. Further testing in production environments with complex camera motions is advised before full-scale deployment.

arXiv: 2312.06640
Cover for Upscale-A-Video: Temporal-Consistent Diffusion Model for Real-World Video Super-Resolution

Abstract

Text-based diffusion models have exhibited remarkable success in generation and editing, showing great promise for enhancing visual content with their generative prior. However, applying these models to video super-resolution remains challenging due to the high demands for output fidelity and temporal consistency, which is complicated by the inherent randomness in diffusion models. Our study introduces Upscale-A-Video, a text-guided latent diffusion framework for video upscaling. This framework ensures temporal coherence through two key mechanisms: locally, it integrates temporal layers into U-Net and VAE-Decoder, maintaining consistency within short sequences; globally, without training, a flow-guided recurrent latent propagation module is introduced to enhance overall video stability by propagating and fusing latent across the entire sequences. Thanks to the diffusion paradigm, our model also offers greater flexibility by allowing text prompts to guide texture creation and adjustable noise levels to balance restoration and generation, enabling a trade-off between fidelity and

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Methodology
  • 3.1. Preliminary: Diffusion Models
  • 3.2. Local Consistency within Video Segments
  • 3.3. Global Consistency cross Video Segments
  • 3.4. Inference with Additional Conditions
  • 4. Experiments
  • 4.1. Datasets and Implementation
  • 4.2. Comparisons
  • 4.3. Ablation Study
  • 5. Conclusion
  • References

Knowls

  1. Knowl 1 — Upscale-A-Video Framework Architecture

    model/method

    Upscale-A-Video is a text-guided latent diffusion model (LDM) framework for real-world video super-resolution (VSR). It adapts a pretrained image ×4\times 4 upscaler to process video inputs and enforce temporal consistency across both short local segments and long video sequences.

    The framework processes video inputs X={xi}i=1NX = \{x_i\}_{i=1}^N through a local-global hierarchy:

    1. Local Intra-Segment Processing: The input video is divided into local segments (clips of length 8). Each segment is processed by a temporal U-Net where 2D convolutions are inflated to 3D convolutions and enhanced with temporal attention and 3D residual blocks, maintaining consistency within the segment during iterative latent denoising.

    2. Global Inter-Segment Consistency: During designated global refinement diffusion steps T∗T^*, a training-free flow-guided recurrent latent propagation module propagates and aggregates latent features across all segments bidirectionally to enforce long-term temporal coherence.

    3. Low-Level Consistency and Color Fidelity: A fine-tuned temporal VAE-Decoder with 3D residual blocks and Spatial Feature Transform (SFT) layers decodes the denoised latents back to high-resolution pixel space while eliminating texture flickering and color shifts.

    4. Conditional Flexibility: Texture synthesis is steered by optional text prompts via cross-attention, and the trade-off between restoration fidelity and generative detail is controlled by adjusting the input noise level τ\tau.

  2. Knowl 2 — Temporal U-Net Adaptation and Training

    model/method

    To adapt a 2D image diffusion denoiser (specifically the Stable Diffusion ×4\times 4 Upscaler) to video sequences, the 2D convolutions in the U-Net denoiser fθf_\theta are inflated into 3D convolutions, and temporal attention layers together with 3D residual blocks (ResBlock3D) are interleaved between the pretrained spatial layers.

    The temporal attention layer computes self-attention along the temporal dimension over all frames within the local video clip, equipped with Rotary Position Embeddings (RoPE) to encode positional time information.

    During training, all pretrained spatial layers are kept frozen, and only the newly inserted temporal layers are optimized using the vv-prediction objective: Ez,x,c,t,ϵ[∥v−fθ(zt,xτ;c,t)∥22]\mathbb{E}_{z, x, c, t, \epsilon} \left[ \left\| v - f_\theta(z_t, x_\tau; c, t) \right\|_2^2 \right] where zz is the latent representation of the high-resolution frame, zt=αtz+σtϵz_t = \alpha_t z + \sigma_t \epsilon is the diffused latent at diffusion step tt with Gaussian noise ϵ∼N(0,I)\epsilon \sim \mathcal{N}(0, I), vt≡αtϵ−σtxv_t \equiv \alpha_t \epsilon - \sigma_t x is the target velocity, xτ=ατx+στϵx_\tau = \alpha_\tau x + \sigma_\tau \epsilon is the low-resolution condition image perturbed with noise level τ\tau, and cc denotes conditioning inputs including text prompts and noise level embeddings.

  3. Knowl 3 — Temporally Fine-Tuned VAE-Decoder with Spatial Feature Transform

    model/method

    Standard 2D image VAE-Decoders introduce severe inter-frame flickering when decoding sequences of latent features independently. In Upscale-A-Video, the VAE-Decoder D\mathcal{D} is adapted for video sequences by inserting temporal 3D residual blocks while keeping its pretrained 2D spatial layers frozen.

    To prevent color shifts caused by the latent diffusion denoising process, a Spatial Feature Transform (SFT) layer is introduced at the first layer of the VAE-Decoder. The SFT layer uses the low-resolution input video frames as conditioning inputs to modulate decoder feature maps, directly injecting low-frequency color and structural priors into the decoding pipeline.

    The added temporal layers and SFT modules in the VAE-Decoder are fine-tuned on latent-video pairs using a composite objective: LVAE=L1+LLPIPS+Ladv\mathcal{L}_{\text{VAE}} = \mathcal{L}_1 + \mathcal{L}_{\text{LPIPS}} + \mathcal{L}_{\text{adv}} where L1\mathcal{L}_1 is the pixel-wise mean absolute error, LLPIPS\mathcal{L}_{\text{LPIPS}} is the learned perceptual image patch similarity loss, and Ladv\mathcal{L}_{\text{adv}} is a temporal adversarial loss employing a temporal PatchGAN discriminator.

  4. Knowl 4 — Training-Free Flow-Guided Recurrent Latent Propagation

    algorithm

    To achieve long-range temporal stability across separate video segments without additional training parameters, a flow-guided recurrent propagation and fusion step is executed directly in latent space during diffusion inference.

    Input: Latent sequence at diffusion step tt, {zti}i=1N\{z_t^i\}_{i=1}^N; Low-resolution video frames {xi}i=1N\{x_i\}_{i=1}^N; Global refinement step set T∗T^*; Consistency threshold δ\delta; Fusion weight β∈[0,1]\beta \in [0, 1]
    Output: Refined predicted initial latents {z^~0i}i=1N\{\tilde{\hat{z}}_0^i\}_{i=1}^N
    Compute optical flow fields fi−1→if_{i-1 \to i} and fi→i−1f_{i \to i-1} between consecutive frames of {xi}i=1N\{x_i\}_{i=1}^N using RAFT at latent resolution
    Predict initial latents z^0i\hat{z}_0^i from ztiz_t^i using denoiser fθf_\theta
    if t∈T∗t \in T^* then
        for i=1i = 1 to NN do
            if i=1i = 1 then
                z^~01=z^01\tilde{\hat{z}}_0^1 = \hat{z}_0^1
            else
                Compute forward-backward error Ei−1→i(p)=∥fi−1→i(p)+fi→i−1(p+fi−1→i(p))∥22E_{i-1 \to i}(p) = \| f_{i-1 \to i}(p) + f_{i \to i-1}(p + f_{i-1 \to i}(p)) \|_2^2
                Compute validity mask M(p)=I(Ei−1→i(p)<δ)M(p) = \mathbb{I}(E_{i-1 \to i}(p) < \delta)
                Warp previous latent: zwarp=W(z^~0i−1,fi→i−1)z_{\text{warp}} = \mathcal{W}(\tilde{\hat{z}}_0^{i-1}, f_{i \to i-1})
                z^~0i=[zwarp⋅β+z^0i⋅(1−β)]⊙M+z^0i⊙(1−M)\tilde{\hat{z}}_0^i = [ z_{\text{warp}} \cdot \beta + \hat{z}_0^i \cdot (1 - \beta) ] \odot M + \hat{z}_0^i \odot (1 - M)
            end if
        end for
        Perform backward propagation analogously from frame NN down to 1
    else
        z^~0i=z^0i\tilde{\hat{z}}_0^i = \hat{z}_0^i for all i∈{1,…,N}i \in \{1, \dots, N\}
    end if
    return {z^~0i}i=1N\{\tilde{\hat{z}}_0^i\}_{i=1}^N

    RAFT estimates optical flow directly at the latent feature resolution, avoiding resizing. Forward and backward passes are applied sequentially to distribute context across the entire sequence.

  5. Knowl 5 — Latent Flow Consistency Error and Fusion Equations

    equation

    Let fi−1→if_{i-1 \to i} and fi→i−1f_{i \to i-1} denote the forward and backward optical flow fields between frame i−1i-1 and frame ii, and let pp represent a spatial coordinate in the latent map. The forward-backward consistency error Ei−1→i(p)E_{i-1 \to i}(p) is defined as: Ei−1→i(p)=∥fi−1→i(p)+fi→i−1(p+fi−1→i(p))∥22E_{i-1 \to i}(p) = \left\| f_{i-1 \to i}(p) + f_{i \to i-1}\left(p + f_{i-1 \to i}(p)\right) \right\|_2^2

    Given an error threshold δ\delta, the occlusion mask MM is given by M(p)=I(Ei−1→i(p)<δ)M(p) = \mathbb{I}(E_{i-1 \to i}(p) < \delta), where I(⋅)\mathbb{I}(\cdot) is the indicator function.

    The recurrent update and aggregation of the predicted clean latent z^0i\hat{z}_0^i into updated latent z^~0i\tilde{\hat{z}}_0^i at frame ii is: z^~0i=[W(z^~0i−1,fi→i−1)⋅β+z^0i⋅(1−β)]⊙M+z^0i⊙(1−M)\tilde{\hat{z}}_0^i = \left[ \mathcal{W}\left(\tilde{\hat{z}}_0^{i-1}, f_{i \to i-1}\right) \cdot \beta + \hat{z}_0^i \cdot (1 - \beta) \right] \odot M + \hat{z}_0^i \odot (1 - M) where W(⋅)\mathcal{W}(\cdot) is the optical flow warping operator using nearest-neighbor sampling, β∈[0,1]\beta \in [0, 1] is the fusion weight (default β=0.5\beta = 0.5), and ⊙\odot represents element-wise multiplication.

  6. Knowl 6 — Text Guidance and Controllable Noise-Level Fidelity Modulation

    model/method

    Upscale-A-Video provides two inference-time control handles for flexible video generation and restoration:

    1. Text-Prompt Steering: Text prompts ctextc_{\text{text}} are passed through cross-attention layers in the U-Net. In combination with Classifier-Free Guidance (CFG), descriptive prompts direct the generation of fine-grained domain-specific textures (e.g., animal fur, foliage, architectural surfaces) that are unrecoverable from degraded low-resolution inputs alone.

    2. Noise-Level Trade-off Control: Random Gaussian noise is injected into the condition image xx via xτ=ατx+στϵx_\tau = \alpha_\tau x + \sigma_\tau \epsilon, where τ\tau parameterizes the noise level along early diffusion steps. Lower noise levels (e.g., τ=0\tau = 0) force the diffusion model to prioritize strict restoration fidelity to the input, while higher noise levels (e.g., τ=150\tau = 150) dilute heavy or unseen input degradations and empower the generative prior to produce sharper details.

  7. Knowl 7 — Training Datasets and Implementation Protocol

    experimental setup

    Upscale-A-Video is trained on two video datasets:

    1. WebVid10M subset: Approximately 335K video-text pairs with video resolutions around 336×596336 \times 596.
    2. YouHQ dataset: A dataset of approximately 37K high-definition (1080×19201080 \times 1920) YouTube video clips spanning diverse scenes (street view, landscape, animal, human face, static objects, nighttime).

    Degradation Model: Training pairs are synthesized using the real-world degradation pipeline of RealBasicVSR (combinations of downsampling, blur, noise, compression, and color jitter).

    Training Protocol:

    • Hardware & Optimizer: 32 NVIDIA A100-80G GPUs, Adam optimizer, base learning rate 1×10−41 \times 10^{-4}, effective batch size 384.
    • Temporal U-Net: Trained with spatial crop size 80×8080 \times 80 and clip length of 8 frames. Training runs for 70K iterations on both WebVid10M and YouHQ, followed by 10K iterations on YouHQ alone using null text prompts to allow unprompted inference.
    • Temporal VAE-Decoder: Fine-tuned on 100K synthetic LQ-HQ video pairs generated from WebVid10M and YouHQ using latent codes produced by the fine-tuned U-Net.
  8. Knowl 8 — Quantitative Video Super-Resolution Performance across Benchmarks

    data/table

    The performance of Upscale-A-Video was benchmarked against leading GAN, CNN, and diffusion-based super-resolution methods across synthetic datasets (SPMCS, UDM10, REDS30, YouHQ40), a real-world dataset (VideoLQ), and an AI-generated video dataset (AIGC30). Metrics include PSNR, SSIM, LPIPS, flow warping error Ewarp∗=Ewarp×10−3E_{\text{warp}}^* = E_{\text{warp}} \times 10^{-3}, and non-reference quality metrics CLIP-IQA, MUSIQ, and DOVER.

    Datasets Metrics Real-ESRGAN SD ×4\times 4 Upscaler ResShift StableSR RealVSR DBVSR RealBasicVSR Ours
    SPMCS PSNR ↑\uparrow 22.89 23.19 23.27 22.71 23.88 24.28 24.51 25.32
    SSIM ↑\uparrow 0.669 0.631 0.667 0.657 0.681 0.726 0.717 0.741
    LPIPS ↓\downarrow 0.238 0.304 0.257 0.231 0.437 0.302 0.198 0.222
    Ewarp∗↓E_{\text{warp}}^* \downarrow 1.364 5.008 4.942 4.815 0.294 1.360 0.559 0.367
    UDM10 PSNR ↑\uparrow 27.13 28.07 27.62 26.45 27.38 29.60 29.11 30.79
    SSIM ↑\uparrow 0.843 0.811 0.827 0.825 0.825 0.880 0.876 0.878
    LPIPS ↓\downarrow 0.190 0.186 0.222 0.181 0.278 0.155 0.172 0.133
    Ewarp∗↓E_{\text{warp}}^* \downarrow 1.462 1.710 2.196 2.797 0.531 1.943 0.602 0.446
    REDS30 PSNR ↑\uparrow 22.40 22.98 23.00 23.72 23.05 24.37 23.91 24.41
    SSIM ↑\uparrow 0.591 0.572 0.580 0.635 0.603 0.633 0.636 0.631
    LPIPS ↓\downarrow 0.303 0.399 0.369 0.352 0.658 0.588 0.249 0.335
    Ewarp∗↓E_{\text{warp}}^* \downarrow 3.658 3.753 4.131 1.645 0.378 9.659 1.557 1.278
    YouHQ40 PSNR ↑\uparrow 24.37 19.71 23.77 24.53 24.19 25.37 24.09 25.83
    SSIM ↑\uparrow 0.710 0.579 0.654 0.711 0.695 0.719 0.689 0.733
    LPIPS ↓\downarrow 0.272 0.442 0.376 0.271 0.484 0.430 0.306 0.268
    Ewarp∗↓E_{\text{warp}}^* \downarrow 1.856 3.399 4.426 1.529 0.485 1.149 1.052 0.737
    VideoLQ CLIP-IQA ↑\uparrow 0.360 0.158 0.430 0.344 0.211 0.274 0.387 0.530
    MUSIQ ↑\uparrow 49.48 26.21 40.95 44.23 24.52 29.15 55.33 57.99
    DOVER ↑\uparrow 7.161 2.884 4.679 6.783 2.531 3.628 7.562 7.811
    AIGC30 CLIP-IQA ↑\uparrow 0.430 0.329 0.569 0.467 0.276 0.290 0.565 0.674
    MUSIQ ↑\uparrow 47.09 35.30 43.32 44.93 24.39 27.22 58.87 57.66
    DOVER ↑\uparrow 9.710 5.646 7.042 9.668 3.285 3.523 10.68 11.67

    Upscale-A-Video achieves the highest PSNR across all four synthetic benchmarks (25.32 dB on SPMCS, 30.79 dB on UDM10, 24.41 dB on REDS30, 25.83 dB on YouHQ40) and the lowest LPIPS on UDM10 (0.133) and YouHQ40 (0.268). It achieves top or second-best temporal warping errors (Ewarp∗E_{\text{warp}}^*), markedly outperforming other diffusion baselines (e.g., StableSR at 1.529 vs. Ours at 0.737 on YouHQ40). On real-world and AIGC benchmarks, it achieves the highest CLIP-IQA and DOVER scores.

  9. Knowl 9 — Ablation of Temporal VAE-Decoder and Latent Propagation

    data/table

    An ablation study evaluated the isolated and joint contributions of the fine-tuned temporal VAE-Decoder (ft-VAE-Dec.) and the training-free recurrent latent propagation module (Latent Prop.) on the YouHQ40 test dataset.

    Exp. ft-VAE-Dec. Latent Prop. PSNR ↑\uparrow SSIM ↑\uparrow Ewarp∗↓E_{\text{warp}}^* \downarrow
    (a) 23.82 0.6385 2.398
    (b) ✓ 25.47 0.7215 1.815
    (c) ✓ 25.75 0.7328 0.842
    (d) ✓ ✓ 25.83 0.7326 0.737
    • The baseline model (Exp. a), which uses only the temporal U-Net and a standard 2D image VAE-Decoder without latent propagation, yields high temporal warping error (Ewarp∗=2.398E_{\text{warp}}^* = 2.398) due to low-level texture flickering.
    • Adding the fine-tuned temporal VAE-Decoder (Exp. c) produces the largest single reduction in warping error (from 2.398 to 0.842) and boosts PSNR by 1.93 dB.
    • Enabling both the fine-tuned VAE-Decoder and latent propagation (Exp. d) achieves the best overall temporal stability (Ewarp∗=0.737E_{\text{warp}}^* = 0.737) and reconstruction quality (PSNR = 25.83 dB).

Coverage note — No substantial contributed material was omitted. All primary algorithmic components, mathematical formulations, network adaptations, datasets, experimental results, and ablation studies have been included.

References

  1. 1.Pika Labs. https://www.pika.art/, 2023. 6
  2. 2.Stable Diffusion x4 Upscaler. https://huggingface. co / stabilityai / stable - diffusion - x4 - upscaler, 2023. 2, 3, 4, 6
  3. 3.Zeroscope V2 XL. https : / / replicate . com / anotherjesse/zeroscope-v2-xl, 2023. 6
  4. 4.Omri Avrahami, Dani Lischinski, and Ohad Fried. Blended diffusion for text-driven editing of natural images. In CVPR, 2022. 3
  5. 5.Max Bain, Arsha Nagrani, Gül Varol, and Andrew Zisserman. Frozen in Time: A joint video and image encoder for end-to-end retrieval. In ICCV, 2021. 6
  6. 6.Andreas Blattmann, Robin Rombach, Huan Ling, Tim Dockhorn, Seung Wook Kim, Sanja Fidler, and Karsten Kreis. Align your Latents: High-resolution video synthesis with latent diffusion models. In CVPR, 2023. 2, 3, 4, 6
  7. 7.Jiezhang Cao, Yawei Li, Kai Zhang, and Luc Van Gool. Video super-resolution transformer. arXiv preprint arXiv:2106.06847, 2021. 2
  8. 8.Kelvin CK Chan, Xintao Wang, Ke Yu, Chao Dong, and Chen Change Loy. BasicVSR: The search for essential components in video super-resolution and beyond. In CVPR, 2021. 2
  9. 9.Kelvin CK Chan, Shangchen Zhou, Xiangyu Xu, and Chen Change Loy. BasicVSR++: Improving video super-resolution with enhanced propagation and alignment. In CVPR, 2022. 2, 5
  10. 10.Kelvin CK Chan, Shangchen Zhou, Xiangyu Xu, and Chen Change Loy. Investigating tradeoffs in real-world video super-resolution. In CVPR, 2022. 2, 3, 5, 6, 7
  11. 11.Haoxin Chen, Menghan Xia, Yingqing He, Yong Zhang, Xiaodong Cun, Shaoshu Yang, Jinbo Xing, Yaofang Liu, Qifeng Chen, Xintao Wang, et al. VideoCrafter1: Open diffusion models for high-quality video generation. arXiv preprint arXiv:2310.19512, 2023. 2, 6
  12. 12.Jooyoung Choi, Sungwon Kim, Yonghyun Jeong, Youngjune Gwon, and Sungroh Yoon. ILVR: Conditioning method for denoising diffusion probabilistic models. In ICCV, 2021. 3
  13. 13.Hyungjin Chung, Byeongsu Sim, Dohoon Ryu, and Jong Chul Ye. Improving diffusion models for inverse problems using manifold constraints. In NeurIPS, 2022. 3
  14. 14.Yuren Cong, Mengmeng Xu, Christian Simon, Shoufa Chen, Jiawei Ren, Yanping Xie, Juan-Manuel Perez-Rua, Bodo Rosenhahn, Tao Xiang, and Sen He. FLATTEN: optical flow-guided attention for consistent text-to-video editing. arXiv preprint arXiv:2310.05922, 2023. 2, 3
  15. 15.Patrick Esser, Johnathan Chiu, Parmida Atighehchian, Jonathan Granskog, and Anastasis Germanidis. Structure and content-guided video synthesis with diffusion models. In ICCV, 2023. 3
  16. 16.Rinon Gal, Moab Arar, Yuval Atzmon, Amit H Bermano, Gal Chechik, and Daniel Cohen-Or. Designing an encoder for fast personalization of text-to-image models. arXiv preprint arXiv:2302.12228, 2023. 3
  17. 17.Songwei Ge, Seungjun Nah, Guilin Liu, Tyler Poon, Andrew Tao, Bryan Catanzaro, David Jacobs, Jia-Bin Huang, Ming-Yu Liu, and Yogesh Balaji. Preserve Your Own Correlation: A noise prior for video diffusion models. In ICCV, 2023. 2, 4, 6
  18. 18.Michal Geyer, Omer Bar-Tal, Shai Bagon, and Tali Dekel. TokenFlow: Consistent diffusion features for consistent video editing. arXiv preprint arXiv:2307.10373, 2023. 2
  19. 19.Michal Geyer, Omer Bar-Tal, Shai Bagon, and Tali Dekel. Tokenflow: Consistent diffusion features for consistent video editing. arXiv preprint arxiv:2307.10373, 2023. 3
  20. 20.Shuyang Gu, Dong Chen, Jianmin Bao, Fang Wen, Bo Zhang, Dongdong Chen, Lu Yuan, and Baining Guo. Vector quantized diffusion model for text-to-image synthesis. In CVPR, 2022. 3
  21. 21.Yuwei Guo, Ceyuan Yang, Anyi Rao, Yaohui Wang, Yu Qiao, Dahua Lin, and Bo Dai. AnimateDiff: Animate your personalized text-to-image diffusion models without specific tuning. arXiv preprint arXiv:2307.04725, 2023. 3
  22. 22.Amir Hertz, Ron Mokady, Jay Tenenbaum, Kfir Aberman, Yael Pritch, and Daniel Cohen-Or. Prompt-to-prompt image editing with cross attention control. arXiv preprint arXiv:2208.01626, 2022. 3
  23. 23.Jonathan Ho and Tim Salimans. Classifier-free diffusion guidance. In NeurIPSW, 2022. 5, 8
  24. 24.Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. In NeurIPS, 2020. 2, 3
  25. 25.Jonathan Ho, William Chan, Chitwan Saharia, Jay Whang, Ruiqi Gao, Alexey Gritsenko, Diederik P Kingma, Ben Poole, Mohammad Norouzi, David J Fleet, et al. Imagen Video: High definition video generation with diffusion models. arXiv preprint arXiv:2210.02303, 2022. 2, 3, 6
  26. 26.Jonathan Ho, Tim Salimans, Alexey Gritsenko, William Chan, Mohammad Norouzi, and David J Fleet. Video diffusion models. NeurIPS, 2022.
  27. 27.Yaosi Hu, Zhenzhong Chen, and Chong Luo. LaMD: Latent motion diffusion for video generation. arXiv preprint arXiv:2304.11603, 2023. 3
  28. 28.Takashi Isobe, Xu Jia, Shuhang Gu, Songjiang Li, Shengjin Wang, and Qi Tian. Video super-resolution with recurrent structure-detail network. In ECCV, 2020. 2
  29. 29.Takashi Isobe, Songjiang Li, Xu Jia, Shanxin Yuan, Gregory Slabaugh, Chunjing Xu, Ya-Li Li, Shengjin Wang, and Qi Tian. Video super-resolution with temporal group attention. In CVPR, 2020.
  30. 30.Takashi Isobe, Fang Zhu, Xu Jia, and Shengjin Wang. Revisiting temporal modeling for video super-resolution. BMVC, 2020.
  31. 31.Younghyun Jo, Seoung Wug Oh, Jaeyeon Kang, and Seon Joo Kim. Deep video super-resolution network using dynamic upsampling filters without explicit motion compensation. In CVPR, 2018. 2
  32. 32.Bahjat Kawar, Michael Elad, Stefano Ermon, and Jiaming Song. Denoising diffusion restoration models. In NeurIPS, 2022. 3
  33. 33.Junjie Ke, Qifei Wang, Yilin Wang, Peyman Milanfar, and Feng Yang. MUSIQ: Multi-scale image quality transformer. In ICCV, 2021. 6
  34. 34.Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization. In ICLR, 2015. 6
  35. 35.Wei-Sheng Lai, Jia-Bin Huang, Oliver Wang, Eli Shechtman, Ersin Yumer, and Ming-Hsuan Yang. Learning blind video temporal consistency. In ECCV, 2018. 6
  36. 36.Jingyun Liang, Jiezhang Cao, Yuchen Fan, Kai Zhang, Rakesh Ranjan, Yawei Li, Radu Timofte, and Luc Van Gool. VRT: A video restoration transformer. arXiv preprint arXiv:2201.12288, 2022. 2
  37. 37.Jingyun Liang, Yuchen Fan, Xiaoyu Xiang, Rakesh Ranjan, Eddy Ilg, Simon Green, Jiezhang Cao, Kai Zhang, Radu Timofte, and Luc Van Gool. Recurrent video restoration transformer with guided deformable attention. In NeurIPS, 2022. 2
  38. 38.Xinqi Lin, Jingwen He, Ziyan Chen, Zhaoyang Lyu, Ben Fei, Bo Dai, Wanli Ouyang, Yu Qiao, and Chao Dong. DiffBIR: Towards blind image restoration with generative diffusion prior. arXiv preprint arXiv:2308.15070, 2023. 2
  39. 39.Ce Liu and Deqing Sun. On bayesian adaptive video super resolution. In IEEE TPAMI, 2013. 2
  40. 40.Haoyu Lu, Guoxing Yang, Nanyi Fei, Yuqi Huo, Zhiwu Lu, Ping Luo, and Mingyu Ding. VDT: An empirical study on video diffusion with transformers. arXiv preprint arXiv:2305.13311, 2023. 3
  41. 41.Zhengxiong Luo, Dayou Chen, Yingya Zhang, Yan Huang, Liang Wang, Yujun Shen, Deli Zhao, Jingren Zhou, and Tieniu Tan. VideoFusion: Decomposed diffusion models for high-quality video generation. In CVPR, 2023. 6
  42. 42.Kangfu Mei and Vishal Patel. VIDM: Video implicit diffusion models. In AAAI, 2023. 3
  43. 43.Simon Meister, Junhwa Hur, and Stefan Roth. UnFlow: Unsupervised learning of optical flow with a bidirectional census loss. In AAAI, 2018. 5
  44. 44.Chong Mou, Xintao Wang, Liangbin Xie, Jian Zhang, Zhongang Qi, Ying Shan, and Xiaohu Qie. T2I-Adapter: Learning adapters to dig out more controllable ability for text-to-image diffusion models. arXiv preprint arXiv:2302.08453, 2023. 3
  45. 45.Seungjun Nah, Sungyong Baik, Seokil Hong, Gyeongsik Moon, Sanghyun Son, Radu Timofte, and Kyoung Mu Lee. NTIRE 2019 challenge on video deblurring and super-resolution: Dataset and study. In CVPRW, 2019. 2, 6, 7
  46. 46.Alexander Quinn Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam, Pamela Mishkin, Bob Mcgrew, Ilya Sutskever, and Mark Chen. GLIDE: Towards photorealistic image generation and editing with text-guided diffusion models. In ICML, 2022. 3
  47. 47.Jinshan Pan, Haoran Bai, Jiangxin Dong, Jiawei Zhang, and Jinhui Tang. Deep blind video super-resolution. In ICCV, 2021. 6
  48. 48.Chenyang Qi, Xiaodong Cun, Yong Zhang, Chenyang Lei, Xintao Wang, Ying Shan, and Qifeng Chen. FateZero: Fusing attentions for zero-shot text-based video editing. ICCV, 2023. 3
  49. 49.Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. Hierarchical text-conditional image generation with clip latents. arXiv preprint arXiv:2204.06125, 2022. 2
  50. 50.Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. High-resolution image synthesis with latent diffusion models. In CVPR, 2022. 2, 3
  51. 51.Hshmat Sahak, Daniel Watson, Chitwan Saharia, and David Fleet. Denoising diffusion probabilistic models for robust image super-resolution in the wild. arXiv preprint arXiv:2302.07864, 2023. 2, 3
  52. 52.Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily L Denton, Kamyar Ghasemipour, Raphael Gontijo Lopes, Burcu Karagol Ayan, Tim Salimans, et al. Photorealistic text-to-image diffusion models with deep language understanding. In NeurIPS, 2022. 2
  53. 53.Chitwan Saharia, Jonathan Ho, William Chan, Tim Salimans, David J Fleet, and Mohammad Norouzi. Image super-resolution via iterative refinement. IEEE TPAMI, 2022. 3
  54. 54.Tim Salimans and Jonathan Ho. Progressive distillation for fast sampling of diffusion models. In ICLR, 2022. 4
  55. 55.Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli. Deep unsupervised learning using nonequilibrium thermodynamics. In ICML, 2015. 3
  56. 56.Jiaming Song, Arash Vahdat, Morteza Mardani, and Jan Kautz. Pseudoinverse-guided diffusion models for inverse problems. In ICLR, 2023. 3
  57. 57.Xin Tao, Hongyun Gao, Renjie Liao, Jue Wang, and Jiaya Jia. Detail-revealing deep video super-resolution. In ICCV, 2017. 6, 7
  58. 58.Zachary Teed and Jia Deng. RAFT: Recurrent all-pairs field transforms for optical flow. In ECCV, 2020. 3, 5
  59. 59.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothee Lacroix, Baptiste Roziere, Naman Goyal, Eric Hambro, Faisal Azhar, et al. LLaMA: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023. 4
  60. 60.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. NeurIPS, 2017. 3
  61. 61.Jianyi Wang, Kelvin CK Chan, and Chen Change Loy. Exploring clip for assessing the look and feel of images. In AAAI, 2023. 6
  62. 62.Jianyi Wang, Zongsheng Yue, Shangchen Zhou, Kelvin CK Chan, and Chen Change Loy. Exploiting diffusion prior for real-world image super-resolution. arXiv preprint arXiv:2305.07015, 2023. 2, 3, 5, 6
  63. 63.Xintao Wang, Ke Yu, Chao Dong, and Chen Change Loy. Recovering realistic texture in image super-resolution by deep spatial feature transform. In CVPR, 2018. 5
  64. 64.Xintao Wang, Kelvin CK Chan, Ke Yu, Chao Dong, and Chen Change Loy. EDVR: Video restoration with enhanced deformable convolutional networks. In CVPRW, 2019. 2
  65. 65.Xintao Wang, Liangbin Xie, Chao Dong, and Ying Shan. Real-ESRGAN: Training real-world blind super-resolution with pure synthetic data. In ICCVW, 2021. 6
  66. 66.Xiang Wang, Hangjie Yuan, Shiwei Zhang, Dayou Chen, Jiuniu Wang, Yingya Zhang, Yujun Shen, Deli Zhao, and Jingren Zhou. VideoComposer: Compositional video synthesis with motion controllability. arXiv preprint arXiv:2306.02018, 2023. 2, 4
  67. 67.Yinhuai Wang, Jiwen Yu, and Jian Zhang. Zero-shot image restoration using denoising diffusion null-space model. ICLR, 2022. 3
  68. 68.Yaohui Wang, Xinyuan Chen, Xin Ma, Shangchen Zhou, Ziqi Huang, Yi Wang, Ceyuan Yang, Yinan He, Jiashuo Yu, Peiqing Yang, et al. LaVie: High-quality video generation with cascaded latent diffusion models. arXiv preprint arXiv:2309.15103, 2023. 4, 6
  69. 69.Yinhuai Wang, Jiwen Yu, and Jian Zhang. Zero-shot image restoration using denoising diffusion null-space model. ICLR, 2023. 3
  70. 70.Haoning Wu, Erli Zhang, Liang Liao, Chaofeng Chen, Jingwen Hou, Annan Wang, Wenxiu Sun, Qiong Yan, and Weisi Lin. Exploring video quality assessment on user generated contents from aesthetic and technical perspectives. In ICCV, 2023. 6
  71. 71.Jay Zhangjie Wu, Yixiao Ge, Xintao Wang, Stan Weixian Lei, Yuchao Gu, Yufei Shi, Wynne Hsu, Ying Shan, Xiaohu Qie, and Mike Zheng Shou. Tune-A-Video: One-shot tuning of image diffusion models for text-to-video generation. In ICCV, 2023. 2, 3, 4
  72. 72.Bin Xia, Yulun Zhang, Shiyin Wang, Yitong Wang, Xinglong Wu, Yapeng Tian, Wenming Yang, and Luc Van Gool. DiffIR: Efficient diffusion model for image restoration. ICCV, 2023. 3
  73. 73.Liangbin Xie, Xintao Wang, Shuwei Shi, Jinjin Gu, Chao Dong, and Ying Shan. Mitigating artifacts in real-world video super-resolution models. In AAAI, 2023. 2, 3
  74. 74.Zhen Xing, Qi Dai, Han Hu, Zuxuan Wu, and Yu-Gang Jiang. SimDA: Simple diffusion adapter for efficient video generation. arXiv preprint arXiv:2308.09710, 2023. 3
  75. 75.Tianfan Xue, Baian Chen, Jiajun Wu, Donglai Wei, and William T Freeman. Video enhancement with task-oriented flow. In ICCV, 2019. 2
  76. 76.Peiqing Yang, Shangchen Zhou, Qingyi Tao, and Chen Change Loy. PGDiff: Guiding diffusion models for versatile face restoration via partial guidance. In NeurIPS, 2023. 3
  77. 77.S. Yang, J. Sohl-Dickstein, D. P. Kingma, A. Kumar, S. Ermon, and B. Poole. Score-based generative modeling through stochastic differential equations. In ICLR, 2021. 3
  78. 78.Shuai Yang, Yifan Zhou, Ziwei Liu, , and Chen Change Loy. Rerender A Video: Zero-shot text-guided video-to-video translation. In SIGGRAPH Asia, 2023. 2, 3, 4
  79. 79.Tao Yang, Peiran Ren, Xuansong Xie, and Lei Zhang. Pixel-aware stable diffusion for realistic image super-resolution and personalized stylization. arXiv preprint arXiv:2308.14469, 2023. 3
  80. 80.Xi Yang, Wangmeng Xiang, Hui Zeng, and Lei Zhang. Real-world video super-resolution: A benchmark dataset and a decomposition based learning scheme. In ICCV, 2021. 2, 3, 6
  81. 81.Peng Yi, Zhongyuan Wang, Kui Jiang, Junjun Jiang, and Jiayi Ma. Progressive fusion video super-resolution network via exploiting non-local spatio-temporal correlations. In ICCV, 2019. 6, 7
  82. 82.Peng Yi, Zhongyuan Wang, Kui Jiang, Junjun Jiang, and Jiayi Ma. Progressive fusion video super-resolution network via exploiting non-local spatio-temporal correlations. In ICCV, 2019. 2
  83. 83.Zongsheng Yue, Jianyi Wang, and Chen Change Loy. ResShift: Efficient diffusion model for image super-resolution by residual shifting. In NeurIPS, 2023. 2, 3, 5, 6
  84. 84.David Junhao Zhang, Jay Zhangjie Wu, Jia-Wei Liu, Rui Zhao, Lingmin Ran, Yuchao Gu, Difei Gao, and Mike Zheng Shou. Show-1: Marrying pixel and latent diffusion models for text-to-video generation. arXiv preprint arXiv:2309.15818, 2023. 6
  85. 85.Lvmin Zhang, Anyi Rao, and Maneesh Agrawala. Adding conditional control to text-to-image diffusion models. In ICCV, 2023. 2, 3
  86. 86.Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shechtman, and Oliver Wang. The unreasonable effectiveness of deep features as a perceptual metric. In CVPR, 2018. 5, 6
  87. 87.Daquan Zhou, Weimin Wang, Hanshu Yan, Weiwei Lv, Yizhe Zhu, and Jiashi Feng. MagicVideo: Efficient video generation with latent diffusion models. arXiv preprint arXiv:2211.11018, 2022. 2
  88. 88.Shangchen Zhou, Jiawei Zhang, Jinshan Pan, Haozhe Xie, Wangmeng Zuo, and Jimmy Ren. Spatio-temporal filter adaptive network for video deblurring. In ICCV, 2019. 5
  89. 89.Shangchen Zhou, Chongyi Li, Kelvin CK Chan, and Chen Change Loy. ProPainter: Improving propagation and transformer for video inpainting. In ICCV, 2023. 5

Citation

MLA
Zhou, S., et al. “Upscale-A-Video: Temporal-Consistent Diffusion Model for Real-World Video Super-Resolution”. arXiv, 2023, http://arxiv.org/abs/2312.06640v1.
APA
Zhou, S., Yang, P., Wang, J., Luo, Y., & Loy, C. C. (2023). Upscale-A-Video: Temporal-Consistent Diffusion Model for Real-World Video Super-Resolution. arXiv. http://arxiv.org/abs/2312.06640v1
Chicago
Zhou, S., P. Yang, J. Wang, Y. Luo, and C. C. Loy. 2023. “Upscale-A-Video: Temporal-Consistent Diffusion Model for Real-World Video Super-Resolution”. arXiv. http://arxiv.org/abs/2312.06640v1.
Harvard
Zhou, S. et al. (2023) “Upscale-A-Video: Temporal-Consistent Diffusion Model for Real-World Video Super-Resolution”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2312.06640v1.
Vancouver
1. Zhou S, Yang P, Wang J, Luo Y, Loy CC (2023) Upscale-A-Video: Temporal-Consistent Diffusion Model for Real-World Video Super-Resolution. arXiv

BibTeX

@article{zhou2023upscale,
  title = {Upscale-A-Video: Temporal-Consistent Diffusion Model for Real-World Video Super-Resolution},
  author = {Zhou, Shangchen and Yang, Peiqing and Wang, Jianyi and Luo, Yihang and Loy, Chen Change},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2312.06640v1},
  eprint = {2312.06640}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE