Upscale-A-Video: Temporal-Consistent Diffusion Model for Real-World Video Super-Resolution
Shangchen ZhouPeiqing YangJianyi WangYihang LuoChen Change Loy
Proposes a text-guided latent diffusion framework for real-world video super-resolution that integrates local temporal layers into the U-Net and VAE decoder alongside training-free recurrent latent propagation to produce temporally consistent, high-quality video sequences.
Real-world video super-resolution requires upscaling low-quality footage plagued by complex, unknown degradations such as blur, compression artifacts, noise, and downsampling. Traditional methods based on convolutional neural networks struggle to reconstruct realistic fine details and often produce over-smoothed results. While diffusion models have emerged as a powerful tool for generating highly detailed, photo-realistic imagery, adapting them to video restoration introduces significant challenges: the inherent randomness of diffusion denoising causes temporal instability, severe frame-to-frame flickering, and low-level texture inconsistencies across video sequences.
The article introduces and evaluates Upscale-A-Video, a text-guided latent diffusion framework designed specifically for real-world video super-resolution. The primary objective is to adapt pretrained image diffusion priors to video upscaling while ensuring robust temporal consistency both within short video clips and across entire long sequences.
To achieve this without the prohibitive cost of training large diffusion models from scratch, the authors adapt a pretrained image upscaler using a hybrid local-global strategy. At the local level, they freeze the core spatial layers and train lightweight temporal layers—specifically 3D convolutions and temporal attention—within the denoiser, followed by fine-tuning the latent decoder using temporal blocks and spatial feature transform layers to preserve low-level textures and color fidelity. At the global level, they introduce a training-free, optical-flow-guided recurrent propagation module that bidirectionally aligns and aggregates latent representations across video segments during inference. The framework was trained on roughly 335,000 video clips from public datasets and an additional curated set of 37,000 high-definition YouTube videos, then evaluated across six synthetic, real-world, and artificial intelligence-generated benchmarks.
The evaluation yielded several key findings in order of importance. First, Upscale-A-Video established top-tier performance across diverse benchmarks, achieving the highest peak signal-to-noise ratio on all four synthetic datasets (reaching up to 30.79 dB on UDM10) and outperforming existing methods on perceptual quality and non-reference metrics for real-world and AI-generated videos. Second, the framework substantially improved temporal stability, outperforming baseline diffusion models and achieving flow warping error rates that beat or matched specialized convolutional video networks (e.g., reducing warping error on YouHQ40 from 2.398 to 0.737 in ablation testing). Third, fine-tuning the latent decoder and applying the training-free propagation module proved critical, as omitting them led to severe visual flickering and measurable drops in consistency. Fourth, incorporating text prompts and adjustable input noise levels enabled controllable generation, allowing users to intentionally balance faithful low-level restoration against generative fine detail synthesis.
These findings demonstrate that organizations can successfully leverage large pretrained image generative models for video enhancement tasks without retraining foundational networks from scratch, thereby avoiding massive compute expenditures. The framework enables high-fidelity video restoration for challenging media archives, user-generated mobile content, and AI-generated video workflows where conventional enhancers fail. Furthermore, the ability to steer details via text prompts and tune restoration-generation trade-offs provides operational flexibility depending on whether exact reconstruction fidelity or heightened aesthetic quality is desired.
For practical implementation, organizations upgrading video enhancement pipelines should consider adopting this local-global latent propagation framework. Operators should tune noise parameters to the specific degradation severity of their input material, using lower noise levels for faithful restoration and higher levels when aggressive detail synthesis is needed. While the results are backed by strong empirical metrics across multiple standard benchmarks, potential limitations remain around computational memory constraints during inference and the reliance on optical flow accuracy in scenes with extreme motion or heavy occlusions. Further testing in production environments with complex camera motions is advised before full-scale deployment.
- Paper: High-Resolution Image Synthesis with Latent Diffusion Models, Robin Rombach et al. (2022). It introduces Latent Diffusion Models (LDMs), the foundational framework that Upscale-A-Video directly adopts and adapts for video super-resolution.
- Paper: Align Your Latents: High-Resolution Video Synthesis with Latent Diffusion Models, Andreas Blattmann et al. (2023). It pioneers the paradigm of inserting temporal attention layers into frozen pretrained image latent diffusion models to enable video generation.
- Paper: BasicVSR++: Improving Video Super-Resolution with Enhanced Propagation and Alignment, Kelvin C. K. Chan et al. (2022). It establishes bidirectional optical-flow-guided recurrent feature propagation for video super-resolution, inspiring the global propagation strategy in Upscale-A-Video.
- Paper: Video Diffusion Models, Jonathan Ho et al. (2022). It provides foundational principles for extending diffusion models to the temporal video domain using 3D convolutions and joint image-video conditioning.
- Paper: Real-ESRGAN: Training Real-World Blind Super-Resolution with Pure Synthetic Data, Xintao Wang et al. (2021). It formalizes high-order degradation modeling and blind restoration pipelines essential for training real-world super-resolution systems.
- Paper: EDVR: Video Restoration With Enhanced Deformable Convolutional Networks, Xintao Wang et al. (2019). It defines standard benchmarks and spatio-temporal alignment techniques for deep learning-based video restoration.
- Paper: Denoising Diffusion Probabilistic Models, Jonathan Ho et al. (2020). It introduces the core denoising diffusion probabilistic modeling formulation upon which modern diffusion upscaling architectures rely.
- Paper: StreamingT2V: Consistent, Dynamic, and Extendable Long Video Generation from Text, Roberto Henschel et al. (2025). It builds upon long-range video consistency mechanisms to generate and upscale extended video sequences autoregressively.
- Paper: History-Guided Video Diffusion, Kiwhan Song et al. (2025). It generalizes temporal conditioning in video diffusion models by introducing flexible history guidance and per-frame noise schedules across extended rollouts.
- Paper: RAVE: Randomized Noise Shuffling for Fast and Consistent Video Editing with Diffusion Models, Ozgur Kara et al. (2024). It explores alternative training-free temporal consistency methods for diffusion models using randomized frame shuffling across video grids.
- Paper: Lumiere: A Space-Time Diffusion Model for Video Generation, Omer Bar-Tal et al. (2024). It advances space-time architecture design in video diffusion models by employing a unified Space-Time U-Net for coherent video generation and upsampling.
