Built independently by an author, for readers. Read the story and support ChapterPal

keyword

video enhancement

Video enhancement is the process of improving the visual quality, clarity, and resolution of digital video sequences. It encompasses a range of computational techniques designed to restore degraded footage or augment lower-quality recordings, including video super-resolution, noise reduction, compression artifact removal, and frame rate interpolation. Unlike single-image processing, video enhancement relies heavily on temporal information, utilizing motion estimation and alignment across consecutive frames to reconstruct fine spatial details while maintaining temporal consistency to prevent visual flickering, jitter, and artifacts across the sequence.

2 items

Upscale-A-Video: Temporal-Consistent Diffusion Model for Real-World Video Super-Resolution

Upscale-A-Video: Temporal-Consistent Diffusion Model for Real-World Video Super-Resolution

Shangchen Zhou, Peiqing Yang, Jianyi Wang, Yihang Luo, Chen Change Loy

OrganizationsNanyang Technological University

Why you should read this

Proposes a text-guided latent diffusion framework for real-world video super-resolution that integrates local temporal layers into the U-Net and VAE decoder alongside training-free recurrent latent propagation to produce temporally consistent, high-quality video sequences.

Text-based diffusion models have exhibited remarkable success in generation and editing, showing great promise for enhancing visual content with their generative prior. However, applying these models to video super-resolution remains challenging due to the high demands for output fidelity and temporal consistency, which is complicated by the inherent randomness in diffusion models. Our study introduces Upscale-A-Video, a text-guided latent diffusion framework for video upscaling. This framework ensures temporal coherence through two key mechanisms: locally, it integrates temporal layers into U-Net and VAE-Decoder, maintaining consistency within short sequences; globally, without training, a flow-guided recurrent latent propagation module is introduced to enhance overall video stability by propagating and fusing latent across the entire sequences. Thanks to the diffusion paradigm, our model also offers greater flexibility by allowing text prompts to guide texture creation and adjustable noise levels to balance restoration and generation, enabling a trade-off between fidelity and

Added

2026-09-26