Built independently by an author, for readers. Read the story and support ChapterPal

keyword

space-time video super-resolution

Space-time video super-resolution is a computer vision task that aims to reconstruct high-resolution, high-frame-rate video sequences from low-resolution, low-frame-rate inputs. It combines the objectives of spatial super-resolution, which enhances the pixel clarity and detail of individual frames, and temporal super-resolution or frame interpolation, which synthesizes intermediate frames to increase motion smoothness. Rather than executing spatial enhancement and frame interpolation as separate, sequential stages that can accumulate visual artifacts and alignment errors, space-time video super-resolution models jointly capture spatial textures and temporal motion dynamics to produce temporally coherent and visually sharp video content.

2 items

Video Frame Interpolation Transformer

Video Frame Interpolation Transformer

Zhihao Shi, Xiangyu Xu, Xiaohong Liu, Jun Chen, Ming-Hsuan Yang

Why you should read this

Proposes a lightweight video frame interpolation transformer that adapts local self-attention into a space-time separable mechanism to capture long-range spatial-temporal dependencies with high memory efficiency.

Existing methods for video interpolation heavily rely on deep convolution neural networks, and thus suffer from their intrinsic limitations, such as content-agnostic kernel weights and restricted receptive field. To address these issues, we propose a Transformer-based video interpolation framework that allows content-aware aggregation weights and considers long-range dependencies with the self-attention operations. To avoid the high computational cost of global self-attention, we introduce the concept of local attention into video interpolation and extend it to the spatial-temporal domain. Furthermore, we propose a space-time separation strategy to save memory usage, which also improves performance. In addition, we develop a multi-scale frame synthesis scheme to fully realize the potential of Transformers. Extensive experiments demonstrate the proposed model performs favorably against the state-of-the-art methods both quantitatively and qualitatively on a variety of benchmark datasets. The code and models are released at https://github.com/zhshi0816/Video-Frame-Interpolation-Transformer.

Added

2026-10-05

Recurrent Video Restoration Transformer with Guided Deformable Attention

Recurrent Video Restoration Transformer with Guided Deformable Attention

Jingyun Liang, Yuchen Fan, Xiaoyu Xiang, Rakesh Ranjan, Eddy Ilg, Simon Green, Jiezhang Cao, Kai Zhang, Radu Timofte, Luc Van Gool

OrganizationsETH ZurichMetaUniversity of Würzburg

Why you should read this

Proposes a hybrid video restoration transformer that combines clip-level parallel processing with global recurrence and guided deformable attention, achieving state-of-the-art super-resolution, deblurring, and denoising performance while maintaining low memory consumption and runtime.

Video restoration aims at restoring multiple high-quality frames from multiple low-quality frames. Existing video restoration methods generally fall into two extreme cases, i.e., they either restore all frames in parallel or restore the video frame by frame in a recurrent way, which would result in different merits and drawbacks. Typically, the former has the advantage of temporal information fusion. However, it suffers from large model size and intensive memory consumption; the latter has a relatively small model size as it shares parameters across frames; however, it lacks long-range dependency modeling ability and parallelizability. In this paper, we attempt to integrate the advantages of the two cases by proposing a recurrent video restoration transformer, namely RVRT. RVRT processes local neighboring frames in parallel within a globally recurrent framework which can achieve a good trade-off between model size, effectiveness, and efficiency. Specifically, RVRT divides the video into multiple clips and uses the previously inferred clip feature to estimate the subsequent clip feature. Within each clip, different frame features are jointly updated with implicit feature aggregation. Across different clips, the guided deformable attention is designed for clip-to-clip alignment, which predicts multiple relevant locations from the whole inferred clip and aggregates their features by the attention mechanism. Extensive experiments on video super-resolution, deblurring, and denoising show that the proposed RVRT achieves state-of-the-art performance on benchmark datasets with balanced model size, testing memory and runtime. The codes are available at https://github.com/JingyunLiang/RVRT.

Added

2026-09-26