A recurrent video restoration transformer is a deep learning architecture designed to reconstruct high-quality video frames from degraded video sequences by combining recurrent temporal modeling with transformer-based attention mechanisms. Unlike conventional approaches that operate purely frame by frame or process an entire video simultaneously, this architecture divides a video into localized clips, processing the frames within each clip in parallel while propagating temporal features across consecutive clips in a recurrent fashion. By using attention mechanisms to align and aggregate features across clips, a recurrent video restoration transformer balances computational efficiency, memory usage, and long-range temporal modeling for video enhancement tasks such as super-resolution, denoising, and deblurring.