Scale-Recurrent Network for Deep Image Deblurring
Xin TaoHongyun GaoXiaoyong ShenJue WangJiaya Jia
Proposes a scale-recurrent architecture for single-image deblurring that shares network weights across coarse-to-fine scales, achieving state-of-the-art restoration quality on complex motion blur with significantly fewer parameters.
Motion and focal blur caused by camera shake and moving objects pose a major challenge for computer vision systems and digital photography. Traditional restoration methods rely on hand-crafted mathematical assumptions that often fail in real-world scenarios, while recent deep-learning approaches typically require massive models with independent parameters at every image resolution level. These multi-scale deep learning models are computationally expensive, difficult to train, and susceptible to instability.
The article demonstrates a novel deep learning framework, the Scale-recurrent Network (SRN-DeblurNet), designed to restore sharp images efficiently. The primary objective is to show that sharing model weights across different image scales while passing structural information via recurrent modules produces superior restoration quality with significantly fewer parameters and faster training times than existing methods.
The authors designed a coarse-to-fine restoration architecture where the same neural network is applied across three image resolutions. A Convolutional Long Short-Term Memory module acts as a recurrent hidden state at the bottleneck layer to pass blur and structural information from coarser to finer levels. The framework was trained and evaluated on a standard benchmark dataset of 3,214 high-speed camera blur-sharp image pairs, and further tested against both benchmark and real-world blurred photographs against existing state-of-the-art methods.
The evaluation yielded several key findings. First, the proposed model achieved state-of-the-art restoration quality, surpassing the leading previous multi-scale method on the benchmark dataset with a peak signal-to-noise ratio of 30.26 dB compared to 29.08 dB. Second, sharing network parameters across resolution scales reduced the number of trainable parameters by more than 66% compared to cascaded multi-scale alternatives. Third, the framework reduced the training time required to reach comparable image quality by approximately 75% while processing high-definition test images in about 1.87 seconds, outperforming previous approaches in execution speed. Finally, the internal recurrent mechanism and multi-scale structure proved essential, as single-scale models and architectures without recurrent memory performed substantially worse.
These findings indicate that complex multi-scale restoration tasks do not require separate, parameter-heavy neural networks for each resolution level. Reusing shared weights across scales acts as internal data augmentation, preventing model overfitting and drastically reducing computational costs and memory overhead. For engineering teams and decision-makers, this design enables the deployment of high-performing image restoration tools on tighter hardware constraints, shorter training schedules, and lower operational budgets.
Organizations developing computational photography or computer vision applications should adopt weight-sharing recurrent architectures over independent multi-stage networks for coarse-to-fine processing tasks. As next steps, the authors suggest exploring and applying the scale-recurrent strategy to other multi-scale image processing domains, such as super-resolution, video processing, and general image synthesis.
While the method shows robust performance across synthesized benchmarks and real-world examples, a primary operational constraint is that memory consumption scales with image size, meaning very large images remain bound by available graphics processing memory. Nevertheless, given the consistent quantitative gains and visual fidelity demonstrated across standard datasets and real images, confidence in the architecture's core efficiency and quality advantages remains high.
- Paper: Deep Multi-scale Convolutional Neural Network for Dynamic Scene Deblurring, Seungjun Nah et al. (2017). It introduced the foundational multi-scale convolutional architecture and benchmark dataset for end-to-end dynamic scene deblurring that SRN-DeblurNet directly investigates and simplifies.
- Paper: DeblurGAN: Blind Motion Deblurring Using Conditional Adversarial Networks, Orest Kupyn et al. (2017). It pioneered end-to-end learning for blind motion deblurring without kernel estimation, establishing key benchmarking protocols used to evaluate deep deblurring networks.
- Paper: Deep Laplacian Pyramid Networks for Fast and Accurate Super-Resolution, Wei-Sheng Lai et al. (2017). It demonstrates the coarse-to-fine Laplacian pyramid restoration strategy in deep networks that motivates scale-recurrent architectures for progressive image reconstruction.
- Paper: Image Super-Resolution via Deep Recursive Residual Network, Ying Tai et al. (2017). It illustrates how parameter sharing and recursive neural units can efficiently deepen image restoration networks without inflating parameter counts.
- Paper: Image Restoration Using Very Deep Convolutional Encoder-Decoder Networks with Symmetric Skip Connections, Xiao-Jiao Mao et al. (2016). It establishes the deep convolutional encoder-decoder framework with symmetric skip connections that forms the backbone module operating across scales in SRN-DeblurNet.
- Paper: Multi-Stage Progressive Image Restoration, Syed Waqas Zamir et al. (2021). It advances progressive multi-scale image restoration by introducing cross-stage feature fusion and supervised attention modules to overcome scale-recurrence limitations.
- Paper: Simple Baselines for Image Restoration, Liangyu Chen et al. (2022). It challenges the multi-scale recurrent paradigm by demonstrating that a simple, single-stage baseline without nonlinear activation functions can outperform complex deblurring networks.
- Paper: Uformer: A General U-Shaped Transformer for Image Restoration, Zhendong Wang et al. (2021). It extends hierarchical image restoration architectures by replacing convolutional blocks with locally-enhanced window transformer modules for improved long-range context in deblurring.
- Paper: Restormer: Efficient Transformer for High-Resolution Image Restoration, Syed Waqas Zamir et al. (2022). It builds on multi-scale restoration baselines by designing an efficient transformer architecture tailored for high-resolution motion and defocus deblurring.
