Deeply-Recursive Convolutional Network for Image Super-Resolution
Jiwon KimJung Kwon LeeKyoung Mu Lee
Proposes a deeply-recursive convolutional network for image super-resolution that scales model depth up to 16 recursions without adding new parameters, overcoming gradient instability through recursive supervision and skip connections.
Image super-resolution aims to recover high-resolution details from low-resolution inputs, an ill-posed problem where larger image context helps infer missing information but often requires models with excessive parameters that risk overfitting or become impractical to store and run. This matters now because applications in imaging, video, and restoration demand efficient methods that scale context without added complexity.
The article set out to evaluate whether a convolutional network could exploit very large receptive fields for super-resolution by reusing the same weights recursively, while remaining trainable and compact.
The authors built a basic model with embedding, recursive inference, and reconstruction stages, then added recursive supervision of all intermediate outputs plus skip connections from input to reconstruction. They trained on 91 images using standard gradient methods with these extensions, tested on Set5, Set14, B100, and Urban100 for 2x, 3x, and 4x scaling, and compared against prior methods such as SRCNN and A+.
Performance improved steadily with recursion depth up to 16, and the ensemble of intermediate predictions boosted results further. The final model achieved the highest PSNR and SSIM on every dataset and scale, for example raising Set5 2x PSNR from 36.66 dB to 37.63 dB and producing visibly sharper edges and textures where earlier outputs remained blurred.
These gains show that recursion plus targeted supervision and skips can deliver larger context and better accuracy without increasing parameter count, lowering storage needs and training data requirements while raising output quality for downstream tasks.
Further work should test deeper recursion for full-image context and adapt the approach to related problems such as denoising or artifact removal; pilot studies on larger or domain-specific data would clarify robustness before wider deployment.
The main limitations are the modest training set size, long training time of roughly six days on one GPU, and evaluation restricted to standard benchmarks, so results should be validated on additional real-world imagery before high-stakes use.
- Paper: Learning a Deep Convolutional Network for Image Super-Resolution, Chao Dong et al. (2014). This paper establishes the foundational SRCNN framework for single-image super-resolution that the source directly compares against and extends to very deep, recursive formulations.
- Paper: Image Super-Resolution Using Deep Convolutional Networks, Chao Dong et al. (2014). This work introduces end-to-end convolutional learning for image super-resolution, providing the foundational baseline and problem formulation built upon by the source.
- Paper: Deeply-Supervised Nets, Chen-Yu Lee et al. (2014). This paper introduces deeply-supervised training for intermediate convolutional layers, which directly underpins the recursive intermediate supervision strategy implemented in the source network.
- Paper: Low-Complexity Single-Image Super-Resolution based on Nonnegative Neighbor Embedding, M. Bevilacqua et al. (2012). This study introduces standard single-image super-resolution benchmark datasets and metrics used directly to evaluate the performance of the source method.
- Paper: Single image super-resolution from transformed self-exemplars, Jia-Bin Huang et al. (2015). This paper provides the Urban100 benchmark dataset and classical self-similarity baselines against which the source network is extensively evaluated.
- Paper: Delving Deep into Rectifiers: Surpassing Human-Level Performance on ImageNet Classification, Kaiming He et al. (2015). This paper presents robust weight initialization and parametric rectifier techniques that are essential for successfully stabilizing and training deep convolutional networks such as the source architecture.
- Paper: Image Super-Resolution via Deep Recursive Residual Network, Ying Tai et al. (2017). This paper directly extends the source's recursive super-resolution concept by integrating both local and global residual learning into deep recursive blocks to reach 52 layers without parameter explosion.
- Paper: Accurate Image Super-Resolution Using Very Deep Convolutional Networks, Jiwon Kim et al. (2016). This companion study explores deep residual learning and gradient clipping for super-resolution, complementing the recursive architecture introduced in the source.
- Paper: Deep Laplacian Pyramid Networks for Fast and Accurate Super-Resolution, Wei-Sheng Lai et al. (2017). This work builds on deep and progressive super-resolution architectures by proposing a Laplacian pyramid network with deep supervision and progressive upsampling.
- Paper: Enhanced Deep Residual Networks for Single Image Super-Resolution, Bee Lim et al. (2017). This paper pushes deeper into the super-resolution scaling paradigm by optimizing deep residual architectures, removing batch normalization, and supporting multi-scale reconstruction.
- Paper: Residual Dense Network for Image Super-Resolution, Yulun Zhang et al. (2018). This paper advances deep super-resolution by combining residual and dense connections to fully exploit hierarchical features across all layers.
- Paper: Image Super-Resolution Using Very Deep Residual Channel Attention Networks, Yulun Zhang et al. (2018). This work extends very deep super-resolution architectures by incorporating residual-in-residual structures and channel attention mechanisms to selectively enhance high-frequency detail.
- Paper: Photo-Realistic Single Image Super-Resolution Using a Generative Adversarial Network, Christian Ledig et al. (2017). This study departs from purely PSNR-oriented deep super-resolution by introducing adversarial and perceptual loss formulations for photo-realistic reconstruction.
- Paper: SwinIR: Image Restoration Using Swin Transformer, Jingyun Liang et al. (2021). This research modernizes deep image restoration by adapting Swin Transformers with residual convolutional connections to replace traditional CNN pipelines.
