Deeply-Recursive Convolutional Network for Image Super-Resolution

Jiwon KimJung Kwon LeeKyoung Mu Lee

article2016CVPR2,787 citations

Proposes a deeply-recursive convolutional network for image super-resolution that scales model depth up to 16 recursions without adding new parameters, overcoming gradient instability through recursive supervision and skip connections.

Listen

Image super-resolution aims to recover high-resolution details from low-resolution inputs, an ill-posed problem where larger image context helps infer missing information but often requires models with excessive parameters that risk overfitting or become impractical to store and run. This matters now because applications in imaging, video, and restoration demand efficient methods that scale context without added complexity.

The article set out to evaluate whether a convolutional network could exploit very large receptive fields for super-resolution by reusing the same weights recursively, while remaining trainable and compact.

The authors built a basic model with embedding, recursive inference, and reconstruction stages, then added recursive supervision of all intermediate outputs plus skip connections from input to reconstruction. They trained on 91 images using standard gradient methods with these extensions, tested on Set5, Set14, B100, and Urban100 for 2x, 3x, and 4x scaling, and compared against prior methods such as SRCNN and A+.

Performance improved steadily with recursion depth up to 16, and the ensemble of intermediate predictions boosted results further. The final model achieved the highest PSNR and SSIM on every dataset and scale, for example raising Set5 2x PSNR from 36.66 dB to 37.63 dB and producing visibly sharper edges and textures where earlier outputs remained blurred.

These gains show that recursion plus targeted supervision and skips can deliver larger context and better accuracy without increasing parameter count, lowering storage needs and training data requirements while raising output quality for downstream tasks.

Further work should test deeper recursion for full-image context and adapt the approach to related problems such as denoising or artifact removal; pilot studies on larger or domain-specific data would clarify robustness before wider deployment.

The main limitations are the modest training set size, long training time of roughly six days on one GPU, and evaluation restricted to standard benchmarks, so results should be validated on additional real-world imagery before high-stakes use.

Cover for Deeply-Recursive Convolutional Network for Image Super-Resolution

Abstract

We propose an image super-resolution method (SR) using a deeply-recursive convolutional network (DRCN). Our network has a very deep recursive layer (up to 16 recursions). Increasing recursion depth can improve performance without introducing new parameters for additional convolutions. Albeit advantages, learning a DRCN is very hard with a standard gradient descent method due to exploding/vanishing gradients. To ease the difficulty of training, we propose two extensions: recursive-supervision and skip-connection. Our method outperforms previous methods by a large margin.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Single-Image Super-Resolution
  • 2.2 Recursive Neural Network in Computer Vision
  • 3 Proposed Method
  • 3.1 Basic Model
  • 3.2 Advanced Model
  • 3.3 Training
  • 4 Experimental Results
  • 4.1 Datasets
  • 4.2 Training Setup
  • 4.3 Study of Deep Recursions
  • 4.4 Comparisons with State-of-the-Art Methods
  • 5 Conclusion
  • References

Knowls

  1. Knowl 1 — Deeply-Recursive Convolutional Network (DRCN) Architecture

    model/method

    The Deeply-Recursive Convolutional Network (DRCN) is a deep convolutional architecture for single-image super-resolution that achieves a large effective receptive field without expanding model parameter capacity by repeatedly applying the same convolutional layer. The network takes a low-resolution image pre-upscaled via bicubic interpolation to the target spatial dimension, denoted as xx, and predicts an estimate y^\hat{y} of the ground-truth high-resolution image yy.

    The architecture is decomposed into three functional sub-networks:

    1. Embedding Network (f1f_1): Converts the input image (single-channel luminance or three-channel RGB) into an initial multi-channel feature representation H0H_0.
    2. Inference Network (f2f_2): A recursive sub-network that repeatedly applies a single shared convolutional layer DD times to perform spatial contextual reasoning.
    3. Reconstruction Network (f3f_3): Transforms feature representations back into the target image space.

    When unfolded for D=16D = 16 recursions with 2 convolutional layers in the embedding network, 1 recursive layer repeated 16 times, and 2 convolutional layers in the reconstruction network, the deepest feature path traverses 20 convolutional layers with 3×33 \times 3 filters, providing a receptive field of 41×4141 \times 41 pixels.

  2. Knowl 2 — Embedding and Recursive Inference Sub-Networks

    model/method

    In DRCN, feature extraction and deep recursive processing are structured as follows:

    Embedding Network: The embedding network f1(x)f_1(x) maps the interpolated input image xx to an initial feature map H0H_0 using two convolutional layers with rectified linear unit (ReLU\text{ReLU}) activations:

    H−1=max⁡(0,W−1∗x+b−1)H_{-1} = \max(0, W_{-1} * x + b_{-1})

    H0=max⁡(0,W0∗H−1+b0)H_0 = \max(0, W_0 * H_{-1} + b_0)

    f1(x)=H0f_1(x) = H_0

    where ∗* denotes spatial convolution, W−1,W0W_{-1}, W_0 are weight tensors of size 3×3×F×F3 \times 3 \times F \times F (with F=256F=256), and b−1,b0b_{-1}, b_0 are bias vectors.

    Inference Network: The inference network f2f_2 repeatedly applies a single elementary nonlinear function g(H)g(H) across DD recursion steps using the same shared weight tensor WW and bias vector bb:

    Hd=g(Hd−1)=max⁡(0,W∗Hd−1+b)for d=1,2,…,DH_d = g(H_{d-1}) = \max(0, W * H_{d-1} + b) \quad \text{for } d = 1, 2, \dots, D

    f2(H0)=(g∘g∘⋯∘g⏟D times)(H0)=gD(H0)=HDf_2(H_0) = (\underbrace{g \circ g \circ \dots \circ g}_{D \text{ times}})(H_0) = g^D(H_0) = H_D

    where HdH_d denotes the intermediate hidden feature representation at recursion step dd.

  3. Knowl 3 — Skip-Connection and Reconstruction in DRCN

    model/method

    In single-image super-resolution, the pre-interpolated input image xx and the target high-resolution image yy share substantial low-frequency visual content. To prevent the attenuation of input information across deep recursive layers and avoid using network capacity to learn an identity mapping, DRCN incorporates a global skip-connection from the interpolated input xx directly to the reconstruction sub-network.

    For each recursion step d∈{1,…,D}d \in \{1, \dots, D\}, the intermediate reconstructed output y^d\hat{y}_d is defined as:

    y^d=f3(x,Hd)=x+f3(Hd)\hat{y}_d = f_3(x, H_d) = x + f_3(H_d)

    where f3(Hd)f_3(H_d) is a two-layer convolutional reconstruction network parameterized by shared weights WD+1,WD+2W_{D+1}, W_{D+2} and biases bD+1,bD+2b_{D+1}, b_{D+2} across all recursions:

    Hd,+1=max⁡(0,WD+1∗Hd+bD+1)H_{d,+1} = \max(0, W_{D+1} * H_d + b_{D+1})

    f3(Hd)=max⁡(0,WD+2∗Hd,+1+bD+2)f_3(H_d) = \max(0, W_{D+2} * H_{d,+1} + b_{D+2})

    Adding xx ensures the reconstruction sub-network learns only the residual high-frequency details missing from the interpolated input.

  4. Knowl 4 — Recursive-Supervision and Multi-Recursion Ensemble

    model/method

    To mitigate vanishing and exploding gradients when training deep recursive layers, DRCN introduces recursive-supervision and an ensemble prediction mechanism:

    1. Shared Intermediate Supervision: Rather than supervising only the final recursion state HDH_D, intermediate predictions y^d=f3(x,Hd)\hat{y}_d = f_3(x, H_d) at every recursion level d∈{1,…,D}d \in \{1, \dots, D\} are decoded using the identical, parameter-shared reconstruction network f3f_3 and supervised simultaneously during training. Supervising early recursions provides direct backpropagation paths from the loss layer, alleviating gradient degradation across long recursive chains.

    2. Prediction Ensemble: The final network prediction y^\hat{y} is formed as a learned weighted average of all DD intermediate outputs:

    y^=∑d=1Dwd⋅y^d\hat{y} = \sum_{d=1}^D w_d \cdot \hat{y}_d

    where {wd}d=1D\{w_d\}_{d=1}^D are trainable scalar weighting parameters learned jointly with the network filters.

  5. Knowl 5 — DRCN Multi-Task Training Objective

    equation

    The overall training objective L(θ)\mathcal{L}(\theta) for DRCN balances supervision on intermediate recursion predictions, the ensembled final prediction, and weight decay regularization over a training set of NN image patch pairs {(x(i),y(i))}i=1N\{(x^{(i)}, y^{(i)})\}_{i=1}^N:

    L(θ)=α l1(θ)+(1−α) l2(θ)+β ∥θ∥2\mathcal{L}(\theta) = \alpha \, l_1(\theta) + (1 - \alpha) \, l_2(\theta) + \beta \, \|\theta\|^2

    where θ\theta represents all trainable model parameters, β\beta is the weight decay multiplier, and α∈[0,1]\alpha \in [0, 1] is a companion objective weight that is initialized high for training stability and dynamically decayed during training to prioritize the final ensemble output.

    The intermediate recursion loss l1(θ)l_1(\theta) is:

    l1(θ)=∑d=1D∑i=1N12DN∥y(i)−y^d(i)∥2l_1(\theta) = \sum_{d=1}^D \sum_{i=1}^N \frac{1}{2DN} \|y^{(i)} - \hat{y}_d^{(i)}\|^2

    The ensembled final prediction loss l2(θ)l_2(\theta) is:

    l2(θ)=∑i=1N12N∥y(i)−∑d=1Dwd y^d(i)∥2l_2(\theta) = \sum_{i=1}^N \frac{1}{2N} \left\|y^{(i)} - \sum_{d=1}^D w_d \, \hat{y}_d^{(i)}\right\|^2

    where y^d(i)\hat{y}_d^{(i)} is the reconstructed image from the dd-th recursion for sample ii, and wdw_d is the learned ensemble weight for recursion step dd.

  6. Knowl 6 — Recursive Layer Weight Initialization Scheme

    model/method

    Standard random weight initializations fail to converge in deep recursive convolutional layers (e.g., D=16D=16) because repeated multiplications by the same matrix exacerbate gradient instability.

    DRCN uses a hybrid initialization scheme:

    • Non-recursive layers (embedding and reconstruction sub-networks) are initialized using the robust rectifier initialization of He et al. (zero-mean Gaussian distribution with variance 2/n2/n, where nn is the fan-in).
    • Recursive convolutional layer weights WW are initialized to zero everywhere except for self-connections (the connection from neuron kk to the corresponding neuron kk in the next layer), initializing the recursive convolution as an identity mapping.
    • Biases across all layers are initialized to zero.
  7. Knowl 7 — DRCN Training Setup and Hyperparameters

    experimental setup

    DRCN is trained using the following specifications:

    • Model Configuration: D=16D = 16 recursions; all convolutional layers contain 256 filters of spatial dimensions 3×33 \times 3.
    • Dataset: 91 natural images from Yang et al., decomposed into 41×4141 \times 41 patches with a stride of 21.
    • Optimization: Mini-batch stochastic gradient descent with batch size 64 patches, momentum 0.9, and weight decay multiplier β=0.0001\beta = 0.0001.
    • Learning Rate Schedule: Initial learning rate η=0.01\eta = 0.01, reduced by a factor of 10 whenever the validation error plateaus for 5 consecutive epochs; training is terminated when η<10−6\eta < 10^{-6}.
    • Hardware & Software: Implemented in MatConvNet; training takes approximately 6 days on a single NVIDIA Titan X GPU.
  8. Knowl 8 — Benchmark Super-Resolution Quantitative Performance

    data/table

    DRCN was evaluated on the luminance channel (YY) using Peak Signal-to-Noise Ratio (PSNR, in dB) and Structural Similarity Index (SSIM) across four standard benchmark datasets (Set5, Set14, B100, Urban100) at scaling factors ×2\times 2, imes3 imes 3, and imes4 imes 4. Pixels near image boundaries were cropped to ensure fair comparison with existing methods.

    Dataset Scale Bicubic A+ SRCNN RFL SelfEx DRCN (Ours)
    PSNR/SSIM PSNR/SSIM PSNR/SSIM PSNR/SSIM PSNR/SSIM PSNR/SSIM
    Set5 ×2\times 2 33.66/0.9299 36.54/0.9544 36.66/0.9542 36.54/0.9537 36.49/0.9537 37.63/0.9588
    ×3\times 3 30.39/0.8682 32.58/0.9088 32.75/0.9090 32.43/0.9057 32.58/0.9093 33.82/0.9226
    ×4\times 4 28.42/0.8104 30.28/0.8603 30.48/0.8628 30.14/0.8548 30.31/0.8619 31.53/0.8854
    Set14 ×2\times 2 30.24/0.8688 32.28/0.9056 32.42/0.9063 32.26/0.9040 32.22/0.9034 33.04/0.9118
    ×3\times 3 27.55/0.7742 29.13/0.8188 29.28/0.8209 29.05/0.8164 29.16/0.8196 29.76/0.8311
    ×4\times 4 26.00/0.7027 27.32/0.7491 27.49/0.7503 27.24/0.7451 27.40/0.7518 28.02/0.7670
    B100 ×2\times 2 29.56/0.8431 31.21/0.8863 31.36/0.8879 31.16/0.8840 31.18/0.8855 31.85/0.8942
    ×3\times 3 27.21/0.7385 28.29/0.7835 28.41/0.7863 28.22/0.7806 28.29/0.7840 28.80/0.7963
    ×4\times 4 25.96/0.6675 26.82/0.7087 26.90/0.7101 26.75/0.7054 26.84/0.7106 27.23/0.7233
    Urban100 ×2\times 2 26.88/0.8403 29.20/0.8938 29.50/0.8946 29.11/0.8904 29.54/0.8967 30.75/0.9133
    ×3\times 3 24.46/0.7349 26.03/0.7973 26.24/0.7989 25.86/0.7900 26.44/0.8088 27.15/0.8276
    ×4\times 4 23.14/0.6577 24.32/0.7183 24.52/0.7221 24.19/0.7096 24.79/0.7374 25.14/0.7510

    DRCN outperforms all prior methods across every scale factor and dataset in both PSNR and SSIM. On Set5 at ×4\times 4, DRCN achieves 31.53 dB compared to 30.48 dB for SRCNN (+1.05 dB) and 30.28 dB for A+ (+1.25 dB). On Urban100 at ×2\times 2, DRCN achieves 30.75 dB, exceeding SRCNN (29.50 dB) by 1.25 dB. Running inference takes approximately 1.0 second for a 288×288288 \times 288 image on a Titan X GPU.

  9. Knowl 9 — Effect of Recursion Depth and Intermediate Prediction Ensembling

    empirical result

    Ablation experiments on the DRCN architecture reveal two key properties:

    1. Monotonic Gain with Recursion Depth: When training models with recursion depths D∈{1,6,11,16}D \in \{1, 6, 11, 16\} (which maintain identical parameter counts in the convolutional layers), performance on Set5 (scale ×3\times 3) increases strictly monotonically with depth:
    • D=1D = 1: ≈32.55 dB\approx 32.55\text{ dB} PSNR
    • D=6D = 6: ≈33.15 dB\approx 33.15\text{ dB} PSNR
    • D=11D = 11: ≈33.72 dB\approx 33.72\text{ dB} PSNR
    • D=16D = 16: ≈33.82 dB\approx 33.82\text{ dB} PSNR Expanding the effective receptive field and non-linear depth without adding parameters consistently boosts super-resolution accuracy.
    1. Superiority of Ensemble over Single Intermediate Recursions: Evaluating individual intermediate predictions y^d\hat{y}_d across recursion steps d∈{1,…,16}d \in \{1, \dots, 16\} shows that individual performance peaks at intermediate depths (e.g., d≈5d \approx 5 for ×2\times 2, d≈9d \approx 9 for ×3\times 3, d≈10d \approx 10 for ×4\times 4) and declines for the deepest single steps. However, the learned ensemble y^=∑d=116wdy^d\hat{y} = \sum_{d=1}^{16} w_d \hat{y}_d consistently surpasses every individual intermediate prediction across all scale factors.

Coverage note — Deliberately omitted qualitative visual comparisons from Figures 4, 5, 6, and 7, as their quantitative evaluations are fully captured in the benchmark results table.

References

  1. 1.Y. Bengio, P. Simard, and P. Frasconi. Learning long-term dependencies with gradient descent is difficult. Neural Networks, IEEE Transactions on, 5(2), 1994. 1, 4
  2. 2.M. Bevilacqua, A. Roumy, C. Guillemot, and M.-L. Morel. Super-resolution using neighbor embedding of back-projection residuals. In International Conference on Digital Signal Processing, 2013. 5
  3. 3.C. M. Bishop. Pattern recognition and machine learning. springer, 2006. 5
  4. 4.H. Chang, D.-Y. Yeung, and Y. Xiong. Super-resolution through neighbor embedding. In CVPR, 2004. 2
  5. 5.C. Dong, C. C. Loy, K. He, and X. Tang. Image super-resolution using deep convolutional networks. TPAMI, 2014. 1, 2, 3, 6, 7, 8
  6. 6.D. Eigen, J. Rolfe, R. Fergus, and Y. LeCun. Understanding deep architectures using a recursive convolutional network. In ICLR Workshop, 2014. 2
  7. 7.W. T. Freeman, E. C. Pasztor, and O. T. Carmichael. Learning low-level vision. IJCV, 2000. 2
  8. 8.D. Glasner, S. Bagon, and M. Irani. Super-resolution from a single image. In ICCV, 2009. 2
  9. 9.K. He, X. Zhang, S. Ren, and J. Sun. Delving deep into rectifiers: Surpassing human-level performance on imagenet classification. In ICCV, 2015. 7
  10. 10.J.-B. Huang, A. Singh, and N. Ahuja. Single image super-resolution using transformed self-exemplars. In CVPR, 2015. 2, 6, 7, 8
  11. 11.M. Irani and S. Peleg. Improving resolution by image registration. CVGIP: Graphical models and image processing, 53(3), 1991. 2
  12. 12.K. I. Kim and Y. Kwon. Single-image super-resolution using sparse regression and natural image prior. TPAMI, 2010. 2
  13. 13.A. Krizhevsky, I. Sutskever, and G. E. Hinton. Imagenet classification with deep convolutional neural networks. In NIPS, 2012. 1
  14. 14.Q. V. Le, N. Jaitly, and G. E. Hinton. A simple way to initialize recurrent networks of rectified linear units. arXiv preprint arXiv:1504.00941, 2015. 7
  15. 15.Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner. Gradient-based learning applied to document recognition. Proceedings of the IEEE, 86(11), 1998. 5
  16. 16.C.-Y. Lee, S. Xie, P. Gallagher, Z. Zhang, and Z. Tu. Deeply-supervised nets. arXiv preprint arXiv:1409.5185, 2014. 3, 4
  17. 17.M. Liang and X. Hu. Recurrent convolutional neural network for object recognition. In CVPR, 2015. 2, 4
  18. 18.J. Long, E. Shelhamer, and T. Darrell. Fully convolutional networks for semantic segmentation. arXiv preprint arXiv:1411.4038, 2014. 5
  19. 19.C. G. Marco Bevilacqua, Aline Roumy and M.-L. A. Morel. Low-complexity single-image super-resolution based on nonnegative neighbor embedding. In BMVC, 2012. 2, 5, 7
  20. 20.D. Martin, C. Fowlkes, D. Tal, and J. Malik. A database of human segmented natural images and its application to evaluating segmentation algorithms and measuring ecological statistics. In ICCV, 2001. 7
  21. 21.R. Pascanu, T. Mikolov, and Y. Bengio. On the difficulty of training recurrent neural networks. In ICML, 2013. 4
  22. 22.P. Pinheiro and R. Collobert. Recurrent convolutional neural networks for scene labeling. In Proceedings of The 31st International Conference on Machine Learning, pages 82–90, 2014. 2
  23. 23.S. Schulter, C. Leistner, and H. Bischof. Fast and accurate image upscaling with super-resolution forests. In CVPR, 2015. 2, 6, 7, 8
  24. 24.K. Simonyan and A. Zisserman. Very deep convolutional networks for large-scale image recognition. In ICLR, 2015. 1
  25. 25.R. Socher, B. Huval, B. Bath, C. D. Manning, and A. Y. Ng. Convolutional-recursive deep learning for 3d object classification. In NIPS, 2012. 2
  26. 26.R. Socher, B. Huval, C. D. Manning, and A. Y. Ng. Semantic compositionality through recursive matrix-vector spaces. In EMNLP-CoNLL, 2012. 7
  27. 27.J. Sun, Z. Xu, and H.-Y. Shum. Image super-resolution using gradient profile prior. In CVPR, 2008. 2
  28. 28.R. Timofte, V. De, and L. V. Gool. Anchored neighborhood regression for fast example-based super-resolution. In ICCV, 2013. 2, 5, 7
  29. 29.R. Timofte, V. De Smet, and L. Van Gool. A+: Adjusted anchored neighborhood regression for fast super-resolution. In ACCV, 2014. 2, 5, 6, 7, 8
  30. 30.A. Vedaldi and K. Lenc. Matconvnet – convolutional neural networks for matlab. CoRR, abs/1412.4564, 2014. 5
  31. 31.J. Yang, J. Wright, T. S. Huang, and Y. Ma. Image super-resolution via sparse representation. TIP, 2010. 2, 7
  32. 32.R. Zeyde, M. Elad, and M. Protter. On single image scale-up using sparse-representations. In Curves and Surfaces. Springer, 2012. 2, 7

Citation

MLA
Kim, J., et al. “Deeply-Recursive Convolutional Network for Image Super-Resolution”. 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016, pp. 1637–45, https://doi.org/10.1109/CVPR.2016.181.
APA
Kim, J., Lee, J. K., & Lee, K. M. (2016). Deeply-Recursive Convolutional Network for Image Super-Resolution. 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 1637–1645. https://doi.org/10.1109/CVPR.2016.181
Chicago
Kim, J., J. K. Lee, and K. M. Lee. 2016. “Deeply-Recursive Convolutional Network for Image Super-Resolution”. 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 1637–45. https://doi.org/10.1109/CVPR.2016.181.
Harvard
Kim, J., Lee, J.K. and Lee, K.M. (2016) “Deeply-Recursive Convolutional Network for Image Super-Resolution”, 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp. 1637–1645. Available at: https://doi.org/10.1109/CVPR.2016.181.
Vancouver
1. Kim J, Lee JK, Lee KM (2016) Deeply-Recursive Convolutional Network for Image Super-Resolution. In: 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp 1637–1645

BibTeX

@inproceedings{Kim_2016, title={Deeply-Recursive Convolutional Network for Image Super-Resolution}, url={http://dx.doi.org/10.1109/CVPR.2016.181}, DOI={10.1109/cvpr.2016.181}, booktitle={2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)}, publisher={IEEE}, author={Kim, Jiwon and Lee, Jung Kwon and Lee, Kyoung Mu}, year={2016}, month=June, pages={1637–1645} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF

License: IEEE