Deep Back-Projection Networks for Super-Resolution

Muhammad HarisGreg ShakhnarovichNorimichi Ukita

article2018CVPR1,501 citationsWinner (1st) of NTIRE2018 Competition (Track: x8 Bicubic Downsampling), Winner of PIRM2018 (1st on Region 2, 3rd on Region 1, and 5th on Region 3)

Proposes Deep Back-Projection Networks, an architecture that uses iterative up- and down-sampling stages with an error-feedback mechanism to capture mutual image dependencies and achieve state-of-the-art super-resolution at large scaling factors.

Listen

Generating high-resolution images from low-resolution inputs—known as image super-resolution—is critical for modern computer vision tasks across industries such as security, medical imaging, and media processing. Traditional models often struggle to reconstruct fine details, particularly at large scaling factors, or require excessive computing resources that make real-time deployment difficult.

The article evaluates Deep Back-Projection Networks (DBPN) across several architectural variants to demonstrate how iterative up- and down-sampling with error feedback enhances reconstruction quality while maintaining computational efficiency.

To establish credibility, the authors conducted empirical benchmarking using standard datasets (such as DIV2K, Set5, Set14, BSDS100, Urban100, and Manga109) across 2×, 4×, and 8× enlargement tasks. They compared DBPN variants against leading methods, including LapSRN and EDSR, by measuring visual quality metrics (peak signal-to-noise ratio and structural similarity index) and runtime performance on dedicated graphics hardware.

The analysis reveals several key findings. First, DBPN consistently matches or outperforms competing models across benchmarks, achieving notable image fidelity gains on complex datasets like Urban100 and Manga109. Second, incorporating an error feedback mechanism provides a direct performance boost, improving reconstruction quality on test datasets by 0.26 to 0.53 dB over networks without feedback. Third, the models demonstrate superior computational efficiency; for 8× enlargement, the flagship D-DBPN model processes images in roughly 0.32 seconds compared to about 1.15 seconds for EDSR, while smaller DBPN variants operate in tens of milliseconds. Finally, processing full color RGB channels directly simplifies implementation without sacrificing image quality compared to traditional single-channel luminance approaches.

These results indicate that DBPN variants offer a practical solution for operational environments requiring high-speed, high-fidelity image enlargement. Organizations can deploy smaller variants (such as DBPN-SS or DBPN-S) for low-latency, edge-computing needs or larger variants (D-DBPN) when maximum image sharpness is required at significantly lower computing costs than existing heavy architectures.

For practical implementation, technical teams should select the model variant aligned with their specific hardware and throughput constraints, prioritizing direct RGB processing to streamline pipelines. Before large-scale deployment, organizations should conduct pilot testing on domain-specific imagery, as the evaluations rely primarily on standardized benchmarks and occasionally exhibit minor visual artifacts in challenging pattern reconstructions.

Cover for Deep Back-Projection Networks for Super-Resolution

Abstract

The feed-forward architectures of recently proposed deep super-resolution networks learn representations of low-resolution inputs, and the non-linear mapping from those to high-resolution output. However, this approach does not fully address the mutual dependencies of low- and high-resolution images. We propose Deep Back-Projection Networks (DBPN), that exploit iterative up- and down-sampling layers, providing an error feedback mechanism for projection errors at each stage. We construct mutually-connected up- and down-sampling stages each of which represents different types of image degradation and high-resolution components. We show that extending this idea to allow concatenation of features across up- and down-sampling stages (Dense DBPN) allows us to reconstruct further improve super-resolution, yielding superior results and in particular establishing new state of the art results for large scaling factors such as 8x across multiple data sets.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Image super-resolution using deep networks
  • 2.2 Feedback networks
  • 2.3 Adversarial training
  • 2.4 Back-projection
  • 3 Deep Back-Projection Networks
  • 3.1 Projection units
  • 3.2 Dense projection units
  • 3.3 Network architecture
  • 4 Experimental Results
  • 4.1 Implementation and training details
  • 4.2 Model analysis
  • 4.3 Comparison with the-state-of-the-arts
  • 5 Conclusion
  • References

Knowls

  1. Knowl 1 — Network Architectures and Parameter Configurations of DBPN Variants

    model/method

    The Deep Back-Projection Network (DBPN) family includes six structural variants: DBPN-SS (Super Small), DBPN-S (Small), DBPN-M (Medium), DBPN-L (Large), D-DBPN-L (Dense DBPN-Large), and D-DBPN (Dense DBPN). Each architecture consists of initial feature extraction layers (Feat0\text{Feat0} and Feat1\text{Feat1}), iterative back-projection stages (TT stages of alternating up- and down-projection units), and a final reconstruction layer. In the layer notation conv(f,n,st,pd)\text{conv}(f, n, st, pd), ff denotes the spatial filter size, nn the number of feature filters, stst the stride, and pdpd the spatial padding.

    Scale / Property DBPN-SS DBPN-S DBPN-M DBPN-L D-DBPN-L D-DBPN
    Input/Output Luminance Luminance Luminance Luminance Luminance RGB
    Feat0 conv(3,64,1,1)\text{conv}(3,64,1,1) conv(3,128,1,1)\text{conv}(3,128,1,1) conv(3,128,1,1)\text{conv}(3,128,1,1) conv(3,128,1,1)\text{conv}(3,128,1,1) conv(3,128,1,1)\text{conv}(3,128,1,1) conv(3,256,1,1)\text{conv}(3,256,1,1)
    Feat1 conv(1,18,1,0)\text{conv}(1,18,1,0) conv(1,32,1,0)\text{conv}(1,32,1,0) conv(1,32,1,0)\text{conv}(1,32,1,0) conv(1,32,1,0)\text{conv}(1,32,1,0) conv(1,32,1,0)\text{conv}(1,32,1,0) conv(1,64,1,0)\text{conv}(1,64,1,0)
    Reconstruction conv(1,1,1,0)\text{conv}(1,1,1,0) conv(1,1,1,0)\text{conv}(1,1,1,0) conv(1,1,1,0)\text{conv}(1,1,1,0) conv(1,1,1,0)\text{conv}(1,1,1,0) conv(1,1,1,0)\text{conv}(1,1,1,0) conv(3,3,1,1)\text{conv}(3,3,1,1)
    BP stages (2×2\times) conv(6,18,2,2)\text{conv}(6,18,2,2) conv(6,32,2,2)\text{conv}(6,32,2,2) conv(6,32,2,2)\text{conv}(6,32,2,2) conv(6,32,2,2)\text{conv}(6,32,2,2) conv(6,32,2,2)\text{conv}(6,32,2,2) conv(6,64,2,2)\text{conv}(6,64,2,2)
    BP stages (4×4\times) conv(8,18,4,2)\text{conv}(8,18,4,2) conv(8,32,4,2)\text{conv}(8,32,4,2) conv(8,32,4,2)\text{conv}(8,32,4,2) conv(8,32,4,2)\text{conv}(8,32,4,2) conv(8,32,4,2)\text{conv}(8,32,4,2) conv(8,64,4,2)\text{conv}(8,64,4,2)
    BP stages (8×8\times) conv(12,18,8,2)\text{conv}(12,18,8,2) conv(12,32,8,2)\text{conv}(12,32,8,2) conv(12,32,8,2)\text{conv}(12,32,8,2) conv(12,32,8,2)\text{conv}(12,32,8,2) conv(12,32,8,2)\text{conv}(12,32,8,2) conv(12,64,8,2)\text{conv}(12,64,8,2)
    Parameters (2×2\times) 106k106\text{k} 337k337\text{k} 779k779\text{k} 1221k1221\text{k} 1230k1230\text{k} 5819k5819\text{k}
    Parameters (4×4\times) 188k188\text{k} 595k595\text{k} 1381k1381\text{k} 2168k2168\text{k} 2176k2176\text{k} 10291k10291\text{k}
    Parameters (8×8\times) 421k421\text{k} 1332k1332\text{k} 3101k3101\text{k} 4871k4871\text{k} 4879k4879\text{k} 23071k23071\text{k}
    Depth 12 12 24 36 40 52
    No. of stages (TT) 2 2 4 6 6 7
    Dense connection No No No No Yes Yes

    Dense connection indicates whether feature maps generated by all previous projection units are concatenated and passed into each subsequent unit.

  2. Knowl 2 — Super-Resolution Inference Runtime Comparison

    data/table

    Inference time (in seconds) was evaluated across super-resolution architectures on an Nvidia TITAN X GPU (12 GB12\text{ GB} VRAM) with an input patch size of 64×6464 \times 64 pixels upscaled to 128×128128 \times 128 (2×2\times), 256×256256 \times 256 (4×4\times), and 512×512512 \times 512 (8×8\times). Runtimes represent the average of 10 trials.

    Algorithm 2×2\times (128×128128\times 128) 4×4\times (256×256256\times 256) 8×8\times (512×512512\times 512)
    VDSR 0.02223 0.03225 0.06856
    DRRN 0.25413 0.32893 N.A. (Out of Memory)
    EDSR 0.85790 1.24580 1.14770
    DBPN-SS – 0.01672 0.02692
    DBPN-S – 0.02073 0.03812
    DBPN-M – 0.04511 0.08106
    DBPN-L – 0.06971 0.12635
    D-DBPN 0.15331 0.19396 0.31851

    DBPN-SS and DBPN-S obtain the fastest and second-fastest runtimes for 4×4\times and 8×8\times enlargement. D-DBPN runs faster than EDSR across all scaling factors (0.15331 s0.15331\text{ s} vs. 0.8579 s0.8579\text{ s} for 2×2\times, 0.19396 s0.19396\text{ s} vs. 1.2458 s1.2458\text{ s} for 4×4\times, and 0.31851 s0.31851\text{ s} vs. 1.1477 s1.1477\text{ s} for 8×8\times). VDSR and DRRN only process the luminance channel and require bicubic interpolation preprocessing, which is not included in the network forward time.

  3. Knowl 3 — Ablation of Error Feedback in DBPN-S

    empirical result

    Error feedback (EF) in back-projection units guides early reconstruction by mapping generated high-resolution features back to the low-resolution domain to compute a residual error signal. In the baseline setting without EF, up- and down-projection units are replaced with standard non-iterative deconvolution and convolution layers.

    Evaluating DBPN-S on 4×4\times enlargement demonstrates that error feedback provides substantial quantitative gains:

    Model Set5 (PSNR dB) Set14 (PSNR dB)
    SRCNN 30.49 27.61
    FSRCNN 30.71 27.70
    DBPN-S (Without EF) 31.06 27.95
    DBPN-S (With EF) 31.59 28.21

    Error feedback improves DBPN-S performance by +0.53 dB+0.53\text{ dB} on Set5 and +0.26 dB+0.26\text{ dB} on Set14. Even without EF, the mutual low-to-high and high-to-low projection structure outperforms SRCNN by 0.57 dB0.57\text{ dB} and FSRCNN by 0.35 dB0.35\text{ dB} on Set5.

  4. Knowl 4 — Controlled Capacity and Dataset Comparison Between DBPN-S and LapSRN

    empirical result

    To control for network capacity and training dataset size, DBPN-S (595k595\text{k} parameters) is compared against LapSRN (812k812\text{k} parameters) on 4×4\times super-resolution when both models are trained on the 800-image DIV2K dataset, as well as against the original LapSRN trained on BSDS200 + T91.

    DBPN-S (DIV2K) LapSRN (DIV2K) LapSRN (BSDS200+T91)
    Dataset PSNR (dB) SSIM PSNR (dB) SSIM PSNR (dB) SSIM
    Set5 31.57 0.886 31.64 0.886 31.54 0.885
    Set14 28.20 0.771 28.25 0.772 28.19 0.772
    BSDS100 27.38 0.728 27.36 0.728 27.32 0.728
    Urban100 25.49 0.762 25.37 0.759 25.21 0.756
    Manga109 29.39 0.891 29.24 0.891 29.09 0.890

    Despite having 26.7%26.7\% fewer parameters, DBPN-S outperforms LapSRN trained on the same DIV2K dataset on BSDS100 (+0.02 dB+0.02\text{ dB}), Urban100 (+0.12 dB+0.12\text{ dB}, +0.003+0.003 SSIM), and Manga109 (+0.15 dB+0.15\text{ dB}).

  5. Knowl 5 — Filter Size Selection for Back-Projection Stages

    empirical result

    The choice of convolutional filter spatial dimensions inside the up- and down-projection stages of D-DBPN affects 4×4\times reconstruction quality. Evaluating filter sizes of 6×66 \times 6, 8×88 \times 8, and 10×1010 \times 10 with striding of 4 on Set5 and Set14 benchmarks yields:

    Filter Size Striding Padding Set5 PSNR (dB) Set14 PSNR (dB)
    6×66 \times 6 4 1 32.39 28.78
    8×88 \times 8 4 2 32.47 28.82
    10×1010 \times 10 4 3 32.38 28.79

    A filter size of 8×88 \times 8 provides the best performance, outperforming 6×66 \times 6 by 0.08 dB0.08\text{ dB} on Set5 (0.04 dB0.04\text{ dB} on Set14) and outperforming 10×1010 \times 10 by 0.09 dB0.09\text{ dB} on Set5 (0.03 dB0.03\text{ dB} on Set14).

  6. Knowl 6 — Comparison of RGB and Luminance Color Processing in DBPN-L

    empirical result

    DBPN-L evaluated on 4×4\times image super-resolution shows negligible difference when processing all three RGB color channels directly compared to processing only the luminance (Y) channel:

    Channel Configuration Set5 PSNR (dB) Set14 PSNR (dB)
    RGB 31.88 28.47
    Luminance 31.86 28.47

    Direct RGB processing matches or slightly exceeds luminance PSNR (+0.02 dB+0.02\text{ dB} on Set5, identical on Set14) while simplifying the super-resolution pipeline by eliminating the need for separate bicubic interpolation on chrominance channels.

  7. Knowl 7 — Convergence Characteristics of DBPN Architectures

    empirical result

    Across both 4×4\times and 8×8\times enlargement scales on the Set5 benchmark, DBPN variants demonstrate fast convergence. In particular, D-DBPN reaches a PSNR within 50,000 training iterations that outperforms competing baseline models (DRCN, LapSRN, DRRN, and VDSR). D-DBPN reaches asymptotic performance around 500k500\text{k} to 600k600\text{k} iterations, achieving over 32.4 dB32.4\text{ dB} on 4×4\times and over 27.1 dB27.1\text{ dB} on 8×8\times.

Coverage note — Qualitative visual comparison figures (Figures 4 through 16) displaying specific image crops for 8x super-resolution across Manga109 and Urban100 were omitted because qualitative visual images cannot be represented as standalone quantitative knowls, and their empirical conclusions are captured in the quantitative evaluation tables.

References

  1. 1.P. Arbelaez, M. Maire, C. Fowlkes, and J. Malik. Con- tour detection and hierarchical image segmentation. IEEE transactions on pattern analysis and machine intelligence, 33(5):898–916, 2011. 1
  2. 2.C. Dong, C. C. Loy, K. He, and X. Tang. Image super-resolution using deep convolutional networks. IEEE transactions on pattern analysis and machine intelligence, 38(2):295–307, 2016. 1, 2
  3. 3.C. Dong, C. C. Loy, and X. Tang. Accelerating the super- resolution convolutional neural network. In European Con- ference on Computer Vision, pages 391–407. Springer, 2016. 1, 2
  4. 4.J. Kim, J. Kwon Lee, and K. Mu Lee. Accurate image super- resolution using very deep convolutional networks. In Pro- ceedings of the IEEE Conference on Computer Vision and Pat- tern Recognition, pages 1646–1654, June 2016. 2, 4
  5. 5.W.-S. Lai, J.-B. Huang, N. Ahuja, and M.-H. Yang. Deep laplacian pyramid networks for fast and accurate super- resolution. In IEEE Conferene on Computer Vision and Pat- tern Recognition, 2017. 1, 2, 4
  6. 6.B. Lim, S. Son, H. Kim, S. Nah, and K. M. Lee. Enhanced deep residual networks for single image super-resolution. In The IEEE Conference on Computer Vision and Pattern Recog- nition (CVPR) Workshops, July 2017. 2, 4
  7. 7.Y. Tai, J. Yang, and X. Liu. Image super-resolution via deep recursive residual network. In Proceedings of the IEEE Con- ference on Computer Vision and Pattern Recognition, 2017. 2, 4
  8. 8.R. Timofte, E. Agustsson, L. Van Gool, M.-H. Yang, L. Zhang, B. Lim, S. Son, H. Kim, S. Nah, K. M. Lee, et al. Ntire 2017 challenge on single image super-resolution: Meth- ods and results. In Computer Vision and Pattern Recogni- tion Workshops (CVPRW), 2017 IEEE Conference on, pages 1110–1121. IEEE, 2017. 1
  9. 9.J. Yang, J. Wright, T. S. Huang, and Y. Ma. Image super- resolution via sparse representation. Image Processing, IEEE Transactions on, 19(11):2861–2873, 2010. 1

Citation

MLA
Haris, M., et al. “Deep Back-Projection Networks For Super-Resolution”. arXiv, 2018, http://arxiv.org/abs/1803.02735v1.
APA
Haris, M., Shakhnarovich, G., & Ukita, N. (2018). Deep Back-Projection Networks For Super-Resolution. arXiv. http://arxiv.org/abs/1803.02735v1
Chicago
Haris, M., G. Shakhnarovich, and N. Ukita. 2018. “Deep Back-Projection Networks For Super-Resolution”. arXiv. http://arxiv.org/abs/1803.02735v1.
Harvard
Haris, M., Shakhnarovich, G. and Ukita, N. (2018) “Deep Back-Projection Networks For Super-Resolution”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1803.02735v1.
Vancouver
1. Haris M, Shakhnarovich G, Ukita N (2018) Deep Back-Projection Networks For Super-Resolution. arXiv

BibTeX

@article{haris2018deep,
  title = {Deep Back-Projection Networks For Super-Resolution},
  author = {Haris, Muhammad and Shakhnarovich, Greg and Ukita, Norimichi},
  year = {2018},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1803.02735v1},
  eprint = {1803.02735}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE