Fast, Accurate, and, Lightweight Super-Resolution with Cascading Residual Network

Namhyuk AhnByungkon KangKyung-ah Sohn

article2018ECCV1,446 citations

Proposes a cascading residual network for single-image super-resolution that matches state-of-the-art accuracy while dramatically cutting parameters and computational cost to enable practical real-world deployment.

Listen

Single-image super-resolution reconstructs high-resolution images from low-resolution inputs, playing a vital role in consumer mobile applications, video streaming services, and surveillance systems. While deep learning methods have drastically improved image restoration quality, state-of-the-art architectures demand substantial computational operations and memory. This computational weight creates severe operational bottlenecks, including high latency and excessive battery consumption on edge devices.

The main objective of the article is to develop and evaluate high-performing, lightweight deep learning architectures that significantly reduce computational operations and parameter sizes while preserving image reconstruction accuracy. To accomplish this, the authors introduce the Cascading Residual Network (CARN) and its mobile-optimized variant (CARN-M).

The approach designs neural network blocks that connect intermediate layers across both local and global cascading pathways, upsampling the image only at the final stage to avoid heavy intermediate calculations. To create the mobile variant, the authors incorporate group convolutions into an efficient residual block and share parameters recursively. The models were trained on the standard DIV2K dataset using the L1 loss function and evaluated against established benchmarks across standard test sets, including Set5, Set14, B100, and Urban100, at multiple scaling factors.

The findings demonstrate that CARN outperforms existing models of comparable size (under five million parameters) across all standard benchmarks while maintaining a modest computational burden of 90.9 billion multiply-accumulate operations for 720p resolution at 4x scaling. The streamlined CARN-M variant reduces the parameter count by approximately 74% (from 1,592K to 412K) and cuts computational operations by roughly 64% (from 90.9G to 32.5G) relative to CARN, while incurring only a slight drop of 0.29 dB in peak signal-to-noise ratio. The ablation analysis revealed that combining local and global cascading is essential: using local cascading alone degraded performance due to optimization bottlenecks, whereas global shortcuts restored effective information and gradient flow. Furthermore, multi-scale learning allowed a single trained model to handle multiple upscaling factors simultaneously.

These results show that high visual quality does not require massive computing infrastructure. By decoupling image enhancement from excessive processing demands, organizations can deploy real-time super-resolution on resource-constrained mobile hardware and significantly lower bandwidth and storage expenses in media streaming via on-the-fly decompression. Unlike prior recursive approaches that maintained accuracy only by increasing network depth and latency, the proposed cascading framework achieves high fidelity with high execution efficiency.

For practical implementation, engineering teams should adopt CARN when image fidelity is paramount within a moderate compute budget, and deploy CARN-M when targeting battery-sensitive mobile platforms or ultra-low-latency streaming pipelines. Future development should focus on extending this cascading methodology directly to real-time video streaming architectures to validate end-to-end compression and decompression workflows.

The findings carry high confidence across standard static image benchmarks, though decision-makers should note that evaluations were conducted primarily on standard benchmark image sets. Performance across continuous, real-time video streams under varying network conditions remains subject to future empirical validation.

Cover for Fast, Accurate, and, Lightweight Super-Resolution with Cascading Residual Network

Abstract

In recent years, deep learning methods have been successfully applied to single-image super-resolution tasks. Despite their great performances, deep learning methods cannot be easily applied to real-world applications due to the requirement of heavy computation. In this paper, we address this issue by proposing an accurate and lightweight deep network for image super-resolution. In detail, we design an architecture that implements a cascading mechanism upon a residual network. We also present variant models of the proposed cascading residual network to further improve efficiency. Our extensive experiments show that even with much fewer parameters and operations, our models achieve performance comparable to that of state-of-the-art methods.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Deep Learning Based Image Super-Resolution
  • 2.2 Efficient Neural Network
  • 3 Proposed Method
  • 3.1 Cascading Residual Network
  • 3.2 Efficient Cascading Residual Network
  • 3.3 Comparison to Recent Models
  • 4 Experimental Results
  • 4.1 Datasets
  • 4.2 Implementation and Training Details
  • 4.3 Comparison with State-of-the-art Methods
  • 4.4 Model Analysis
  • Cascading Modules.
  • Efficiency Trade-off.
  • 5 Conclusion
  • References

Knowls

  1. Knowl 1 — Cascading Residual Network (CARN) Architecture

    model/method

    The Cascading Residual Network (CARN) is a deep convolutional neural network designed for fast, accurate, and lightweight single-image super-resolution (SISR). Instead of upsampling the input low-resolution (LR) image at the beginning of the network, CARN extracts and processes features entirely at the LR resolution and upsamples to high resolution (HR) at the end of the network using sub-pixel convolution (pixel shuffle).

    The CARN architecture consists of three main parts:

    1. Initial Feature Extraction: A single convolutional layer extracts shallow feature representations from the input LR image XX.
    2. Cascading Backbone (Global and Local Cascading): A sequence of B=3B=3 cascading blocks with global cascading connections connecting the output of the initial convolution and each preceding cascading block into subsequent cascading blocks. Inside each cascading block, U=3U=3 residual blocks are arranged with local cascading connections. At both global and local cascading junctions, intermediate multi-level representations are concatenated channel-wise and compressed using a 1×11\times 1 convolutional layer.
    3. Multi-Scale Reconstruction and Upsampling: The aggregated multi-level feature map is fed into scale-specific upsampling modules composed of convolutions and pixel shuffle operations to generate the HR image for scaling factors ×2\times 2, ×3\times 3, or ×4\times 4.

    By integrating cascading mechanisms at both global (layer-wise) and local (block-wise) levels, CARN facilitates multi-level representation learning and creates multiple shortcut paths that assist gradient flow during optimization.

  2. Knowl 2 — Mathematical Formulation of Multi-Level Cascading

    equation

    Let f(⋅;W)f(\cdot; W) denote a convolution operation parameterized by weights WW, and τ(⋅)\tau(\cdot) denote the ReLU activation function.

    A standard residual block RiR_i with parameter set WRi={WRi,1,WRi,2}W_R^i = \{W_R^{i,1}, W_R^{i,2}\} operating on an input feature map Hi−1H^{i-1} is defined as: Ri(Hi−1;WRi)=τ(f(τ(f(Hi−1;WRi,1));WRi,2)+Hi−1)R_i(H^{i-1}; W_R^i) = \tau\left(f\left(\tau\left(f\left(H^{i-1}; W_R^{i,1}\right)\right); W_R^{i,2}\right) + H^{i-1}\right)

    The ii-th local cascading block Blocali(Hi−1;Wli)≡Bi,UB_{\text{local}}^i(H^{i-1}; W_l^i) \equiv B^{i,U} containing UU residual units is defined recursively for u=1,…,Uu = 1, \dots, U by: Bi,0=Hi−1B^{i,0} = H^{i-1} Bi,u=f([Bi,0,Bi,1,…,Bi,u−1,Ru(Bi,u−1;WRu)];Wci,u)B^{i,u} = f\left(\left[B^{i,0}, B^{i,1}, \dots, B^{i,u-1}, R^u\left(B^{i,u-1}; W_R^u\right)\right]; W_c^{i,u}\right) where [⋅][\cdot] represents channel-wise feature concatenation and Wci,uW_c^{i,u} denotes the parameters of the 1×11\times 1 convolution compressing the concatenated features.

    The global cascading process over BB cascading blocks computes intermediate representations HbH^b for b=1,…,Bb = 1, \dots, B from an input image XX as: H0=f(X;Wc)H^0 = f(X; W_c) Hb=f([H0,H1,…,Hb−1,Blocalb(Hb−1;WBb)];Wgb)H^b = f\left(\left[H^0, H^1, \dots, H^{b-1}, B_{\text{local}}^b\left(H^{b-1}; W_B^b\right)\right]; W_g^b\right) where WcW_c is the parameter set of the initial convolution layer and WgbW_g^b represents the parameters of the 1×11\times 1 convolution merging all preceding global representations. In CARN, U=3U = 3 and B=3B = 3.

  3. Knowl 3 — Efficient Residual Block (residual-E) and Computation Reduction Ratio

    equation

    The efficient residual block (residual-E) replaces standard convolutions in a residual block with group convolutions and a pointwise (1×11\times 1) convolution to make computational cost tunable.

    A residual-E block consists of two 3×33\times 3 group convolutions with group size GG, followed by one 1×11\times 1 pointwise convolution. Let KK denote the kernel size (K=3K=3), CinC_{\text{in}} and CoutC_{\text{out}} denote the number of input and output channels (kept constant at 64), and F×FF \times F denote the spatial resolution of the input and output feature maps.

    The computational cost (in multiply-accumulate operations) of a standard residual block is: Coststandard=2×(K⋅K⋅Cin⋅Cout⋅F⋅F)\text{Cost}_{\text{standard}} = 2 \times \left(K \cdot K \cdot C_{\text{in}} \cdot C_{\text{out}} \cdot F \cdot F\right)

    The computational cost of a residual-E block is: Costresidual-E=2×(K⋅K⋅Cin⋅CoutG⋅F⋅F)+Cin⋅Cout⋅F⋅F\text{Cost}_{\text{residual-E}} = 2 \times \left(K \cdot K \cdot C_{\text{in}} \cdot \frac{C_{\text{out}}}{G} \cdot F \cdot F\right) + C_{\text{in}} \cdot C_{\text{out}} \cdot F \cdot F

    The theoretical computation reduction ratio of the residual-E block relative to the standard residual block is: Costresidual-ECoststandard=2×(K⋅K⋅Cin⋅CoutG⋅F⋅F)+Cin⋅Cout⋅F⋅F2×(K⋅K⋅Cin⋅Cout⋅F⋅F)=1G+12K2\frac{\text{Cost}_{\text{residual-E}}}{\text{Cost}_{\text{standard}}} = \frac{2 \times \left(K \cdot K \cdot C_{\text{in}} \cdot \frac{C_{\text{out}}}{G} \cdot F \cdot F\right) + C_{\text{in}} \cdot C_{\text{out}} \cdot F \cdot F}{2 \times \left(K \cdot K \cdot C_{\text{in}} \cdot C_{\text{out}} \cdot F \cdot F\right)} = \frac{1}{G} + \frac{1}{2K^2}

    For K=3K=3, the reduction ratio is 1G+118\frac{1}{G} + \frac{1}{18}, reducing convolution computation by a factor between 1.8×1.8\times (for G=2G=2) and 14×14\times (for G=64G=64).

  4. Knowl 4 — CARN-Mobile (CARN-M) Architecture

    model/method

    CARN-Mobile (CARN-M) is an ultra-lightweight variant of the Cascading Residual Network designed to optimize both parameter footprint and floating-point operations (Mult-Adds) for mobile and real-time super-resolution applications.

    CARN-M alters the base CARN design via two modifications:

    1. Efficient Residual Blocks (residual-E): Standard residual blocks inside all cascading blocks are replaced with residual-E blocks, setting the group convolution group size to G=4G=4.
    2. Recursive Parameter Sharing: The parameters across all cascading blocks are shared, converting the global cascading backbone into a recursive cascading architecture. Parameter sharing reduces the parameter count of the cascading blocks by up to three times.

    Through these modifications, CARN-M reduces parameter count from 1,592K (CARN) down to 412K (a nearly 4×4\times reduction) and cuts Mult-Adds by approximately 2.5×2.5\times to 2.8×2.8\times across ×2\times 2, ×3\times 3, and ×4\times 4 scales, with only a modest 0.29 dB average PSNR penalty relative to CARN.

  5. Knowl 5 — Uniform Weight Initialization for Cascading Networks with 1x1 Convolutions

    model/method

    Standard weight initialization routines (such as Xavier/Glorot or He initialization) assign excessively large initial weight values to narrow 1×11\times 1 convolutional layers present in cascading shortcuts, which causes numerical instability and training failure in cascading architectures.

    To ensure stable convergence, all weights and biases θ\theta across CARN and CARN-M are initialized from a uniform distribution: θ∼U(−k,k),with k=1cin\theta \sim \mathcal{U}(-k, k), \quad \text{with } k = \frac{1}{\sqrt{c_{\text{in}}}} where cinc_{\text{in}} is the number of input channels to the respective convolutional layer. This initialization prevents exploding activation and gradient values during backpropagation through cascading connections.

  6. Knowl 6 — Quantitative SISR Benchmark Comparison

    data/table

    CARN and CARN-M were benchmarked against state-of-the-art deep learning SISR models on Set5, Set14, B100, and Urban100 datasets across scaling factors ×2\times 2, ×3\times 3, and ×4\times 4. Multiply-accumulate operations (Mult-Adds) are calculated assuming a 720p720\text{p} (1280×7201280 \times 720) output image.

    Scale Model Params Mult-Adds Set5 PSNR/SSIM Set14 PSNR/SSIM B100 PSNR/SSIM
    ×2\times 2 SRCNN 57K 52.7G 36.66 / 0.9542 32.42 / 0.9063 31.36 / 0.8879
    ×2\times 2 FSRCNN 12K 6.0G 37.00 / 0.9558 32.63 / 0.9088 31.53 / 0.8920
    ×2\times 2 VDSR 665K 612.6G 37.53 / 0.9587 33.03 / 0.9124 31.90 / 0.8960
    ×2\times 2 DRCN 1,774K 17,974.3G 37.63 / 0.9588 33.04 / 0.9118 31.85 / 0.8942
    ×2\times 2 DRRN 297K 6,796.9G 37.74 / 0.9591 33.23 / 0.9136 32.05 / 0.8973
    ×2\times 2 MemNet 677K 2,662.4G 37.78 / 0.9597 33.28 / 0.9142 32.08 / 0.8978
    ×2\times 2 SelNet 974K 225.7G 37.89 / 0.9598 33.61 / 0.9160 32.08 / 0.8984
    ×2\times 2 CARN (ours) 1,592K 222.8G 37.76 / 0.9590 33.52 / 0.9166 32.09 / 0.8978
    ×2\times 2 CARN-M (ours) 412K 91.2G 37.53 / 0.9583 33.26 / 0.9141 31.92 / 0.8960
    ×3\times 3 SRCNN 57K 52.7G 32.75 / 0.9090 29.28 / 0.8209 28.41 / 0.7863
    ×3\times 3 FSRCNN 12K 5.0G 33.16 / 0.9140 29.43 / 0.8242 28.53 / 0.7910
    ×3\times 3 VDSR 665K 612.6G 33.66 / 0.9213 29.77 / 0.8314 28.82 / 0.7976
    ×3\times 3 DRRN 297K 6,796.9G 34.03 / 0.9244 29.96 / 0.8349 28.95 / 0.8004
    ×3\times 3 SelNet 1,159K 120.0G 34.27 / 0.9257 30.30 / 0.8399 28.97 / 0.8025
    ×3\times 3 CARN (ours) 1,592K 118.8G 34.29 / 0.9255 30.29 / 0.8407 29.06 / 0.8034
    ×3\times 3 CARN-M (ours) 412K 46.1G 33.99 / 0.9236 30.08 / 0.8367 28.91 / 0.8000
    ×4\times 4 SRCNN 57K 52.7G 30.48 / 0.8628 27.49 / 0.7503 26.90 / 0.7101
    ×4\times 4 FSRCNN 12K 4.6G 30.71 / 0.8657 27.59 / 0.7535 26.98 / 0.7150
    ×4\times 4 VDSR 665K 612.6G 31.35 / 0.8838 28.01 / 0.7674 27.29 / 0.7251
    ×4\times 4 DRRN 297K 6,796.9G 31.68 / 0.8888 28.21 / 0.7720 27.38 / 0.7284
    ×4\times 4 SelNet 1,417K 83.1G 32.00 / 0.8931 28.49 / 0.7783 27.44 / 0.7325
    ×4\times 4 SRDenseNet 2,015K 389.9G 32.02 / 0.8934 28.50 / 0.7782 27.53 / 0.7337
    ×4\times 4 CARN (ours) 1,592K 90.9G 32.13 / 0.8937 28.60 / 0.7806 27.58 / 0.7349
    ×4\times 4 CARN-M (ours) 412K 32.5G 31.92 / 0.8903 28.42 / 0.7762 27.44 / 0.7304

    On Urban100 (×4\times 4), CARN achieves 26.07 dB/0.783726.07\text{ dB} / 0.7837 SSIM with 90.9G90.9\text{G} Mult-Adds, outperforming SRDenseNet (26.05 dB/0.781926.05\text{ dB} / 0.7819 at 389.9G389.9\text{G} Mult-Adds) and DRRN (25.44 dB/0.763825.44\text{ dB} / 0.7638 at 6,796.9G6,796.9\text{G} Mult-Adds). CARN-M achieves 25.62 dB/0.769425.62\text{ dB} / 0.7694 SSIM on Urban100 (×4\times 4) with only 32.5G32.5\text{G} Mult-Adds and 412K412\text{K} parameters.

  7. Knowl 7 — Ablation of Local and Global Cascading Modules

    data/table

    An ablation study on Set14 (×4\times 4 scale) isolates the contribution of local and global cascading mechanisms relative to a plain ResNet baseline.

    Architecture Baseline (ResNet) CARN-NL CARN-NG CARN
    Local Cascading - - ✓ ✓
    Global Cascading - ✓ - ✓
    # Parameters 1,444K 1,481K 1,555K 1,592K
    PSNR (dB) 28.43 28.45 28.42 28.52

    Key findings:

    1. Local cascading alone degrades performance: CARN-NG (local cascading only) drops to 28.42 dB PSNR, performing worse than the plain baseline (28.43 dB). Multiplicative 1×11\times 1 convolutions placed along local shortcut paths hinder information propagation and gradient flow during training.
    2. Global cascading aids multi-level propagation: CARN-NL (global cascading only) achieves 28.45 dB PSNR by effectively transferring multi-level frequency signals from shallow to deep layers.
    3. Dual cascading synergy: Combining local and global cascading (CARN) attains the highest performance (28.52 dB PSNR). Global cascading connections provide direct bypasses that resolve the optimization bottleneck caused by local 1×11\times 1 convolutions, allowing the network to fully leverage multi-level local feature representations.
  8. Knowl 8 — Group Size and Recursion Trade-offs in CARN-M

    empirical result

    Analyzing group sizes G∈{2,4,8,16,32,64}G \in \{2, 4, 8, 16, 32, 64\} in efficient residual blocks (residual-E) across non-recursive and recursive configurations on Set14 (×4\times 4) demonstrates the following:

    1. Group Size vs. Performance: Increasing group size GG reduces Mult-Adds and parameters monotonically. For example, G=64G=64 achieves a 5×5\times reduction in operations and parameters compared to G=1G=1, but incurs noticeable PSNR degradation.
    2. Effect of Recursive Sharing: Applying recursive weight sharing across cascading blocks reduces the parameter count by up to 3×3\times without changing the total Mult-Adds. For a fixed parameter budget, recursive models consistently achieve higher PSNR than non-recursive models with larger group sizes.
    3. Optimal Configuration (G4R): Setting group size G=4G=4 with recursive cascading blocks (designated G4R) delivers the optimal trade-off. Relative to full CARN, CARN-M (G4R) reduces parameters by 3.8×3.8\times (from 1,592K to 412K) and operations by 2.8×2.8\times (from 90.9G to 32.5G Mult-Adds) with a minor PSNR loss of 0.29 dB.
  9. Knowl 9 — Training Methodology and Multi-Scale Training Setup

    experimental setup

    CARN and CARN-M are trained under the following unified experimental protocol:

    • Dataset: DIV2K dataset (800 training high-resolution images). Training inputs are 64×6464 \times 64 RGB patches randomly cropped from LR images, augmented with random horizontal flips and 90∘90^\circ rotations.
    • Loss Function: L1L_1 loss between reconstructed and ground truth HR patches. Compared to L2L_2 loss, L1L_1 provides superior convergence stability and higher reconstruction quality in residual SISR architectures.
    • Optimizer and Hyperparameters: ADAM optimizer with β1=0.9\beta_1 = 0.9, β2=0.999\beta_2 = 0.999, and ϵ=10−8\epsilon = 10^{-8}. Minibatch size is 64. The model is trained for 6×1056 \times 10^5 iterations with an initial learning rate of 10−410^{-4}, halved every 4×1054 \times 10^5 iterations.
    • Multi-Scale Learning: Mini-batches cycle through scaling factors ×2\times 2, ×3\times 3, and ×4\times 4. A single trained CARN(-M) network shares feature extraction layers across all scales and activates the scale-specific sub-pixel upsampling head corresponding to the target scale, producing a unified multi-scale model file of only 1.6 MB for CARN-M.

Coverage note — Visual qualitative comparisons (Fig. 6) and background comparisons to SRDenseNet/MemNet were omitted as standalone knowls as their core insights are subsumed by the benchmark data and model architecture knowls.

References

  1. 1.Agustsson, E., Timofte, R.: Ntire 2017 challenge on single image super-resolution: Dataset and study. In: Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR) Workshops (2017)
  2. 2.Arbelaez, P., Maire, M., Fowlkes, C., Malik, J.: Contour detection and hierarchical image segmentation. IEEE transactions on pattern analysis and machine intelligence 33(5), 898–916 (2011)
  3. 3.Bevilacqua, M., Roumy, A., Guillemot, C., Alberi-Morel, M.: Low-complexity single-image super-resolution based on nonnegative neighbor embedding. In: Proceedings of the British Machine Vision Conference (BMVC) (2012)
  4. 4.Choi, J.S., Kim, M.: A deep convolutional neural network with selection units for super-resolution. In: Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR) Workshops (2017)
  5. 5.Deng, J., Dong, W., Socher, R., Li, L.J., Li, K., Fei-Fei, L.: Imagenet: A large-scale hierarchical image database. In: Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR) (2009)
  6. 6.Dong, C., Loy, C.C., He, K., Tang, X.: Learning a deep convolutional network for image super-resolution. In: Proceedings of the European Conference on Computer Vision (ECCV) (2014)
  7. 7.Dong, C., Loy, C.C., Tang, X.: Accelerating the super-resolution convolutional neural network. In: Proceedings of the European Conference on Computer Vision (ECCV) (2016)
  8. 8.Fan, Y., Shi, H., Yu, J., Liu, D., Han, W., Yu, H., Wang, Z., Wang, X., Huang, T.S.: Balanced two-stage residual networks for image super-resolution. In: Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR) Workshops (2017)
  9. 9.Girshick, R.: Fast r-cnn. In: Proceedings of the International Conference on Computer Vision (ICCV) (2015)
  10. 10.Glorot, X., Bengio, Y.: Understanding the difficulty of training deep feedforward neural networks. In: Proceedings of the International Conference on Artificial Intelligence and Statistics (2010)
  11. 11.Han, S., Mao, H., Dally, W.J.: Deep compression: Compressing deep neural networks with pruning, trained quantization and huffman coding. Proceedings of the International Conference on Learning Representations (ICLR) (2016)
  12. 12.He, K., Zhang, X., Ren, S., Sun, J.: Delving deep into rectifiers: Surpassing human-level performance on imagenet classification. In: Proceedings of the International Conference on Computer Vision (ICCV) (2015)
  13. 13.He, K., Zhang, X., Ren, S., Sun, J.: Deep residual learning for image recognition. In: Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR) (2016)
  14. 14.He, K., Zhang, X., Ren, S., Sun, J.: Deep residual learning for image recognition. In: Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR) (2016)
  15. 15.He, K., Zhang, X., Ren, S., Sun, J.: Identity mappings in deep residual networks. In: Proceedings of the European Conference on Computer Vision (ECCV) (2016)
  16. 16.Howard, A.G., Zhu, M., Chen, B., Kalenichenko, D., Wang, W., Weyand, T., Andreetto, M., Adam, H.: Mobilenets: Efficient convolutional neural networks for mobile vision applications. arXiv preprint arXiv:1704.04861 (2017)
  17. 17.Huang, G., Liu, Z., van der Maaten, L., Weinberger, K.Q.: Densely connected convolutional networks. In: Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR) (2017)
  18. 18.Huang, J.B., Singh, A., Ahuja, N.: Single image super-resolution from transformed self-exemplars. In: Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR) (2015)
  19. 19.Iandola, F.N., Han, S., Moskewicz, M.W., Ashraf, K., Dally, W.J., Keutzer, K.: Squeezenet: Alexnet-level accuracy with 50x fewer parameters and¡ 0.5 mb model size. arXiv preprint arXiv:1602.07360 (2016)
  20. 20.Kim, J., Kwon Lee, J., Mu Lee, K.: Accurate image super-resolution using very deep convolutional networks. In: Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR) (2016)
  21. 21.Kim, J., Kwon Lee, J., Mu Lee, K.: Deeply-recursive convolutional network for image super-resolution. In: Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR) (2016)
  22. 22.Kingma, D.P., Ba, J.: Adam: A method for stochastic optimization. Proceedings of the International Conference on Learning Representations (ICLR) (2015)
  23. 23.Krizhevsky, A., Sutskever, I., Hinton, G.E.: Imagenet classification with deep convolutional neural networks. In: Proceedings of the Conference on Neural Information Processing Systems (NIPS) (2012)
  24. 24.Lai, W.S., Huang, J.B., Ahuja, N., Yang, M.H.: Deep laplacian pyramid networks for fast and accurate super-resolution. In: Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR) (2017)
  25. 25.Lee, J., Nam, J.: Multi-level and multi-scale feature aggregation using pretrained convolutional neural networks for music auto-tagging. IEEE Signal Processing Letters 24(8), 1208–1212 (2017)
  26. 26.Lim, B., Son, S., Kim, H., Nah, S., Lee, K.M.: Enhanced deep residual networks for single image super-resolution. In: Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR) Workshops (2017)
  27. 27.Liu, W., Anguelov, D., Erhan, D., Szegedy, C., Reed, S., Fu, C.Y., Berg, A.C.: Ssd: Single shot multibox detector. In: Proceedings of the European Conference on Computer Vision (ECCV) (2016)
  28. 28.Long, J., Shelhamer, E., Darrell, T.: Fully convolutional networks for semantic segmentation. In: Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR) (2015)
  29. 29.Martin, D., Fowlkes, C., Tal, D., Malik, J.: A database of human segmented natural images and its application to evaluating segmentation algorithms and measuring ecological statistics. In: Proceedings of the International Conference on Computer Vision (ICCV) (2001)
  30. 30.Noh, H., Hong, S., Han, B.: Learning deconvolution network for semantic segmentation. In: Proceedings of the International Conference on Computer Vision (ICCV) (2015)
  31. 31.Ren, H., El-Khamy, M., Lee, J.: Image super resolution based on fusing multiple convolution neural networks. In: Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR) Workshops (2017)
  32. 32.Ronneberger, O., Fischer, P., Brox, T.: U-net: Convolutional networks for biomedical image segmentation. In: Proceedings of the International Conference on Medical Image Computing and Computer-Assisted Intervention (2015)
  33. 33.Shi, W., Caballero, J., Huszar, F., Totz, J., Aitken, A.P., Bishop, R., Rueckert, D., Wang, Z.: Real-time single image and video super-resolution using an efficient sub-pixel convolutional neural network. In: Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR) (2016)
  34. 34.Sifre, L., Mallat, S.: Rigid-motion scattering for image classification. Ph.D. thesis, Citeseer (2014)
  35. 35.Tai, Y., Yang, J., Liu, X.: Image super-resolution via deep recursive residual network. In: Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR) (2017)
  36. 36.Tai, Y., Yang, J., Liu, X., Xu, C.: Memnet: A persistent memory network for image restoration. In: Proceedings of the International Conference on Computer Vision (ICCV) (2017)
  37. 37.Tong, T., Li, G., Liu, X., Gao, Q.: Image super-resolution using dense skip connections. In: Proceedings of the International Conference on Computer Vision (ICCV) (2017)
  38. 38.Wang, Z., Bovik, A.C., Sheikh, H.R., Simoncelli, E.P.: Image quality assessment: from error visibility to structural similarity. IEEE transactions on image processing 13(4), 600–612 (2004)
  39. 39.Yang, J., Wright, J., Huang, T.S., Ma, Y.: Image super-resolution via sparse representation. IEEE transactions on image processing 19(11), 2861–2873 (2010)
  40. 40.Zhang, R., Isola, P., Efros, A.A.: Colorful image colorization. In: Proceedings of the European Conference on Computer Vision (ECCV) (2016)

Citation

MLA
Ahn, N., et al. “Fast, Accurate, and Lightweight Super-Resolution with Cascading Residual Network”. arXiv, 2018, http://arxiv.org/abs/1803.08664v5.
APA
Ahn, N., Kang, B., & Sohn, K.-A. (2018). Fast, Accurate, and Lightweight Super-Resolution with Cascading Residual Network. arXiv. http://arxiv.org/abs/1803.08664v5
Chicago
Ahn, N., B. Kang, and K.-A. Sohn. 2018. “Fast, Accurate, and Lightweight Super-Resolution with Cascading Residual Network”. arXiv. http://arxiv.org/abs/1803.08664v5.
Harvard
Ahn, N., Kang, B. and Sohn, K.-A. (2018) “Fast, Accurate, and Lightweight Super-Resolution with Cascading Residual Network”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1803.08664v5.
Vancouver
1. Ahn N, Kang B, Sohn K-A (2018) Fast, Accurate, and Lightweight Super-Resolution with Cascading Residual Network. arXiv

BibTeX

@article{ahn2018fast,
  title = {Fast, Accurate, and Lightweight Super-Resolution with Cascading Residual Network},
  author = {Ahn, Namhyuk and Kang, Byungkon and Sohn, Kyung-Ah},
  year = {2018},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1803.08664v5},
  eprint = {1803.08664}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF