Deep Learning for Image Super-Resolution: A Survey

Zhihao WangJian ChenSteven C. H. Hoi

article2019TPAMI1,853 citations

Presents a comprehensive taxonomy of deep learning-based image super-resolution by categorizing supervised, unsupervised, and domain-specific models while systematically reviewing benchmark datasets, evaluation metrics, and critical open challenges.

Listen

Digital imaging systems across healthcare, surveillance, and automated analysis frequently produce low-resolution visual inputs due to hardware limitations, transmission bottlenecks, and environmental noise. Enhancing these low-quality images into high-resolution outputs is a fundamental yet mathematically ill-posed challenge, because multiple high-resolution interpretations can correspond to a single degraded input. While classical upscaling methods rely on fixed mathematical interpolations or hand-crafted statistical rules, deep learning has emerged as a transformative mechanism to reconstruct intricate, natural high-frequency details. Understanding how different neural architectures, optimization functions, and training strategies operate is essential for selecting and deploying practical visual enhancement solutions.

The article systematically reviews modern deep learning approaches for image super-resolution to provide a structural taxonomy of supervised, unsupervised, and domain-specific techniques. The authors conduct a comprehensive literature synthesis analyzing dozens of deep learning architectures, standardized benchmark datasets, evaluation frameworks, and downstream tasks such as face reconstruction, medical imaging, and video processing.

The review yields several critical findings regarding system design and real-world deployment. First, model frameworks that execute feature extraction in the low-dimensional input space before upsampling at the final network stages achieve substantial computational savings and memory efficiency compared to older pre-upsampling pipelines. Second, there is a fundamental and mathematically proven trade-off between mathematical distortion metrics and human perceptual quality; optimizing strictly for pixel fidelity yields blurry visuals, whereas generative adversarial techniques produce rich textures and superior perceptual scores at the cost of lower numerical pixel accuracy. Third, structural innovations such as residual learning, dense feature reuse, and attention mechanisms significantly enhance detail recovery without disproportionately ballooning parameter counts. Finally, conventional models trained on synthetic, bicubic downsampling degrade considerably when deployed on real-world imagery containing optical blur, sensor noise, and compression artifacts.

These findings indicate that technology leaders must balance computational resources against visual realism depending on the end application. Operational deployment in resource-constrained environmentssuch as mobile devices or live video feedsrequires lightweight designs like group convolutions or progressive upscaling rather than computationally heavy baseline networks that can take tens of seconds per frame. Furthermore, developers must select loss functions aligned with operational goals: pixel-fidelity metrics suit automated diagnostic systems where artificial artifacts introduce liability, while perceptual and adversarial losses better serve consumer-facing media and human surveillance.

Organizations pursuing super-resolution deployments should prioritize the development of unsupervised and weakly-supervised workflows that learn true camera degradation rather than artificial downscaling. Teams should pilot automated neural architecture search to find compact, real-time models and adopt unified, multi-faceted evaluation protocols combining structural similarity, blind quality evaluators, and task-specific downstream accuracy. Further work remains necessary to develop standardized blind assessment metrics and stabilize adversarial training regimes before deploying automated enhancement pipelines in high-stakes environments.

arXiv: 1902.06068
Cover for Deep Learning for Image Super-Resolution: A Survey

Abstract

Image Super-Resolution (SR) is an important class of image processing techniques to enhance the resolution of images and videos in computer vision. Recent years have witnessed remarkable progress of image super-resolution using deep learning techniques. This article aims to provide a comprehensive survey on recent advances of image super-resolution using deep learning approaches. In general, we can roughly group the existing studies of SR techniques into three major categories: supervised SR, unsupervised SR, and domain-specific SR. In addition, we also cover some other important issues, such as publicly available benchmark datasets and performance evaluation metrics. Finally, we conclude this survey by highlighting several future directions and open issues which should be further addressed by the community in the future.

Table of Contents

  • I Introduction
  • II Problem Setting and Terminology
  • II-A Problem Definitions
  • II-B Datasets for Super-resolution
  • II-C Image Quality Assessment
  • II-C1 Peak Signal-to-Noise Ratio
  • II-C2 Structural Similarity
  • II-C3 Mean Opinion Score
  • II-C4 Learning-based Perceptual Quality
  • II-C5 Task-based Evaluation
  • II-C6 Other IQA Methods
  • II-D Operating Channels
  • II-E Super-resolution Challenges
  • III Supervised Super-resolution
  • III-A Super-resolution Frameworks
  • III-A1 Pre-upsampling Super-resolution
  • III-A2 Post-upsampling Super-resolution
  • III-A3 Progressive Upsampling Super-resolution
  • III-A4 Iterative Up-and-down Sampling Super-resolution
  • III-B Upsampling Methods
  • III-B1 Interpolation-based Upsampling
  • III-B2 Learning-based Upsampling
  • III-C Network Design
  • III-C1 Residual Learning
  • III-C2 Recursive Learning
  • III-C3 Multi-path Learning
  • III-C4 Dense Connections
  • III-C5 Attention Mechanism
  • III-C6 Advanced Convolution
  • III-C7 Region-recursive Learning
  • III-C8 Pyramid Pooling
  • III-C9 Wavelet Transformation
  • III-C10 Desubpixel
  • III-C11 xUnit
  • III-D Learning Strategies
  • III-D1 Loss Functions
  • III-D2 Batch Normalization
  • III-D3 Curriculum Learning
  • III-D4 Multi-supervision
  • III-E Other Improvements
  • III-E1 Context-wise Network Fusion
  • III-E2 Data Augmentation
  • III-E3 Multi-task Learning
  • III-E4 Network Interpolation
  • III-E5 Self-Ensemble
  • III-F State-of-the-art Super-resolution Models
  • IV Unsupervised Super-resolution
  • IV-A Zero-shot Super-resolution
  • IV-B Weakly-supervised Super-resolution
  • IV-C Deep Image Prior
  • V Domain-Specific Applications
  • V-A Depth Map Super-resolution
  • V-B Face Image Super-resolution
  • V-C Hyperspectral Image Super-resolution
  • V-D Real-world Image Super-resolution
  • V-E Video Super-resolution
  • V-F Other Applications
  • VI Conclusion and Future Directions
  • VI-A Network Design
  • VI-B Learning Strategies
  • VI-C Evaluation Metrics
  • VI-D Unsupervised Super-resolution
  • VI-E Towards Real-world Scenarios
  • References

Knowls

  1. Knowl 1 — Mathematical Formulation of Image Super-Resolution and Degradation Models

    definition

    Image super-resolution (SR) aims to reconstruct a high-resolution (HR) image approximation I^y\hat{I}_y from a low-resolution (LR) image IxI_x. The LR image is modeled as the output of an unknown degradation process D\mathcal{D} with degradation parameters δ\delta applied to the ground truth HR image IyI_y:

    Ix=D(Iy;δ)I_x = \mathcal{D}(I_y; \delta)

    In standard benchmark settings, degradation is modeled as a uniform downsampling operation s\downarrow_s by a scaling factor ss (typically using bicubic interpolation with anti-aliasing):

    D(Iy;δ)=(Iy)s,{s}δ\mathcal{D}(I_y; \delta) = (I_y) \downarrow_s, \quad \{s\} \subset \delta

    In more realistic scenarios closer to physical camera acquisitions, degradation is formulated as a composite mapping involving a blur convolution kernel κ\kappa, downsampling by factor ss, and additive white Gaussian noise nςn_\varsigma with standard deviation ς\varsigma:

    D(Iy;δ)=(Iyκ)s+nς,{κ,s,ς}δ\mathcal{D}(I_y; \delta) = (I_y \otimes \kappa) \downarrow_s + n_\varsigma, \quad \{\kappa, s, \varsigma\} \subset \delta

    The learning objective for an SR neural network F\mathcal{F} with learnable parameter set θ\theta is given by:

    θ^=argminθL(I^y,Iy)+λΦ(θ)\hat{\theta} = \arg\min_{\theta} \mathcal{L}(\hat{I}_y, I_y) + \lambda \Phi(\theta)

    where I^y=F(Ix;θ)\hat{I}_y = \mathcal{F}(I_x; \theta), L\mathcal{L} denotes the reconstruction loss function between the predicted HR image I^y\hat{I}_y and the ground truth IyI_y, Φ(θ)\Phi(\theta) is a regularization term on the model parameters, and λ\lambda is the trade-off hyperparameter.

  2. Knowl 2 — Deep Learning Super-Resolution Structural Frameworks

    model/method

    Deep learning-based super-resolution architectures are categorized into four fundamental frameworks based on where and how upsampling operations occur:

    1. Pre-upsampling Super-Resolution: The LR image is first upsampled to the target HR spatial dimensions using a non-learnable interpolation algorithm (e.g., bicubic interpolation). Deep convolutional neural networks (CNNs) then process the coarse HR image solely to reconstruct missing high-frequency details. While this framework handles arbitrary scales and dimensions with ease, performing convolutions in high-dimensional space causes substantial computational and memory overhead and may amplify interpolation artifacts.

    2. Post-upsampling Super-Resolution: Feature extraction and non-linear transformations are performed entirely in the low-dimensional LR feature space. Upsampling is deferred to the final stage of the network using learnable upsampling layers (such as transposed convolutions or sub-pixel layers). This drastically reduces computational complexity and memory usage, making it the most prevalent framework for state-of-the-art models.

    3. Progressive Upsampling Super-Resolution: The network reconstructs higher resolutions progressively in a cascade of stages (e.g., 2×4×8×2\times \to 4\times \to 8\times). At each stage, features or intermediate images are upscaled and refined. This decomposes large-scale SR problems into smaller sub-tasks, reducing learning difficulty and supporting multi-scale outputs within a unified network.

    4. Iterative Up-and-Down Sampling Super-Resolution: Inspired by classical iterative back-projection, this framework alternates between upsampling and downsampling operations across modular blocks to compute reconstruction errors iteratively, feeding intermediate feedback residuals back into the network to refine HR predictions.

  3. Knowl 3 — Learnable and Non-Learnable Upsampling Mechanisms in Super-Resolution

    model/method

    Super-resolution networks rely on distinct upsampling mechanisms to increase spatial feature map dimensions:

    • Interpolation-Based Upsampling: Non-learnable geometric interpolation methods include Nearest-Neighbor, Bilinear (operating over a 2×22 \times 2 pixel neighborhood), and Bicubic (operating over a 4×44 \times 4 pixel neighborhood). They do not introduce new learned information and can introduce computational overhead when placed at network inputs.

    • Transposed Convolution Layer (Deconvolution): Reverses the normal convolution operation by inserting zeros between feature pixels and applying a standard convolution with learnable kernels. While end-to-end differentiable and maintaining standard convolutional connectivity, it frequently causes uneven receptive field overlap, resulting in checkerboard artifacts.

    • Sub-Pixel Convolution Layer (Pixel Shuffle): Applies standard convolutions in low-dimensional space to project an input feature map of size h×w×ch \times w \times c to h×w×(s2c)h \times w \times (s^2 c), where ss is the upsampling scaling factor. A periodic shuffling operator then rearranges the s2cs^2 c channels into an upscaled spatial grid of dimensions sh×sw×csh \times sw \times c. This offers a larger receptive field and avoids zero padding, though boundary artifacts between blocky receptive field regions can still occur.

    • Meta Upscale Module: Solves arbitrary, non-integer, and continuous scaling factor super-resolution within a single model. For each target pixel on the HR grid, it projects coordinates back to the LR feature map, and a multi-layer perceptron (meta-network) dynamically predicts convolution kernel weights based on coordinate offsets and scaling parameters to compute target HR pixels on the fly.

  4. Knowl 4 — Loss Function Formulations in Super-Resolution

    equation

    Super-resolution networks utilize several mathematical loss functions to guide optimization, where I^\hat{I} is the predicted HR image, II is the ground truth HR image, and h,w,ch, w, c denote image height, width, and channels:

    • Pixel L1L_1 Loss: Lpixel_l1(I^,I)=1hwci=1hj=1wk=1cI^i,j,kIi,j,k\mathcal{L}_{\text{pixel\_l1}}(\hat{I}, I) = \frac{1}{hwc} \sum_{i=1}^h \sum_{j=1}^w \sum_{k=1}^c |\hat{I}_{i,j,k} - I_{i,j,k}|

    • Pixel L2L_2 Loss (MSE): Lpixel_l2(I^,I)=1hwci=1hj=1wk=1c(I^i,j,kIi,j,k)2\mathcal{L}_{\text{pixel\_l2}}(\hat{I}, I) = \frac{1}{hwc} \sum_{i=1}^h \sum_{j=1}^w \sum_{k=1}^c (\hat{I}_{i,j,k} - I_{i,j,k})^2

    • Charbonnier Loss (differentiable variant of L1L_1 with stabilizing constant ϵ103\epsilon \approx 10^{-3}): Lpixel_Cha(I^,I)=1hwci=1hj=1wk=1c(I^i,j,kIi,j,k)2+ϵ2\mathcal{L}_{\text{pixel\_Cha}}(\hat{I}, I) = \frac{1}{hwc} \sum_{i=1}^h \sum_{j=1}^w \sum_{k=1}^c \sqrt{(\hat{I}_{i,j,k} - I_{i,j,k})^2 + \epsilon^2}

    • Content Loss (Perceptual Loss): Measures semantic feature Euclidean distance extracted from layer ll of a pre-trained classification network ϕ\phi with dimensions hl,wl,clh_l, w_l, c_l: Lcontent(I^,I;ϕ,l)=1hlwlcli,j,k(ϕi,j,k(l)(I^)ϕi,j,k(l)(I))2\mathcal{L}_{\text{content}}(\hat{I}, I; \phi, l) = \frac{1}{h_l w_l c_l} \sum_{i,j,k} \left( \phi^{(l)}_{i,j,k}(\hat{I}) - \phi^{(l)}_{i,j,k}(I) \right)^2

    • Texture Loss (Style Loss): Computes squared Frobenius distance between Gram matrices G(l)Rcl×clG^{(l)} \in \mathbb{R}^{c_l \times c_l} of feature maps ϕ(l)\phi^{(l)}, where Gi,j(l)(I)=vec(ϕi(l)(I))vec(ϕj(l)(I))G^{(l)}_{i,j}(I) = \text{vec}(\phi^{(l)}_i(I)) \cdot \text{vec}(\phi^{(l)}_j(I)): Ltexture(I^,I;ϕ,l)=1cl2i=1clj=1cl(Gi,j(l)(I^)Gi,j(l)(I))2\mathcal{L}_{\text{texture}}(\hat{I}, I; \phi, l) = \frac{1}{c_l^2} \sum_{i=1}^{c_l} \sum_{j=1}^{c_l} \left( G^{(l)}_{i,j}(\hat{I}) - G^{(l)}_{i,j}(I) \right)^2

    • Adversarial Loss (Cross-Entropy Form) for generator F\mathcal{F} producing I^\hat{I} and discriminator DD with real samples IsI_s: Lgan_ce_g(I^;D)=logD(I^)\mathcal{L}_{\text{gan\_ce\_g}}(\hat{I}; D) = -\log D(\hat{I}) Lgan_ce_d(I^,Is;D)=logD(Is)log(1D(I^))\mathcal{L}_{\text{gan\_ce\_d}}(\hat{I}, I_s; D) = -\log D(I_s) - \log(1 - D(\hat{I}))

    • Total Variation (TV) Loss: Enforces spatial smoothness by summing horizontal and vertical neighbor differences: LTV(I^)=1hwci=1h1j=1w1k=1c(I^i,j+1,kI^i,j,k)2+(I^i+1,j,kI^i,j,k)2\mathcal{L}_{\text{TV}}(\hat{I}) = \frac{1}{hwc} \sum_{i=1}^{h-1} \sum_{j=1}^{w-1} \sum_{k=1}^c \sqrt{(\hat{I}_{i,j+1,k} - \hat{I}_{i,j,k})^2 + (\hat{I}_{i+1,j,k} - \hat{I}_{i,j,k})^2}

    • Cycle Consistency Loss: Ensures reconstructed LR image II' downsampled from I^\hat{I} matches input II: Lcycle(I,I)=1hwci,j,k(Ii,j,kIi,j,k)2\mathcal{L}_{\text{cycle}}(I', I) = \frac{1}{hwc} \sum_{i,j,k} (I'_{i,j,k} - I_{i,j,k})^2

  5. Knowl 5 — Architectural Design Strategies for Super-Resolution Networks

    model/method

    Super-resolution networks employ key modular design paradigms to improve feature extraction, gradient flow, and parameter efficiency:

    • Global and Local Residual Learning: Global residual learning directly adds the input LR image (or its interpolated version) to the output via a global skip connection, requiring the deep network to learn only the sparse high-frequency residual map IyIxI_y - I_x. Local residual learning places skip connections within internal modules (e.g., ResBlocks) to mitigate vanishing gradients in very deep networks.

    • Recursive Learning: Weights of convolutional units are tied and applied iteratively over multiple steps (e.g., 16-recursive DRCN or 25-recursive DRRN). This expands the effective receptive field to large dimensions without increasing parameter count, though computational cost per forward pass remains high.

    • Multi-Path Learning: Processes features simultaneously through diverse architectural branches—such as parallel convolutions with different kernel sizes (3×33 \times 3 and 5×55 \times 5 in MSRN), dual LR/HR information states (DSRN), or scale-specific input/output modules sharing a central trunk (MDSR).

    • Dense Connections: Employs dense blocks where each layer receives concatenated feature maps from all previous layers, generating l(l1)/2l(l-1)/2 connections for an ll-layer block. This enables multi-level feature reuse and allows smaller growth rates (channel counts), reducing parameter sizes (e.g., SRDenseNet, RDN).

    • Channel and Non-Local Attention Mechanisms: Channel attention explicitly models interdependencies across feature channels using global average pooling followed by gating MLPs (RCAN) or second-order feature statistics (SAN/SOCA). Non-local attention calculates pairwise spatial affinities across all pixel locations via self-attention mechanisms to capture long-range structural correlations.

  6. Knowl 6 — Quantitative and Perceptual Image Quality Assessment Metrics

    definition

    Super-resolution performance is evaluated using objective mathematical fidelity metrics, subjective human testing, and learned perceptual measures:

    • Peak Signal-to-Noise Ratio (PSNR): Measures pixel-level reconstruction fidelity using maximum pixel intensity LL (255 for 8-bit images) and mean squared error (MSE) across NN pixels: PSNR=10log10(L21Ni=1N(I(i)I^(i))2)\text{PSNR} = 10 \cdot \log_{10} \left( \frac{L^2}{\frac{1}{N} \sum_{i=1}^N (I(i) - \hat{I}(i))^2} \right) Because PSNR operates strictly on per-pixel differences, maximizing PSNR often produces overly smoothed images lacking high-frequency textures.

    • Structural Similarity Index (SSIM): Evaluates structural degradation based on human visual system (HVS) characteristics, decomposing similarity into luminance ClC_l, contrast CcC_c, and structure CsC_s: SSIM(I,I^)=[Cl(I,I^)]α[Cc(I,I^)]β[Cs(I,I^)]γ\text{SSIM}(I, \hat{I}) = [C_l(I, \hat{I})]^\alpha [C_c(I, \hat{I})]^\beta [C_s(I, \hat{I})]^\gamma where for image means μI,μI^\mu_I, \mu_{\hat{I}}, standard deviations σI,σI^\sigma_I, \sigma_{\hat{I}}, covariance σII^\sigma_{I\hat{I}}, and stability constants C1,C2,C3C_1, C_2, C_3: Cl(I,I^)=2μIμI^+C1μI2+μI^2+C1,Cc(I,I^)=2σIσI^+C2σI2+σI^2+C2,Cs(I,I^)=σII^+C3σIσI^+C3C_l(I, \hat{I}) = \frac{2\mu_I\mu_{\hat{I}} + C_1}{\mu_I^2 + \mu_{\hat{I}}^2 + C_1}, \quad C_c(I, \hat{I}) = \frac{2\sigma_I\sigma_{\hat{I}} + C_2}{\sigma_I^2 + \sigma_{\hat{I}}^2 + C_2}, \quad C_s(I, \hat{I}) = \frac{\sigma_{I\hat{I}} + C_3}{\sigma_I\sigma_{\hat{I}} + C_3}

    • Mean Opinion Score (MOS): Subjective testing where human evaluators score images on an absolute 1 (bad) to 5 (good) scale.

    • Learning-Based and Blind IQA: No-reference metrics (e.g., Ma, NIMA, NIQE) measure deviations from natural image statistical distributions, while deep feature metrics (e.g., LPIPS) assess distance in learned convolutional feature spaces.

  7. Knowl 7 — The Perception-Distortion Trade-off and Network Interpolation

    theoretical result

    Perceptual quality and objective distortion metrics in super-resolution are fundamentally at odds (the perception-distortion tradeoff proven by Blau and Michaeli, 2018). As mathematical distortion (such as MSE or PSNR) decreases, perceptual quality (measured by MOS, NIQE, or Ma) must strictly worsen beyond a certain threshold. Conversely, models trained with adversarial and perceptual losses achieve significantly higher visual realism and texture fidelity but suffer lower PSNR and may introduce hallucinated high-frequency artifacts.

    To balance this trade-off without retraining separate models, network interpolation combines a PSNR-optimized model with parameters θPSNR\theta_{\text{PSNR}} and a GAN-based perceptual model with parameters θGAN\theta_{\text{GAN}} (initialized by fine-tuning the PSNR model):

    θinterp=(1α)θPSNR+αθGAN\theta_{\text{interp}} = (1 - \alpha) \theta_{\text{PSNR}} + \alpha \theta_{\text{GAN}}

    where α[0,1]\alpha \in [0, 1] is an interpolation weight. Adjusting α\alpha enables smooth transitions between distortion minimization and perceptual realism while suppressing unnatural noise artifacts.

  8. Knowl 8 — Unsupervised and Zero-Shot Super-Resolution Paradigms

    model/method

    Supervised SR models trained on synthetic bicubic downsampling perform poorly on real-world images with unknown, complex degradations. Unsupervised and internal-learning methods address this limitation:

    • Zero-Shot Super-Resolution (ZSSR): Exploits cross-scale internal patch recurrence within a single image. At test time, an image-specific blur degradation kernel is estimated from the input LR image, which is used to construct a downscaled internal dataset. A small CNN is trained on the fly on this single-image dataset to perform the test inference, outperforming generic models on non-ideal degradations by 1–2 dB at the cost of high per-image test-time computation.

    • Weakly-Supervised Learned Degradation: A two-stage framework where an HR-to-LR GAN is first trained on unpaired LR and HR datasets to model the true real-world degradation distribution. The trained generator is then used to synthesize realistic LR-HR pairs to train a standard LR-to-HR super-resolution network.

    • Cycle-in-Cycle SR (CinCGAN): Employs nested CycleGAN architectures to map between unpaired domain distributions (e.g., noisy LR \to clean LR \to clean HR) with cycle-consistency losses, bypassing the need for paired training examples.

    • Deep Image Prior (DIP): Leverages the architecture of an un-trained, randomly initialized CNN as an explicit handcrafted prior for low-level image statistics. Optimizing the network weights to reconstruct an HR image whose downscaled version matches the input LR image yields reconstructions exceeding standard bicubic interpolation by ~1 dB without any training datasets.

  9. Knowl 9 — Suboptimality of Batch Normalization in Image Super-Resolution

    limitation

    While Batch Normalization (BN) is standard in high-level computer vision tasks for stabilizing mini-batch statistics and accelerating training, it is sub-optimal for image super-resolution. BN normalizes feature activations across a batch, which strips the network of absolute scale, luminance, and contrast information and restricts feature range flexibility.

    Removing Batch Normalization layers from residual building blocks yields substantial benefits for SR:

    1. It preserves pixel-level intensity and range statistics necessary for exact reconstruction.
    2. It reduces GPU memory consumption by up to approximately 40% during training.
    3. The saved memory allows building substantially deeper and wider networks (e.g., EDSR, RCAN, ESRGAN) under identical hardware constraints, resulting in marked performance gains.
  10. Knowl 10 — Summary of Representative Super-Resolution Architectures and Methodologies

    data/table

    The table below synthesizes the core methodologies, architectural framework placements, upsampling layers, structural components, and loss functions employed across landmark deep learning super-resolution models:

    Method Framework Upsampling Rec. Res. Dense Att. L1L_1 L2L_2 Primary Keywords / Innovations
    SRCNN Pre Bicubic No No No No No Yes First 3-layer CNN for SR
    DRCN Pre Bicubic Yes Yes No No No Yes 16-recursive layers, multi-supervision
    FSRCNN Post Deconv No No No No No Yes Fast post-upsampling, lightweight design
    ESPCN Post Sub-Pixel No No No No No Yes Sub-pixel convolution (pixel shuffle)
    LapSRN Progressive Bicubic No Yes No No Yes No Laplacian pyramid, Charbonnier loss
    DRRN Pre Bicubic Yes Yes No No No Yes 25-recursion residual unit
    SRResNet Post Sub-Pixel No Yes No No No Yes Deep ResNet architecture, MSE loss
    SRGAN Post Sub-Pixel No Yes No No No No Adversarial training, perceptual content loss
    EDSR Post Sub-Pixel No Yes No No Yes No BN removed, large/compact residual model
    EnhanceNet Pre Bicubic No Yes No No No No Automated texture synthesis, Gram loss
    MemNet Pre Bicubic Yes Yes Yes No No Yes Memory blocks, long/short term persistence
    SRDenseNet Post Deconv No No Yes No No Yes Layer-level and block-level dense connections
    DBPN Iterative Deconv Yes No Yes No Yes No Iterative up/down back-projection
    DSRN Pre Deconv Yes Yes No No No Yes Dual-state recurrent LR/HR exchange
    RDN Post Sub-Pixel No Yes Yes No Yes No Residual dense blocks (RDB)
    CARN Post Sub-Pixel Yes Yes Yes No Yes No Cascading residual network for efficiency
    MSRN Post Sub-Pixel No Yes No No Yes No Multi-scale residual blocks (3×33\times 3, 5×55\times 5)
    RCAN Post Sub-Pixel No Yes No Yes Yes No Residual in residual, channel attention
    ESRGAN Post Sub-Pixel No Yes Yes No No No Relativistic GAN, perceptual feature loss
    RNAN Post Sub-Pixel No Yes No Yes Yes No Residual non-local spatial/channel attention
    Meta-RDN Post Meta Upscale No Yes Yes No Yes No Meta learning for arbitrary scaling factors
    SAN Post Sub-Pixel No Yes No Yes Yes No Second-order channel attention (SOCA)
    SRFBN Post Deconv Yes Yes Yes No Yes No Feedback network with iterative refinement

    In this comparison, "Rec.", "Res.", "Dense", and "Att." indicate whether the model incorporates recursive learning, residual connections, dense connections, or attention mechanisms, respectively. "Pre", "Post", "Progressive", and "Iterative" denote the primary upsampling framework classification.

  11. Knowl 11 — Domain-Specific Super-Resolution Methodologies

    model/method

    Specialized computer vision domains require tailored super-resolution formulations:

    • Face Image Super-Resolution (Face Hallucination): Incorporates structured domain priors such as facial landmark heatmaps, semantic parsing maps, component dictionary matching (for eyes, nose, mouth), and identity-preserving loss functions (e.g., Super-FAN, FSRNet, SICNN, LCGE).

    • Depth Map Super-Resolution: Overcomes sensor noise and missing values in low-resolution range measurements by using registered high-resolution RGB guidance images. Multi-scale guidance networks and shape-from-shading variational formulations transfer RGB edge boundaries to guide depth upsampling.

    • Hyperspectral Image (HSI) Super-Resolution: Fuses low-spatial-resolution HSIs containing hundreds of spectral bands with high-spatial-resolution RGB or panchromatic (PAN) images. Models jointly optimize camera spectral response (CSR) functions and enforce spectral angle consistency.

    • Real-World / RAW Image Super-Resolution: Operates directly on uncompressed 12-bit or 14-bit RAW sensor data rather than 8-bit RGB images generated by camera Image Signal Processors (ISPs). This circumvents destructive demosaicing, denoising, and compression artifacts.

    • Video Super-Resolution (VSR): Exploits inter-frame temporal dependencies via explicit optical flow estimation (VESPCN), spatio-temporal deformable convolutions, bidirectional recurrent states (BRCN, FRVSR), or recurrent back-projection networks (RBPN).

  12. Knowl 12 — Inference Self-Ensemble and Multi-Supervision Training Techniques

    model/method

    Super-resolution performance and optimization stability are frequently augmented using two orthogonal training and inference strategies:

    • Self-Ensemble (Enhanced Prediction): An inference-time augmentation technique that transforms a single input LR image IxI_x into eight distinct variations using rotations of 0,90,180,2700^\circ, 90^\circ, 180^\circ, 270^\circ and horizontal flipping. Each transformed image is passed through the trained SR model F\mathcal{F}, the inverse geometric transformations are applied to the resulting HR outputs, and the eight images are aggregated via pixel-wise mean or median to produce the final output I^y\hat{I}_y. This consistently yields modest PSNR gains without architectural modifications.

    • Multi-Supervision: A training technique that introduces auxiliary loss terms at intermediate layers or recursion steps of progressive or recursive networks (e.g., DRCN, LapSRN). Intermediate representations are supervised against ground truth images downsampled to intermediate target resolutions, preventing gradient vanishing/explosion in deep recursive graphs and enforcing coarse-to-fine feature convergence.

Coverage note — Deliberately omitted general introductory historical summaries of non-deep-learning SR methods (e.g., neighbor embedding, sparse representation patch dictionaries) and generic dataset attribute listings from Table 1, focusing entirely on the core contributed taxonomies, modular components, loss equations, theoretical trade-offs, and empirical framework comparisons of deep learning SR.

References

  1. 1.H. Greenspan, “Super-resolution in medical imaging,” The Computer Journal, vol. 52, 2008.
  2. 2.J. S. Isaac and R. Kulkarni, “Super resolution techniques for medical image processing,” in ICTSD, 2015.
  3. 3.Y. Huang, L. Shao, and A. F. Frangi, “Simultaneous super-resolution and cross-modality synthesis of 3d medical images using weakly-supervised joint convolutional sparse coding,” in CVPR, 2017.
  4. 4.L. Zhang, H. Zhang, H. Shen, and P. Li, “A super-resolution reconstruction algorithm for surveillance images,” Elsevier Signal Processing, vol. 90, 2010.
  5. 5.P. Rasti, T. Uiboupin, S. Escalera, and G. Anbarjafari, “Convolutional neural network super resolution for face recognition in surveillance monitoring,” in AMDO, 2016.
  6. 6.D. Dai, Y. Wang, Y. Chen, and L. Van Gool, “Is image super-resolution helpful for other vision tasks?” in WACV, 2016.
  7. 7.M. Haris, G. Shakhnarovich, and N. Ukita, “Task-driven super resolution: Object detection in low-resolution images,” Arxiv:1803.11316, 2018.
  8. 8.M. S. Sajjadi, B. Scholkopf, and M. Hirsch, “Enhancenet: Single ¨ image super-resolution through automated texture synthesis,” in ICCV, 2017.
  9. 9.Y. Zhang, Y. Bai, M. Ding, and B. Ghanem, “Sod-mtgan: Small object detection via multi-task generative adversarial network,” in ECCV, 2018.
  10. 10.R. Keys, “Cubic convolution interpolation for digital image processing,” IEEE Transactions on Acoustics, Speech, and Signal Processing, vol. 29, 1981.
  11. 11.C. E. Duchon, “Lanczos filtering in one and two dimensions,” Journal of Applied Meteorology, vol. 18, 1979.
  12. 12.M. Irani and S. Peleg, “Improving resolution by image registration,” CVGIP: Graphical Models and Image Processing, vol. 53, 1991.
  13. 13.G. Freedman and R. Fattal, “Image and video upscaling from local self-examples,” TOG, vol. 30, 2011.
  14. 14.J. Sun, Z. Xu, and H.-Y. Shum, “Image super-resolution using gradient profile prior,” in CVPR, 2008.
  15. 15.K. I. Kim and Y. Kwon, “Single-image super-resolution using sparse regression and natural image prior,” TPAMI, vol. 32, 2010.
  16. 16.Z. Xiong, X. Sun, and F. Wu, “Robust web image/video super-resolution,” IEEE Transactions on Image Processing, vol. 19, 2010.
  17. 17.W. T. Freeman, T. R. Jones, and E. C. Pasztor, “Example-based super-resolution,” IEEE Computer Graphics and Applications, vol. 22, 2002.
  18. 18.H. Chang, D.-Y. Yeung, and Y. Xiong, “Super-resolution through neighbor embedding,” in CVPR, 2004.
  19. 19.D. Glasner, S. Bagon, and M. Irani, “Super-resolution from a single image,” in ICCV, 2009.
  20. 20.Y. Jianchao, J. Wright, T. Huang, and Y. Ma, “Image super-resolution as sparse representation of raw image patches,” in CVPR, 2008.
  21. 21.J. Yang, J. Wright, T. S. Huang, and Y. Ma, “Image super-resolution via sparse representation,” IEEE Transactions on Image Processing, vol. 19, 2010.
  22. 22.C. Dong, C. C. Loy, K. He, and X. Tang, “Learning a deep convolutional network for image super-resolution,” in ECCV, 2014.
  23. 23.——, “Image super-resolution using deep convolutional networks,” TPAMI, vol. 38, 2016.
  24. 24.I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, “Generative adversarial nets,” in NIPS, 2014.
  25. 25.C. Ledig, L. Theis, F. Huszar, J. Caballero, A. Cunningham, ´ A. Acosta, A. P. Aitken, A. Tejani, J. Totz, Z. Wang et al., “Photorealistic single image super-resolution using a generative adversarial network,” in CVPR, 2017.
  26. 26.J. Kim, J. Kwon Lee, and K. Mu Lee, “Accurate image super-resolution using very deep convolutional networks,” in CVPR, 2016.
  27. 27.W.-S. Lai, J.-B. Huang, N. Ahuja, and M.-H. Yang, “Deep laplacian pyramid networks for fast and accurate superresolution,” in CVPR, 2017.
  28. 28.N. Ahn, B. Kang, and K.-A. Sohn, “Fast, accurate, and lightweight super-resolution with cascading residual network,” in ECCV, 2018.
  29. 29.J. Johnson, A. Alahi, and L. Fei-Fei, “Perceptual losses for realtime style transfer and super-resolution,” in ECCV, 2016.
  30. 30.A. Bulat and G. Tzimiropoulos, “Super-fan: Integrated facial landmark localization and super-resolution of real-world low resolution faces in arbitrary poses with gans,” in CVPR, 2018.
  31. 31.B. Lim, S. Son, H. Kim, S. Nah, and K. M. Lee, “Enhanced deep residual networks for single image super-resolution,” in CVPRW, 2017.
  32. 32.Y. Wang, F. Perazzi, B. McWilliams, A. Sorkine-Hornung, O. Sorkine-Hornung, and C. Schroers, “A fully progressive approach to single-image super-resolution,” in CVPRW, 2018.
  33. 33.S. C. Park, M. K. Park, and M. G. Kang, “Super-resolution image reconstruction: A technical overview,” IEEE Signal Processing Magazine, vol. 20, 2003.
  34. 34.K. Nasrollahi and T. B. Moeslund, “Super-resolution: A comprehensive survey,” Machine Vision and Applications, vol. 25, 2014.
  35. 35.J. Tian and K.-K. Ma, “A survey on super-resolution imaging,” Signal, Image and Video Processing, vol. 5, 2011.
  36. 36.J. Van Ouwerkerk, “Image super-resolution survey,” Image and Vision Computing, vol. 24, 2006.
  37. 37.C.-Y. Yang, C. Ma, and M.-H. Yang, “Single-image super-resolution: A benchmark,” in ECCV, 2014.
  38. 38.D. Thapa, K. Raahemifar, W. R. Bobier, and V. Lakshminarayanan, “A performance comparison among different super-resolution techniques,” Computers & Electrical Engineering, vol. 54, 2016.
  39. 39.K. Zhang, W. Zuo, and L. Zhang, “Learning a single convolutional super-resolution network for multiple degradations,” in CVPR, 2018.
  40. 40.D. Martin, C. Fowlkes, D. Tal, and J. Malik, “A database of human segmented natural images and its application to evaluating segmentation algorithms and measuring ecological statistics,” in ICCV, 2001.
  41. 41.P. Arbelaez, M. Maire, C. Fowlkes, and J. Malik, “Contour detection and hierarchical image segmentation,” TPAMI, vol. 33, 2011.
  42. 42.E. Agustsson and R. Timofte, “Ntire 2017 challenge on single image super-resolution: Dataset and study,” in CVPRW, 2017.
  43. 43.C. Dong, C. C. Loy, and X. Tang, “Accelerating the super-resolution convolutional neural network,” in ECCV, 2016.
  44. 44.R. Timofte, R. Rothe, and L. Van Gool, “Seven ways to improve example-based single image super resolution,” in CVPR, 2016.
  45. 45.A. Fujimoto, T. Ogawa, K. Yamamoto, Y. Matsui, T. Yamasaki, and K. Aizawa, “Manga109 dataset and creation of metadata,” in MANPU, 2016.
  46. 46.X. Wang, K. Yu, C. Dong, and C. C. Loy, “Recovering realistic texture in image super-resolution by deep spatial feature transform,” 2018.
  47. 47.Y. Blau, R. Mechrez, R. Timofte, T. Michaeli, and L. Zelnik-Manor, “2018 pirm challenge on perceptual image super-resolution,” in ECCV Workshop, 2018.
  48. 48.M. Bevilacqua, A. Roumy, C. Guillemot, and M. L. Alberi-Morel, “Low-complexity single-image super-resolution based on non-negative neighbor embedding,” in BMVC, 2012.
  49. 49.R. Zeyde, M. Elad, and M. Protter, “On single image scale-up using sparse-representations,” in International Conference on Curves and Surfaces, 2010.
  50. 50.J.-B. Huang, A. Singh, and N. Ahuja, “Single image super-resolution from transformed self-exemplars,” in CVPR, 2015.
  51. 51.J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in CVPR, 2009.
  52. 52.T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollar, and C. L. Zitnick, “Microsoft coco: Common objects in ´ context,” in ECCV, 2014.
  53. 53.M. Everingham, S. A. Eslami, L. Van Gool, C. K. Williams, J. Winn, and A. Zisserman, “The pascal visual object classes challenge: A retrospective,” IJCV, vol. 111, 2015.
  54. 54.Z. Liu, P. Luo, X. Wang, and X. Tang, “Deep learning face attributes in the wild,” in ICCV, 2015.
  55. 55.Y. Tai, J. Yang, X. Liu, and C. Xu, “Memnet: A persistent memory network for image restoration,” in ICCV, 2017.
  56. 56.Y. Tai, J. Yang, and X. Liu, “Image super-resolution via deep recursive residual network,” in CVPR, 2017.
  57. 57.M. Haris, G. Shakhnarovich, and N. Ukita, “Deep backp-rojection networks for super-resolution,” in CVPR, 2018.
  58. 58.Z. Wang, A. C. Bovik, H. R. Sheikh, and E. P. Simoncelli, “Image quality assessment: From error visibility to structural similarity,” IEEE Transactions on Image Processing, vol. 13, 2004.
  59. 59.Z. Wang, A. C. Bovik, and L. Lu, “Why is image quality assessment so difficult?” in ICASSP, 2002.
  60. 60.H. R. Sheikh, M. F. Sabir, and A. C. Bovik, “A statistical evaluation of recent full reference image quality assessment algorithms,” IEEE Transactions on Image Processing, vol. 15, 2006.
  61. 61.Z. Wang and A. C. Bovik, “Mean squared error: Love it or leave it? a new look at signal fidelity measures,” IEEE Signal Processing Magazine, vol. 26, 2009.
  62. 62.Z. Wang, D. Liu, J. Yang, W. Han, and T. Huang, “Deep networks for image super-resolution with sparse prior,” in ICCV, 2015.
  63. 63.X. Xu, D. Sun, J. Pan, Y. Zhang, H. Pfister, and M.-H. Yang, “Learning to super-resolve blurry face and text images,” in ICCV, 2017.
  64. 64.R. Dahl, M. Norouzi, and J. Shlens, “Pixel recursive super resolution,” in ICCV, 2017.
  65. 65.W.-S. Lai, J.-B. Huang, N. Ahuja, and M.-H. Yang, “Fast and accurate image super-resolution with deep laplacian pyramid networks,” TPAMI, 2018.
  66. 66.C. Ma, C.-Y. Yang, X. Yang, and M.-H. Yang, “Learning a no-reference quality metric for single-image super-resolution,” Computer Vision and Image Understanding, 2017.
  67. 67.H. Talebi and P. Milanfar, “Nima: Neural image assessment,” IEEE Transactions on Image Processing, vol. 27, 2018.
  68. 68.J. Kim and S. Lee, “Deep learning of human visual sensitivity in image quality assessment framework,” in CVPR, 2017.
  69. 69.R. Zhang, P. Isola, A. A. Efros, E. Shechtman, and O. Wang, “The unreasonable effectiveness of deep features as a perceptual metric,” in CVPR, 2018.
  70. 70.Y. Zhang, K. Li, K. Li, L. Wang, B. Zhong, and Y. Fu, “Image super-resolution using very deep residual channel attention networks,” in ECCV, 2018.
  71. 71.C. Fookes, F. Lin, V. Chandran, and S. Sridharan, “Evaluation of image resolution and super-resolution on face recognition performance,” Journal of Visual Communication and Image Representation, vol. 23, 2012.
  72. 72.K. Zhang, Z. ZHANG, C.-W. Cheng, W. Hsu, Y. Qiao, W. Liu, and T. Zhang, “Super-identity convolutional neural network for face hallucination,” in ECCV, 2018.
  73. 73.Y. Chen, Y. Tai, X. Liu, C. Shen, and J. Yang, “Fsrnet: End-to-end learning face super-resolution with facial priors,” in CVPR, 2018.
  74. 74.Z. Wang, E. Simoncelli, A. Bovik et al., “Multi-scale structural similarity for image quality assessment,” in Asilomar Conference on Signals, Systems, and Computers, 2003.
  75. 75.L. Zhang, L. Zhang, X. Mou, D. Zhang et al., “Fsim: a feature similarity index for image quality assessment,” IEEE transactions on Image Processing, vol. 20, 2011.
  76. 76.A. Mittal, R. Soundararajan, and A. C. Bovik, “Making a completely blind image quality analyzer,” IEEE Signal Processing Letters, 2013.
  77. 77.Y. Blau and T. Michaeli, “The perception-distortion tradeoff,” in CVPR, 2018.
  78. 78.X. Mao, C. Shen, and Y.-B. Yang, “Image restoration using very deep convolutional encoder-decoder networks with symmetric skip connections,” in NIPS, 2016.
  79. 79.T. Tong, G. Li, X. Liu, and Q. Gao, “Image super-resolution using dense skip connections,” in ICCV, 2017.
  80. 80.R. Timofte, E. Agustsson, L. Van Gool, M.-H. Yang, L. Zhang, B. Lim, S. Son, H. Kim, S. Nah, K. M. Lee et al., “Ntire 2017 challenge on single image super-resolution: Methods and results,” in CVPRW, 2017.
  81. 81.A. Ignatov, R. Timofte, T. Van Vu, T. Minh Luu, T. X Pham, C. Van Nguyen, Y. Kim, J.-S. Choi, M. Kim, J. Huang et al., “Pirm challenge on perceptual image enhancement on smartphones: report,” in ECCV Workshop, 2018.
  82. 82.J. Kim, J. Kwon Lee, and K. Mu Lee, “Deeply-recursive convolutional network for image super-resolution,” in CVPR, 2016.
  83. 83.A. Shocher, N. Cohen, and M. Irani, “zero-shot super-resolution using deep internal learning,” in CVPR, 2018.
  84. 84.W. Shi, J. Caballero, F. Huszar, J. Totz, A. P. Aitken, R. Bishop, ´ D. Rueckert, and Z. Wang, “Real-time single image and video super-resolution using an efficient sub-pixel convolutional neural network,” in CVPR, 2016.
  85. 85.W. Han, S. Chang, D. Liu, M. Yu, M. Witbrock, and T. S. Huang, “Image super-resolution via dual-state recurrent networks,” in CVPR, 2018.
  86. 86.Z. Li, J. Yang, Z. Liu, X. Yang, G. Jeon, and W. Wu, “Feedback network for image super-resolution,” in CVPR, 2019.
  87. 87.M. Haris, G. Shakhnarovich, and N. Ukita, “Recurrent back-projection network for video super-resolution,” in CVPR, 2019.
  88. 88.R. Timofte, V. De Smet, and L. Van Gool, “A+: Adjusted anchored neighborhood regression for fast super-resolution,” in ACCV, 2014.
  89. 89.S. Schulter, C. Leistner, and H. Bischof, “Fast and accurate image upscaling with super-resolution forests,” in CVPR, 2015.
  90. 90.M. D. Zeiler, D. Krishnan, G. W. Taylor, and R. Fergus, “Deconvolutional networks,” in CVPRW, 2010.
  91. 91.M. D. Zeiler and R. Fergus, “Visualizing and understanding convolutional networks,” in ECCV, 2014.
  92. 92.A. Odena, V. Dumoulin, and C. Olah, “Deconvolution and checkerboard artifacts,” Distill, 2016.
  93. 93.Y. Zhang, Y. Tian, Y. Kong, B. Zhong, and Y. Fu, “Residual dense network for image super-resolution,” in CVPR, 2018.
  94. 94.H. Gao, H. Yuan, Z. Wang, and S. Ji, “Pixel transposed convolutional networks,” TPAMI, 2019.
  95. 95.X. Hu, H. Mu, X. Zhang, Z. Wang, T. Tan, and J. Sun, “Meta-sr: A magnification-arbitrary network for super-resolution,” in CVPR, 2019.
  96. 96.K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in CVPR, 2016.
  97. 97.R. Timofte, V. De Smet, and L. Van Gool, “Anchored neighborhood regression for fast example-based super-resolution,” in ICCV, 2013.
  98. 98.Z. Hui, X. Wang, and X. Gao, “Fast and accurate single image super-resolution via information distillation network,” in CVPR, 2018.
  99. 99.J. Li, F. Fang, K. Mei, and G. Zhang, “Multi-scale residual network for image super-resolution,” in ECCV, 2018.
  100. 100.H. Ren, M. El-Khamy, and J. Lee, “Image super resolution based on fusing multiple convolution neural networks,” in CVPRW, 2017.
  101. 101.C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich, “Going deeper with convolutions,” in CVPR, 2015.
  102. 102.G. Huang, Z. Liu, L. Van Der Maaten, and K. Q. Weinberger, “Densely connected convolutional networks,” in CVPR, 2017.
  103. 103.X. Wang, K. Yu, S. Wu, J. Gu, Y. Liu, C. Dong, C. C. Loy, Y. Qiao, and X. Tang, “Esrgan: Enhanced super-resolution generative adversarial networks,” in ECCV Workshop, 2018.
  104. 104.J. Hu, L. Shen, and G. Sun, “Squeeze-and-excitation networks,” in CVPR, 2018.
  105. 105.T. Dai, J. Cai, Y. Zhang, S.-T. Xia, and L. Zhang, “Second-order attention network for single image super-resolution,” in CVPR, 2019.
  106. 106.Y. Zhang, K. Li, K. Li, B. Zhong, and Y. Fu, “Residual non-local attention networks for image restoration,” ICLR, 2019.
  107. 107.K. Zhang, W. Zuo, S. Gu, and L. Zhang, “Learning deep cnn denoiser prior for image restoration,” in CVPR, 2017.
  108. 108.S. Xie, R. Girshick, P. Dollar, Z. Tu, and K. He, “Aggregated ´ residual transformations for deep neural networks,” in CVPR, 2017.
  109. 109.F. Chollet, “Xception: Deep learning with depthwise separable convolutions,” in CVPR, 2017.
  110. 110.A. G. Howard, M. Zhu, B. Chen, D. Kalenichenko, W. Wang, T. Weyand, M. Andreetto, and H. Adam, “Mobilenets: Efficient convolutional neural networks for mobile vision applications,” Arxiv:1704.04861, 2017.
  111. 111.A. van den Oord, N. Kalchbrenner, L. Espeholt, O. Vinyals, A. Graves et al., “Conditional image generation with pixelcnn decoders,” in NIPS, 2016.
  112. 112.J. Najemnik and W. S. Geisler, “Optimal eye movement strategies in visual search,” Nature, vol. 434, 2005.
  113. 113.Q. Cao, L. Lin, Y. Shi, X. Liang, and G. Li, “Attention-aware face hallucination via deep reinforcement learning,” in CVPR, 2017.
  114. 114.K. He, X. Zhang, S. Ren, and J. Sun, “Spatial pyramid pooling in deep convolutional networks for visual recognition,” in ECCV, 2014.
  115. 115.H. Zhao, J. Shi, X. Qi, X. Wang, and J. Jia, “Pyramid scene parsing network,” in CVPR, 2017.
  116. 116.D. Park, K. Kim, and S. Y. Chun, “Efficient module based single image super resolution for multiple problems,” in CVPRW, 2018.
  117. 117.I. Daubechies, Ten lectures on wavelets. SIAM, 1992.
  118. 118.S. Mallat, A wavelet tour of signal processing. Elsevier, 1999.
  119. 119.W. Bae, J. J. Yoo, and J. C. Ye, “Beyond deep residual learning for image restoration: Persistent homology-guided manifold simplification,” in CVPRW, 2017.
  120. 120.T. Guo, H. S. Mousavi, T. H. Vu, and V. Monga, “Deep wavelet prediction for image super-resolution,” in CVPRW, 2017.
  121. 121.H. Huang, R. He, Z. Sun, T. Tan et al., “Wavelet-srnet: A wavelet-based cnn for multi-scale face super resolution,” in ICCV, 2017.
  122. 122.P. Liu, H. Zhang, K. Zhang, L. Lin, and W. Zuo, “Multi-level wavelet-cnn for image restoration,” in CVPRW, 2018.
  123. 123.T. Vu, C. Van Nguyen, T. X. Pham, T. M. Luu, and C. D. Yoo, “Fast and efficient image quality enhancement via desubpixel convolutional neural networks,” in ECCV Workshop, 2018.
  124. 124.I. Kligvasser, T. Rott Shaham, and T. Michaeli, “xunit: Learning a spatial activation function for efficient image restoration,” in CVPR, 2018.
  125. 125.A. Bruhn, J. Weickert, and C. Schnorr, “Lucas/kanade meets ¨ horn/schunck: Combining local and global optic flow methods,” IJCV, vol. 61, 2005.
  126. 126.H. Zhao, O. Gallo, I. Frosio, and J. Kautz, “Loss functions for image restoration with neural networks,” IEEE Transactions on Computational Imaging, vol. 3, 2017.
  127. 127.A. Dosovitskiy and T. Brox, “Generating images with perceptual similarity metrics based on deep networks,” in NIPS, 2016.
  128. 128.K. Simonyan and A. Zisserman, “Very deep convolutional networks for large-scale image recognition,” in ICLR, 2015.
  129. 129.L. Gatys, A. S. Ecker, and M. Bethge, “Texture synthesis using convolutional neural networks,” in NIPS, 2015.
  130. 130.L. A. Gatys, A. S. Ecker, and M. Bethge, “Image style transfer using convolutional neural networks,” in CVPR, 2016.
  131. 131.Y. Yuan, S. Liu, J. Zhang, Y. Zhang, C. Dong, and L. Lin, “Unsupervised image super-resolution using cycle-in-cycle generative adversarial networks,” in CVPRW, 2018.
  132. 132.X. Mao, Q. Li, H. Xie, R. Y. Lau, Z. Wang, and S. P. Smolley, “Least squares generative adversarial networks,” in ICCV, 2017.
  133. 133.S.-J. Park, H. Son, S. Cho, K.-S. Hong, and S. Lee, “Srfeat: Single image super resolution with feature discrimination,” in ECCV, 2018.
  134. 134.A. Jolicoeur-Martineau, “The relativistic discriminator: a key element missing from standard gan,” Arxiv:1807.00734, 2018.
  135. 135.M. Arjovsky, S. Chintala, and L. Bottou, “Wasserstein generative adversarial networks,” in ICML, 2017.
  136. 136.I. Gulrajani, F. Ahmed, M. Arjovsky, V. Dumoulin, and A. C. Courville, “Improved training of wasserstein gans,” in NIPS, 2017.
  137. 137.T. Miyato, T. Kataoka, M. Koyama, and Y. Yoshida, “Spectral normalization for generative adversarial networks,” in ICLR, 2018.
  138. 138.J.-Y. Zhu, T. Park, P. Isola, and A. A. Efros, “Unpaired image-to-image translation using cycle-consistent adversarial networks,” in ICCV, 2017.
  139. 139.L. I. Rudin, S. Osher, and E. Fatemi, “Nonlinear total variation based noise removal algorithms,” Physica D: Nonlinear Phenomena, vol. 60, 1992.
  140. 140.H. A. Aly and E. Dubois, “Image up-sampling using total-variation regularization with a new observation model,” IEEE Transactions on Image Processing, vol. 14, 2005.
  141. 141.Y. Guo, Q. Chen, J. Chen, J. Huang, Y. Xu, J. Cao, P. Zhao, and M. Tan, “Dual reconstruction nets for image super-resolution with gradient sensitive loss,” arXiv:1809.07099, 2018.
  142. 142.S. Vasu, N. T. Madam et al., “Analyzing perception-distortion tradeoff using enhanced perceptual super-resolution network,” in ECCV Workshop, 2018.
  143. 143.M. Cheon, J.-H. Kim, J.-H. Choi, and J.-S. Lee, “Generative adversarial network-based image super-resolution using perceptual content losses,” in ECCV Workshop, 2018.
  144. 144.J.-H. Choi, J.-H. Kim, M. Cheon, and J.-S. Lee, “Deep learning-based image super-resolution considering quantitative and perceptual quality,” in ECCV Workshop, 2018.
  145. 145.I. Sergey and S. Christian, “Batch normalization: Accelerating deep network training by reducing internal covariate shift,” in ICML, 2015.
  146. 146.C. K. Sønderby, J. Caballero, L. Theis, W. Shi, and F. Huszar, ´ “Amortised map inference for image super-resolution,” in ICLR, 2017.
  147. 147.R. Chen, Y. Qu, K. Zeng, J. Guo, C. Li, and Y. Xie, “Persistent memory residual network for single image super resolution,” in CVPRW, 2018.
  148. 148.Y. Bengio, J. Louradour, R. Collobert, and J. Weston, “Curriculum learning,” in ICML, 2009.
  149. 149.Y. Bei, A. Damian, S. Hu, S. Menon, N. Ravi, and C. Rudin, “New techniques for preserving global structure and denoising with low information loss in single-image super-resolution,” in CVPRW, 2018.
  150. 150.N. Ahn, B. Kang, and K.-A. Sohn, “Image super-resolution via progressive cascading residual network,” in CVPRW, 2018.
  151. 151.T. Karras, T. Aila, S. Laine, and J. Lehtinen, “Progressive growing of gans for improved quality, stability, and variation,” in ICLR, 2018.
  152. 152.R. Caruana, “Multitask learning,” Machine Learning, vol. 28, 1997.
  153. 153.K. He, G. Gkioxari, P. Dollar, and R. Girshick, “Mask r-cnn,” in ´ ICCV, 2017.
  154. 154.Z. Zhang, P. Luo, C. C. Loy, and X. Tang, “Facial landmark detection by deep multi-task learning,” in ECCV, 2014.
  155. 155.X. Wang, K. Yu, C. Dong, X. Tang, and C. C. Loy, “Deep network interpolation for continuous imagery effect transition,” in CVPR, 2019.
  156. 156.J. Caballero, C. Ledig, A. P. Aitken, A. Acosta, J. Totz, Z. Wang, and W. Shi, “Real-time video super-resolution with spatio-temporal networks and motion compensation,” in CVPR, 2017.
  157. 157.L. Zhu, “pytorch-opcounter,” https://github.com/Lyken17/pytorch-OpCounter, 2019.
  158. 158.T. Michaeli and M. Irani, “Nonparametric blind super-resolution,” in ICCV, 2013.
  159. 159.A. Bulat, J. Yang, and G. Tzimiropoulos, “To learn image super-resolution, use a gan to learn how to do image degradation first,” in ECCV, 2018.
  160. 160.D. Ulyanov, A. Vedaldi, and V. Lempitsky, “Deep image prior,” in CVPR, 2018.
  161. 161.J. Shotton, A. Fitzgibbon, M. Cook, T. Sharp, M. Finocchio, R. Moore, A. Kipman, and A. Blake, “Real-time human pose recognition in parts from single depth images,” in CVPR, 2011.
  162. 162.G. Moon, J. Yong Chang, and K. Mu Lee, “V2v-posenet: Voxel-to-voxel prediction network for accurate 3d hand and human pose estimation from a single depth map,” in CVPR, 2018.
  163. 163.S. Gupta, R. Girshick, P. Arbelaez, and J. Malik, “Learning rich ´ features from rgb-d images for object detection and segmentation,” in ECCV, 2014.
  164. 164.W. Wang and U. Neumann, “Depth-aware cnn for rgb-d segmentation,” in ECCV, 2018.
  165. 165.X. Song, Y. Dai, and X. Qin, “Deep depth super-resolution: Learning depth super-resolution using deep convolutional neural network,” in ACCV, 2016.
  166. 166.T.-W. Hui, C. C. Loy, and X. Tang, “Depth map super-resolution by deep multi-scale guidance,” in ECCV, 2016.
  167. 167.B. Haefner, Y. Queau, T. M ´ ollenhoff, and D. Cremers, “Fight ¨ ill-posedness with ill-posedness: Single-shot variational depth super-resolution from shading,” in CVPR, 2018.
  168. 168.G. Riegler, M. Ruther, and H. Bischof, “Atgv-net: Accurate depth ¨ super-resolution,” in ECCV, 2016.
  169. 169.J.-S. Park and S.-W. Lee, “An example-based face hallucination method for single-frame, low-resolution facial images,” IEEE Transactions on Image Processing, vol. 17, 2008.
  170. 170.S. Zhu, S. Liu, C. C. Loy, and X. Tang, “Deep cascaded bi-network for face hallucination,” in ECCV, 2016.
  171. 171.X. Yu, B. Fernando, B. Ghanem, F. Porikli, and R. Hartley, “Face super-resolution guided by facial component heatmaps,” in ECCV, 2018.
  172. 172.X. Yu and F. Porikli, “Face hallucination with tiny unaligned images by transformative discriminative neural networks,” in AAAI, 2017.
  173. 173.M. Jaderberg, K. Simonyan, A. Zisserman et al., “Spatial transformer networks,” in NIPS, 2015.
  174. 174.X. Yu and F. Porikli, “Hallucinating very low-resolution unaligned and noisy face images by transformative discriminative autoencoders,” in CVPR, 2017.
  175. 175.Y. Song, J. Zhang, S. He, L. Bao, and Q. Yang, “Learning to hallucinate face images via component generation and enhancement,” in IJCAI, 2017.
  176. 176.C.-Y. Yang, S. Liu, and M.-H. Yang, “Hallucinating compressed face images,” IJCV, vol. 126, 2018.
  177. 177.X. Yu and F. Porikli, “Ultra-resolving face images by discriminative generative networks,” in ECCV, 2016.
  178. 178.C.-H. Lee, K. Zhang, H.-C. Lee, C.-W. Cheng, and W. Hsu, “Attribute augmented convolutional neural network for face hallucination,” in CVPRW, 2018.
  179. 179.X. Yu, B. Fernando, R. Hartley, and F. Porikli, “Super-resolving very low-resolution face images with supplementary attributes,” in CVPR, 2018.
  180. 180.M. Mirza and S. Osindero, “Conditional generative adversarial nets,” Arxiv:1411.1784, 2014.
  181. 181.M. Fauvel, Y. Tarabalka, J. A. Benediktsson, J. Chanussot, and J. C. Tilton, “Advances in spectral-spatial classification of hyperspectral images,” Proceedings of the IEEE, vol. 101, 2013.
  182. 182.Y. Fu, Y. Zheng, I. Sato, and Y. Sato, “Exploiting spectral-spatial correlation for coded hyperspectral image restoration,” in CVPR, 2016.
  183. 183.B. Uzkent, A. Rangnekar, and M. J. Hoffman, “Aerial vehicle tracking by adaptive fusion of hyperspectral likelihood maps,” in CVPRW, 2017.
  184. 184.G. Masi, D. Cozzolino, L. Verdoliva, and G. Scarpa, “Pansharpening by convolutional neural networks,” Remote Sensing, vol. 8, 2016.
  185. 185.Y. Qu, H. Qi, and C. Kwan, “Unsupervised sparse dirichlet-net for hyperspectral image super-resolution,” in CVPR, 2018.
  186. 186.Y. Fu, T. Zhang, Y. Zheng, D. Zhang, and H. Huang, “Hyperspectral image super-resolution with optimized rgb guidance,” in CVPR, 2019.
  187. 187.C. Chen, Z. Xiong, X. Tian, Z.-J. Zha, and F. Wu, “Camera lens super-resolution,” in CVPR, 2019.
  188. 188.X. Zhang, Q. Chen, R. Ng, and V. Koltun, “Zoom to learn, learn to zoom,” in CVPR, 2019.
  189. 189.X. Xu, Y. Ma, and W. Sun, “Towards real scene super-resolution with raw images,” in CVPR, 2019.
  190. 190.R. Liao, X. Tao, R. Li, Z. Ma, and J. Jia, “Video super-resolution via deep draft-ensemble learning,” in ICCV, 2015.
  191. 191.A. Kappeler, S. Yoo, Q. Dai, and A. K. Katsaggelos, “Video super-resolution with convolutional neural networks,” IEEE Transactions on Computational Imaging, vol. 2, 2016.
  192. 192.——, “Super-resolution of compressed videos using convolutional neural networks,” in ICIP, 2016.
  193. 193.M. Drulea and S. Nedevschi, “Total variation regularization of local-global optical flow,” in ITSC, 2011.
  194. 194.D. Liu, Z. Wang, Y. Fan, X. Liu, Z. Wang, S. Chang, and T. Huang, “Robust video super-resolution with learned temporal dynamics,” in ICCV, 2017.
  195. 195.D. Liu, Z. Wang, Y. Fan, X. Liu, Z. Wang, S. Chang, X. Wang, and T. S. Huang, “Learning temporal dynamics for video super-resolution: A deep learning approach,” IEEE Transactions on Image Processing, vol. 27, 2018.
  196. 196.X. Tao, H. Gao, R. Liao, J. Wang, and J. Jia, “Detail-revealing deep video super-resolution,” in ICCV, 2017.
  197. 197.Y. Huang, W. Wang, and L. Wang, “Bidirectional recurrent convolutional networks for multi-frame super-resolution,” in NIPS, 2015.
  198. 198.——, “Video super-resolution via bidirectional recurrent convolutional networks,” TPAMI, vol. 40, 2018.
  199. 199.J. Guo and H. Chao, “Building an end-to-end spatial-temporal convolutional network for video super-resolution,” in AAAI, 2017.
  200. 200.A. Graves, S. Fernandez, and J. Schmidhuber, “Bidirectional lstm ´ networks for improved phoneme classification and recognition,” in ICANN, 2005.
  201. 201.M. S. Sajjadi, R. Vemulapalli, and M. Brown, “Frame-recurrent video super-resolution,” in CVPR, 2018.
  202. 202.S. Li, F. He, B. Du, L. Zhang, Y. Xu, and D. Tao, “Fast spatio-temporal residual network for video super-resolution,” in CVPR, 2019.
  203. 203.Z. Zhang and V. Sze, “Fast: A framework to accelerate super-resolution processing on compressed videos,” in CVPRW, 2017.
  204. 204.Y. Jo, S. W. Oh, J. Kang, and S. J. Kim, “Deep video super-resolution network using dynamic upsampling filters without explicit motion compensation,” in CVPR, 2018.
  205. 205.J. Li, X. Liang, Y. Wei, T. Xu, J. Feng, and S. Yan, “Perceptual generative adversarial networks for small object detection,” in CVPR, 2017.
  206. 206.W. Tan, B. Yan, and B. Bare, “Feature super-resolution: Make machine see more clearly,” in CVPR, 2018.
  207. 207.D. S. Jeon, S.-H. Baek, I. Choi, and M. H. Kim, “Enhancing the spatial resolution of stereo images using a parallax prior,” in CVPR, 2018.
  208. 208.L. Wang, Y. Wang, Z. Liang, Z. Lin, J. Yang, W. An, and Y. Guo, “Learning parallax attention for stereo image super-resolution,” in CVPR, 2019.
  209. 209.Y. Li, V. Tsiminaki, R. Timofte, M. Pollefeys, and L. V. Gool, “3d appearance super-resolution with deep learning,” in CVPR, 2019.
  210. 210.S. Zhang, Y. Lin, and H. Sheng, “Residual networks for light field image super-resolution,” in CVPR, 2019.
  211. 211.C. Ancuti, C. O. Ancuti, R. Timofte, L. Van Gool, L. Zhang, M.-H. Yang, V. M. Patel, H. Zhang, V. A. Sindagi, R. Zhao et al., “Ntire 2018 challenge on image dehazing: Methods and results,” in CVPRW, 2018.
  212. 212.H. Pham, M. Y. Guan, B. Zoph, Q. V. Le, and J. Dean, “Efficient neural architecture search via parameter sharing,” in ICML, 2018.
  213. 213.H. Liu, K. Simonyan, and Y. Yang, “Darts: Differentiable architecture search,” ICLR, 2019.
  214. 214.Y. Guo, Y. Zheng, M. Tan, Q. Chen, J. Chen, P. Zhao, and J. Huang, “Nat: Neural architecture transformer for accurate and compact architectures,” in NIPS, 2019, pp. 735–747.

Citation

MLA
Wang, Z., et al. “Deep Learning for Image Super-resolution: A Survey”. arXiv, 2019, http://arxiv.org/abs/1902.06068v2.
APA
Wang, Z., Chen, J., & Hoi, S. C. H. (2019). Deep Learning for Image Super-resolution: A Survey. arXiv. http://arxiv.org/abs/1902.06068v2
Chicago
Wang, Z., J. Chen, and S. C. H. Hoi. 2019. “Deep Learning for Image Super-resolution: A Survey”. arXiv. http://arxiv.org/abs/1902.06068v2.
Harvard
Wang, Z., Chen, J. and Hoi, S.C.H. (2019) “Deep Learning for Image Super-resolution: A Survey”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1902.06068v2.
Vancouver
1. Wang Z, Chen J, Hoi SCH (2019) Deep Learning for Image Super-resolution: A Survey. arXiv

BibTeX

@article{wang2019deep,
  title = {Deep Learning for Image Super-resolution: A Survey},
  author = {Wang, Zhihao and Chen, Jian and Hoi, Steven C. H.},
  year = {2019},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1902.06068v2},
  eprint = {1902.06068}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF