Second-Order Attention Network for Single Image Super-Resolution

Tao DaiJianrui CaiYongbing ZhangShutao XiaLei Zhang

article2019CVPR1,840 citations

Proposes a second-order attention network that exploits higher-order feature statistics and non-local spatial context to capture inter-channel dependencies, achieving superior image super-resolution performance over standard deep CNNs.

Listen

Single image super-resolution—the process of reconstructing high-resolution images from low-resolution inputs—is a fundamental challenge in digital imaging and computer vision. While deep convolutional neural networks have achieved strong results in this area, existing approaches have largely relied on building deeper or wider networks. In doing so, they have neglected intermediate feature correlations and failed to fully exploit information directly available in the original low-resolution inputs, constraining overall image restoration quality.

The article evaluates whether incorporating second-order feature statistics and regional spatial relationships into a deep network can improve reconstruction accuracy and visual quality. To demonstrate this, the authors develop a novel architecture termed the Second-order Attention Network (SAN).

The approach introduces a second-order channel attention mechanism that adaptively rescales feature channels by computing covariance statistics rather than traditional first-order averages. To ensure computational efficiency during graphics processor training, the method incorporates a fast iterative matrix normalization technique. The architecture also introduces a non-locally enhanced residual group structure that uses regional non-local operations to capture wide spatial context, paired with shared skip connections that bypass low-frequency information from the low-resolution input. The authors trained the network on 800 high-resolution images from the standard DIV2K dataset and evaluated performance across five standard benchmark datasets under bicubic and blur-downscale degradation models across multiple scaling factors.

The experimental findings show that the proposed network consistently outperforms state-of-the-art super-resolution methods across the evaluated benchmarks. Ablation experiments demonstrate that second-order channel attention delivers measurable accuracy improvements over first-order alternatives, and the regional non-local design effectively boosts reconstruction performance. In visual assessments, the model produces sharper structural boundaries and restores complex high-frequency details with fewer blurring artifacts compared to existing techniques. Furthermore, the model achieves these gains while requiring approximately 15.7 million parameters, fewer than comparable leading models such as RDN (22.3 million) and EDSR (43 million).

These results demonstrate that capturing higher-order feature statistics and spatial context offers a more parameter-efficient path to enhancing image quality than simply expanding network depth. For organizations deploying imaging pipelines, this approach reduces model size trade-offs while improving visual fidelity, particularly in texture-rich scenes. However, while the model excels on high-order patterns like textures, performance gains were less pronounced on simple geometric edges compared to select channel-attention models. Decision-makers should evaluate this architecture for tasks requiring fine detail recovery, while validating performance against specific real-world operational degradation types.

No sufficiently relevant recommendations were found.

Cover for Second-Order Attention Network for Single Image Super-Resolution

Abstract

Recently, deep convolutional neural networks (CNNs) have been widely explored in single image super-resolution (SISR) and obtained remarkable performance. However, most of the existing CNN-based SISR methods mainly focus on wider or deeper architecture design, neglecting to explore the feature correlations of intermediate layers, hence hindering the representational power of CNNs. To address this issue, in this paper, we propose a second-order attention network (SAN) for more powerful feature expression and feature correlation learning. Specifically, a novel trainable second-order channel attention (SOCA) module is developed to adaptively rescale the channel-wise features by using second-order feature statistics for more discriminative representations. Furthermore, we present a non-locally enhanced residual group (NLRG) structure, which not only incorporates non-local operations to capture long-distance spatial contextual information, but also contains repeated local-source residual attention groups (LSRAG) to learn increasingly abstract feature representations. Experimental results demonstrate the superiority of our SAN network over state-of-the-art SISR methods in terms of both quantitative metrics and visual quality.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Second-order Attention Network (SAN)
  • 3.1. Network Framework
  • 3.2. Non-locally Enhanced Residual Group (NLRG)
  • 3.3. Second-order Channel Attention (SOCA)
  • 3.4. Covariance Normalization Acceleration
  • 3.5. Implementations
  • 3.6. Discussions
  • 4. Experiments
  • 4.1. Setup
  • 4.2. Ablation Study
  • 4.3. Results with Bicubic Degradation (BI)
  • 4.4. Results with Blur-downscale Degradation (BD)
  • 4.5. Model Size Analyses
  • 5. Conclusions
  • References

Knowls

  1. Knowl 1 — Second-order Attention Network (SAN) Architecture for SISR

    model/method

    The Second-order Attention Network (SAN) is an end-to-end deep convolutional neural network for single image super-resolution (SISR) consisting of four functional stages: shallow feature extraction, deep feature extraction via a Non-locally Enhanced Residual Group (NLRG), an upscale module, and an image reconstruction layer.

    Given a low-resolution input image ILRI_{LR}, a single convolutional layer HSF(⋅)H_{SF}(\cdot) with C=64C = 64 filters of size 3×33 \times 3 extracts an initial shallow feature map: F0=HSF(ILR)F_0 = H_{SF}(I_{LR})

    The shallow feature F0F_0 is fed into the NLRG module HNLRG(⋅)H_{NLRG}(\cdot) to extract high-level discriminative deep representations: FDF=HNLRG(F0)F_{DF} = H_{NLRG}(F_0)

    The extracted deep feature FDFF_{DF} is upscaled to the target spatial resolution using an efficient sub-pixel convolution module H↑(⋅)H_{\uparrow}(\cdot) (ESPCNN): F↑=H↑(FDF)F_{\uparrow} = H_{\uparrow}(F_{DF})

    Finally, a reconstruction convolutional layer HR(⋅)H_R(\cdot) with 3 filters of size 3×33 \times 3 transforms the upscaled feature into the final super-resolved RGB image ISR=HSAN(ILR)I_{SR} = H_{SAN}(I_{LR}): ISR=HR(F↑)=HSAN(ILR)I_{SR} = H_R(F_{\uparrow}) = H_{SAN}(I_{LR})

    The network parameters Θ\Theta are optimized across NN training pairs {ILRi,IHRi}i=1N\{I_{LR}^i, I_{HR}^i\}_{i=1}^N using the L1L_1 loss function: L(Θ)=1N∑i=1N∥HSAN(ILRi)−IHRi∥1\mathcal{L}(\Theta) = \frac{1}{N} \sum_{i=1}^N \|H_{SAN}(I_{LR}^i) - I_{HR}^i\|_1

    Training is performed with the ADAM optimizer (β1=0.9,β2=0.99,ϵ=10−8\beta_1 = 0.9, \beta_2 = 0.99, \epsilon = 10^{-8}) using minibatches of 8 color patches of size 48×4848 \times 48, an initial learning rate of 10−410^{-4} reduced by half every 200 epochs, and data augmentation via random rotations (90∘,180∘,270∘90^\circ, 180^\circ, 270^\circ) and horizontal flips.

  2. Knowl 2 — Second-order Channel Attention (SOCA) Mechanism

    model/method

    The Second-Order Channel Attention (SOCA) module adaptively rescales channel-wise feature responses by exploiting second-order feature statistics rather than first-order global pooling, enhancing the network's discriminative representation of high-frequency textures.

    Given an intermediate feature map F=[f1,f2,…,fC]∈RH×W×CF = [f_1, f_2, \dots, f_C] \in \mathbb{R}^{H \times W \times C} with CC channels and spatial dimensions H×WH \times W, FF is reshaped into a matrix X∈RC×sX \in \mathbb{R}^{C \times s} where s=HWs = HW. The sample covariance matrix Σ∈RC×C\Sigma \in \mathbb{R}^{C \times C} is computed as: Σ=XIˉXT\Sigma = X \bar{I} X^T where Iˉ=1s(I−1s1)\bar{I} = \frac{1}{s}\left(I - \frac{1}{s}\mathbf{1}\right), with II being the s×ss \times s identity matrix and 1\mathbf{1} the s×ss \times s matrix of all ones.

    Covariance normalization is applied to nonlinearly scale the eigenvalues Λ=diag⁡(λ1,…,λC)\Lambda = \operatorname{diag}(\lambda_1, \dots, \lambda_C) of Σ=UΛUT\Sigma = U\Lambda U^T (with orthogonal matrix UU) using a power α=1/2\alpha = 1/2: Y^=Σ1/2=UΛ1/2UT\hat{Y} = \Sigma^{1/2} = U \Lambda^{1/2} U^T

    Global Covariance Pooling (GCP) shrinks the normalized covariance matrix Y^=[y1,…,yC]\hat{Y} = [y_1, \dots, y_C] into a channel descriptor vector z∈RC×1z \in \mathbb{R}^{C \times 1}, where the cc-th channel statistic is the column-wise mean: zc=HGCP(yc)=1C∑i=1Cyc(i)z_c = H_{GCP}(y_c) = \frac{1}{C} \sum_{i=1}^C y_c(i)

    A gating mechanism with channel reduction ratio r=16r = 16 computes the final channel attention weights w∈RC×1w \in \mathbb{R}^{C \times 1}: w=f(WUδ(WDz))w = f(W_U \delta(W_D z)) where WD∈RCr×CW_D \in \mathbb{R}^{\frac{C}{r} \times C} and WU∈RC×CrW_U \in \mathbb{R}^{C \times \frac{C}{r}} are convolutional weight matrices with 1×11 \times 1 filters, δ(⋅)\delta(\cdot) is the ReLU activation function, and f(⋅)f(\cdot) is the sigmoid gating function.

    The input channel feature maps are rescaled element-wise by the attention weights: f^c=wc⋅fc,c∈{1,…,C}\hat{f}_c = w_c \cdot f_c, \quad c \in \{1, \dots, C\}

  3. Knowl 3 — Covariance Normalization Acceleration via Newton-Schulz Iteration

    algorithm

    Because exact eigenvalue decomposition (EIG) on GPU platforms is computationally slow, the matrix square root normalization Σ1/2\Sigma^{1/2} in the second-order channel attention (SOCA) module is approximated iteratively using the Newton-Schulz iterative method.

    To ensure convergence, the covariance matrix Σ∈RC×C\Sigma \in \mathbb{R}^{C \times C} is pre-normalized by its trace tr⁡(Σ)=∑i=1Cλi\operatorname{tr}(\Sigma) = \sum_{i=1}^C \lambda_i, ensuring that ∥Σ^−I∥2<1\|\hat{\Sigma} - I\|_2 < 1. After NN iterations (where N≤5N \le 5 in practice), post-compensation restores the data magnitude.

    Input: Sample covariance matrix Σ∈RC×C\Sigma \in \mathbb{R}^{C \times C}, number of iterations N≤5N \le 5
    Output: Normalized covariance matrix Y^≈Σ1/2∈RC×C\hat{Y} \approx \Sigma^{1/2} \in \mathbb{R}^{C \times C}
    1. Compute trace: t=tr⁡(Σ)=∑i=1CΣi,it = \operatorname{tr}(\Sigma) = \sum_{i=1}^C \Sigma_{i,i}
    2. Pre-normalize covariance matrix: Y0=1tΣY_0 = \frac{1}{t} \Sigma
    3. Initialize inverse approximation: Z0=ICZ_0 = I_C (where ICI_C is the C×CC \times C identity matrix)
    4. for n=1n = 1 to NN do:
    5. Yn=12Yn−1(3IC−Zn−1Yn−1)Y_n = \frac{1}{2} Y_{n-1} (3 I_C - Z_{n-1} Y_{n-1})
    6. Zn=12(3IC−Zn−1Yn−1)Zn−1Z_n = \frac{1}{2} (3 I_C - Z_{n-1} Y_{n-1}) Z_{n-1}
    7. end for
    8. Compensate scale: Y^=tYN\hat{Y} = \sqrt{t} Y_N
    9. return Y^\hat{Y}
  4. Knowl 4 — Share-Source Residual Group (SSRG) with Share-Source Skip Connections

    model/method

    The Share-Source Residual Group (SSRG) structure stacks G=20G = 20 Local-Source Residual Attention Groups (LSRAG) coupled with share-source skip connections (SSC) to ease information flow and bypass abundant low-frequency details from the original low-resolution input.

    Let F0F_0 denote the shallow feature extracted from the input image. The feature transformation for the gg-th LSRAG (g∈{1,…,G}g \in \{1, \dots, G\}) is defined by: Fg=WSSCF0+Hg(Fg−1)F_g = W_{SSC} F_0 + H_g(F_{g-1}) where Hg(⋅)H_g(\cdot) denotes the operation of the gg-th LSRAG, Fg−1F_{g-1} and FgF_g are its input and output, and WSSCW_{SSC} is a learnable convolutional weight matrix. WSSCW_{SSC} is initialized to 0 and gradually learns to weight the contribution of the shallow feature F0F_0.

    After passing through all GG groups, the deep feature FDFF_{DF} produced by the SSRG is obtained as: FDF=WSSCF0+FGF_{DF} = W_{SSC} F_0 + F_G

    This continuous share-source routing prevents gradient vanishing and exploding in networks with more than 400 convolutional layers while enabling the deeper blocks to concentrate on residual high-frequency details.

  5. Knowl 5 — Region-Level Non-Local (RL-NL) Module

    model/method

    The Region-Level Non-Local (RL-NL) module captures spatial correlations and self-similarities within localized image regions while avoiding the quadratic computational and memory footprint of full-image non-local operations.

    Given an intermediate feature map of spatial dimensions H×WH \times W with CC channels, the RL-NL module partitions the feature map into a uniform grid of k×kk \times k rectangular regions (yielding k2k^2 sub-blocks of spatial size Hk×Wk\frac{H}{k} \times \frac{W}{k} with CC channels, where k=2k = 2). Non-local self-attention operations are then executed independently within each sub-block.

    Two RL-NL modules (k=2k = 2) are incorporated into the Non-locally Enhanced Residual Group (NLRG): one immediately before the Share-Source Residual Group (SSRG) to enrich shallow representations with spatial structure cues, and one immediately following the SSRG to integrate spatial contextual information into the deep features before upscaling.

  6. Knowl 6 — Local-Source Residual Attention Group (LSRAG)

    model/method

    A Local-Source Residual Attention Group (LSRAG) serves as the primary building unit of the Share-Source Residual Group (SSRG). Each of the G=20G = 20 LSRAGs contains M=10M = 10 simplified residual blocks followed by a single Second-Order Channel Attention (SOCA) module.

    Within the gg-th LSRAG, the mm-th residual block (m∈{1,…,M}m \in \{1, \dots, M\}) transforms its input Fg,m−1F_{g,m-1} according to: Fg,m=Hg,m(Fg,m−1)F_{g,m} = H_{g,m}(F_{g,m-1}) where each simplified residual block consists of two 3×33 \times 3 convolutional layers with C=64C = 64 filters separated by a ReLU activation function.

    A local-source skip connection combines the group input Fg−1F_{g-1} with the output of the final residual block Fg,MF_{g,M}: Fg=WgFg−1+Fg,MF_g = W_g F_{g-1} + F_{g,M} where WgW_g is a learnable weighting parameter. The resulting feature is then recalibrated by the SOCA module at the tail of the LSRAG.

  7. Knowl 7 — Ablation Study on Architecture Modules

    data/table

    An ablation study on the Set5 dataset (4×4\times super-resolution evaluated after 5.6×1055.6 \times 10^5 training iterations) isolates the individual and cumulative performance contributions of the Region-Level Non-Local (RL-NL) modules, the Share-Source Skip Connection (SSC), First-Order Channel Attention (FOCA), and Second-Order Channel Attention (SOCA).

    Module Base RaR_a RbR_b RcR_c RdR_d ReR_e RfR_f RgR_g RhR_h RiR_i
    RL-NL (before SSRG) ✓ ✓ ✓ ✓ ✓
    RL-NL (after SSRG) ✓ ✓ ✓ ✓ ✓
    Share-source skip connection (SSC) ✓ ✓ ✓ ✓
    First-order channel attention (FOCA) ✓ ✓
    Second-order channel attention (SOCA) ✓ ✓
    PSNR (dB) on Set5 (4×4\times) 32.00 32.04 32.06 32.07 32.12 32.16 32.08 32.10 32.14 32.20

    The baseline model (Base) contains 20 LSRAGs and 10 residual blocks per group (>400 convolutional layers) with conventional skip connections, yielding 32.00 dB. Adding individual modules increases PSNR to 32.04 dB (RaR_a), 32.06 dB (RbR_b), and 32.07 dB (RcR_c). Second-order channel attention (ReR_e, 32.16 dB) outperforms first-order channel attention (RdR_d, 32.12 dB). Combining all proposed modules (RiR_i) achieves the highest PSNR of 32.20 dB.

  8. Knowl 8 — Quantitative Evaluation under Bicubic Degradation

    data/table

    Quantitative benchmark comparison under standard Bicubic (BI) downscaling degradation across five benchmark datasets (Set5, Set14, BSD100, Urban100, Manga109) for scaling factors 2×,3×,4×,2\times, 3\times, 4\times, and 8×8\times. Evaluation metrics are PSNR (dB) and SSIM evaluated on the luminance (Y) channel of the YCbCr color space. SAN+ denotes the model evaluated with self-ensemble.

    Method Scale Set5 Set14 BSD100 Urban100 Manga109
    PSNR / SSIM PSNR / SSIM PSNR / SSIM PSNR / SSIM PSNR / SSIM
    Bicubic 2×2\times 33.66 / 0.9299 30.24 / 0.8688 29.56 / 0.8431 26.88 / 0.8403 30.80 / 0.9339
    RCAN 2×2\times 38.27 / 0.9614 34.11 / 0.9216 32.41 / 0.9026 33.34 / 0.9384 39.43 / 0.9786
    SAN 2×2\times 38.31 / 0.9620 34.07 / 0.9213 32.42 / 0.9028 33.10 / 0.9370 39.32 / 0.9792
    SAN+ 2×2\times 38.35 / 0.9619 34.44 / 0.9244 32.50 / 0.9038 33.73 / 0.9416 39.72 / 0.9797
    Bicubic 3×3\times 30.39 / 0.8682 27.55 / 0.7742 27.21 / 0.7385 24.46 / 0.7349 26.95 / 0.8556
    EDSR 3×3\times 34.65 / 0.9280 30.52 / 0.8462 29.25 / 0.8093 28.80 / 0.8653 34.17 / 0.9476
    RDN 3×3\times 34.71 / 0.9296 30.57 / 0.8468 29.26 / 0.8093 28.80 / 0.8653 34.13 / 0.9484
    RCAN 3×3\times 34.74 / 0.9299 30.64 / 0.8481 29.32 / 0.8111 29.08 / 0.8702 34.43 / 0.9498
    SAN 3×3\times 34.75 / 0.9300 30.59 / 0.8476 29.33 / 0.8112 28.93 / 0.8671 34.30 / 0.9494
    SAN+ 3×3\times 34.89 / 0.9306 30.77 / 0.8498 29.38 / 0.8121 29.29 / 0.8730 34.74 / 0.9512
    Bicubic 4×4\times 28.42 / 0.8104 26.00 / 0.7027 25.96 / 0.6675 23.14 / 0.6577 24.89 / 0.7866
    EDSR 4×4\times 32.46 / 0.8968 28.80 / 0.7876 27.71 / 0.7420 26.64 / 0.8033 31.02 / 0.9148
    RDN 4×4\times 32.47 / 0.8990 28.81 / 0.7871 27.72 / 0.7419 26.61 / 0.8028 31.00 / 0.9151
    RCAN 4×4\times 32.62 / 0.9001 28.86 / 0.7888 27.76 / 0.7435 26.82 / 0.8087 31.21 / 0.9172
    SAN 4×4\times 32.64 / 0.9003 28.92 / 0.7888 27.78 / 0.7436 26.79 / 0.8068 31.18 / 0.9169
    SAN+ 4×4\times 32.70 / 0.9013 29.05 / 0.7921 27.86 / 0.7457 27.23 / 0.8169 31.66 / 0.9222
    Bicubic 8×8\times 24.40 / 0.6580 23.10 / 0.5660 23.67 / 0.5480 20.74 / 0.5160 21.47 / 0.6500
    EDSR 8×8\times 26.96 / 0.7762 24.91 / 0.6420 24.81 / 0.5985 22.51 / 0.6221 24.69 / 0.7841
    DBPN 8×8\times 27.21 / 0.7840 25.13 / 0.6480 24.88 / 0.6010 22.73 / 0.6312 25.14 / 0.7987
    SAN 8×8\times 27.22 / 0.7829 25.14 / 0.6476 24.88 / 0.6011 22.70 / 0.6314 24.85 / 0.7906
    SAN+ 8×8\times 27.30 / 0.7849 25.23 / 0.6493 24.97 / 0.6031 22.91 / 0.6369 25.17 / 0.7964

    SAN+ achieves the highest quantitative scores across all scales and datasets. SAN demonstrates advantages on datasets rich in high-order texture patterns (Set5, Set14, BSD100), while performing competitively with first-order attention methods on datasets dominated by repeated sharp edges (Urban100, Manga109).

  9. Knowl 9 — Quantitative Evaluation under Blur-Downscale Degradation

    data/table

    Quantitative comparison under the Blur-Downscale (BD) degradation model for 3×3\times super-resolution across Set5, Set14, BSD100, Urban100, and Manga109 datasets. PSNR and SSIM are evaluated on the Y channel in YCbCr space.

    Method Set5 Set14 BSD100 Urban100 Manga109
    PSNR / SSIM PSNR / SSIM PSNR / SSIM PSNR / SSIM PSNR / SSIM
    Bicubic 28.78 / 0.8308 26.38 / 0.7271 26.33 / 0.6918 23.52 / 0.6862 25.46 / 0.8149
    SPMSR 32.21 / 0.9001 28.89 / 0.8105 28.13 / 0.7740 25.84 / 0.7856 29.64 / 0.9003
    SRCNN 32.05 / 0.8944 28.80 / 0.8074 28.13 / 0.7736 25.70 / 0.7770 29.47 / 0.8924
    FSRCNN 26.23 / 0.8124 24.44 / 0.7106 24.86 / 0.6832 22.04 / 0.6745 23.04 / 0.7927
    VDSR 33.25 / 0.9150 29.46 / 0.8244 28.57 / 0.7893 26.61 / 0.8136 31.06 / 0.9234
    IRCNN 33.38 / 0.9182 29.63 / 0.8281 28.65 / 0.7922 26.77 / 0.8154 31.15 / 0.9245
    SRMD 34.01 / 0.9242 30.11 / 0.8364 28.98 / 0.8009 27.50 / 0.8370 32.97 / 0.9391
    RDN 34.58 / 0.9280 30.53 / 0.8447 29.23 / 0.8079 28.46 / 0.8582 33.97 / 0.9465
    RCAN 34.70 / 0.9288 30.63 / 0.8462 29.32 / 0.8093 28.81 / 0.8645 34.38 / 0.9483
    SAN 34.75 / 0.9290 30.68 / 0.8466 29.33 / 0.8101 28.83 / 0.8646 34.46 / 0.9487
    SAN+ 34.86 / 0.9297 30.77 / 0.8481 29.39 / 0.8112 29.03 / 0.8674 34.76 / 0.9501

    SAN consistently outperforms all compared methods across all five datasets under the BD degradation setting without requiring self-ensemble, achieving up to a 0.37 dB to 0.49 dB PSNR gain over RDN on the Urban100 and Manga109 benchmarks.

  10. Knowl 10 — Model Complexity and Parameter Efficiency Analysis

    data/table

    Comparison of parameter count (in millions/thousands) and reconstruction performance (PSNR in dB on Set5 for 2×2\times super-resolution) between SAN and contemporary deep CNN super-resolution architectures.

    Metric EDSR MemNet NLRG DBPN RDN RCAN SAN
    Parameters 43M 677k 330k 10M 22.3M 16M 15.7M
    PSNR (dB) 38.11 37.78 38.00 38.09 38.24 38.27 38.31

    SAN achieves a higher PSNR (38.31 dB) than EDSR (38.11 dB, 43M parameters), RDN (38.24 dB, 22.3M parameters), and RCAN (38.27 dB, 16.0M parameters) while utilizing fewer parameters (15.7M), demonstrating superior parameter efficiency and representational capacity.

Coverage note — No substantial contributed material was omitted. All structural modules (SAN architecture, SOCA, SSRG, RL-NL, LSRAG), the Newton-Schulz matrix square root acceleration algorithm, and full quantitative evaluations (ablation studies, BI and BD degradation benchmarks, and parameter analysis) are completely covered.

References

  1. 1.Chao Dong, Chen Change Loy, Kaiming He, and Xiaoou Tang. Learning a deep convolutional network for image super-resolution. In ECCV, 2014. 7
  2. 2.Chao Dong, Chen Change Loy, Kaiming He, and Xiaoou Tang. Image super-resolution using deep convolutional networks. TPAMI, 2016. 1, 2, 3, 7, 8
  3. 3.Chao Dong, Chen Change Loy, and Xiaoou Tang. Accelerating the super-resolution convolutional neural network. In ECCV. Springer, 2016. 3, 7, 8
  4. 4.Weisheng Dong, Lei Zhang, Guangming Shi, and Xiaolin Wu. Image deblurring and super-resolution by adaptive sparse domain selection and adaptive regularization. TIP, 2011. 1
  5. 5.William T Freeman, Egon C Pasztor, and Owen T Carmichael. Learning low-level vision. IJCV, 40(1):25–47, 2000. 1
  6. 6.Muhammad Haris, Greg Shakhnarovich, and Norimichi Ukita. Deep backprojection networks for super-resolution. In CVPR, 2018. 2, 3, 6, 7
  7. 7.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In CVPR, 2016. 1
  8. 8.Nicholas J Higham. Functions of matrices: theory and computation. SIAM, 2008. 5
  9. 9.Jie Hu, Li Shen, and Gang Sun. Squeeze-and-excitation networks. In CVPR, 2018. 2, 4, 5, 6
  10. 10.Gao Huang, Zhuang Liu, Laurens Van Der Maaten, and Kilian Q Weinberger. Densely connected convolutional networks. In CVPR, 2017. 2
  11. 11.Justin Johnson, Alexandre Alahi, and Li Fei-Fei. Perceptual losses for real-time style transfer and super-resolution. In ECCV, 2016. 3
  12. 12.Jiwon Kim, Jung Kwon Lee, and Kyoung Mu Lee. Accurate image super-resolution using very deep convolutional networks. In CVPR, 2016. 1, 2, 3, 7, 8
  13. 13.Jiwon Kim, Jung Kwon Lee, and Kyoung Mu Lee. Deeply-recursive convolutional network for image super-resolution. In CVPR, 2016. 2
  14. 14.Wei-Sheng Lai, Jia-Bin Huang, Narendra Ahuja, and Ming-Hsuan Yang. Deep laplacian pyramid networks for fast and accurate superresolution. In CVPR, 2017. 1, 2, 3, 7
  15. 15.Wei-Sheng Lai, Jia-Bin Huang, Narendra Ahuja, and Ming-Hsuan Yang. Fast and accurate image super-resolution with deep laplacian pyramid networks. arXiv preprint arXiv:1710.01992, 2017. 3
  16. 16.Christof Koch Laurent Itti and Ernst Niebur. A model of saliency-based visual attention for rapid scene analysis. PAMI, 1998. 2
  17. 17.Christian Ledig, Lucas Theis, Ferenc Huszar, Jose Caballero, ´ Andrew Cunningham, Alejandro Acosta, Andrew P Aitken, Alykhan Tejani, Johannes Totz, Zehan Wang, et al. Photorealistic single image super-resolution using a generative adversarial network. In CVPR, 2017. 2
  18. 18.Peihua Li, Jiangtao Xie, Qilong Wang, and Zilin Gao. Towards faster training of global covariance pooling networks by iterative matrix square root normalization. In CVPR, 2018. 5
  19. 19.Peihua Li, Jiangtao Xie, Qilong Wang, and Wangmeng Zuo. Is second-order information helpful for large-scale visual recognition. In ICCV, 2017. 4, 5
  20. 20.Bee Lim, Sanghyun Son, Heewon Kim, Seungjun Nah, and Kyoung Mu Lee. Enhanced deep residual networks for single image super-resolution. In CVPRW, 2017. 2, 3, 5, 6, 7
  21. 21.Tsung-Yu Lin, Aruni RoyChowdhury, and Subhransu Maji. Bilinear cnn models for fine-grained visual recognition. In ICCV, 2015. 4
  22. 22.Ding Liu, Bihan Wen, Yuchen Fan, Chen Change Loy, and Thomas S Huang. Non-local recurrent network for image restoration. In NIPS, 2018. 2, 4, 5, 7
  23. 23.Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer. Automatic differentiation in pytorch. 2017. 6
  24. 24.Tomer Peleg and Michael Elad. A statistical prediction model based on sparse representations for single image super-resolution. TIP, 2014. 8
  25. 25.Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. Faster r-cnn: Towards real-time object detection with region proposal networks. In NIPS, 2015. 1
  26. 26.Mehdi SM Sajjadi, Bernhard Scholkopf, and Michael ¨ Hirsch. Enhancenet: Single image super-resolution through automated texture synthesis. In ICCV, 2017. 3
  27. 27.Jorge Sanchez, Florent Perronnin, Thomas Mensink, and ´ Jakob Verbeek. Image classification with the fisher vector: Theory and practice. IJCV. 4
  28. 28.Wenzhe Shi, Jose Caballero, Ferenc Huszar, Johannes Totz, ´ Andrew P Aitken, Rob Bishop, Daniel Rueckert, and Zehan Wang. Real-time single image and video super-resolution using an efficient sub-pixel convolutional neural network. In CVPR, 2016. 3, 5
  29. 29.Ying Tai, Jian Yang, and Xiaoming Liu. Image super-resolution via deep recursive residual network. In CVPR, 2017. 2, 3
  30. 30.Ying Tai, Jian Yang, Xiaoming Liu, and Chunyan Xu. Memnet: A persistent memory network for image restoration. In CVPR, pages 4539–4547, 2017. 2, 3, 7
  31. 31.Radu Timofte, Eirikur Agustsson, Luc Van Gool, Ming-Hsuan Yang, Lei Zhang, Bee Lim, Sanghyun Son, Heewon Kim, Seungjun Nah, Kyoung Mu Lee, et al. Ntire 2017 challenge on single image super-resolution: Methods and results. In CVPRW, 2017. 6
  32. 32.Shenlong Wang, Lei Zhang, Yan Liang, and Quan Pan. Semi-coupled dictionary learning with applications to image super-resolution and photo-sketch synthesis. In CVPR, 2012. 1
  33. 33.Xiaolong Wang, Ross Girshick, Abhinav Gupta, and Kaiming He. Non-local neural networks. In CVPR, 2018. 2, 3
  34. 34.K. Zhang, X. Gao, D. Tao, and X. Li. Single image super-resolution with non-local means and steering kernel regression. TIP, 21(11):4544–4556, 2012. 1, 2
  35. 35.Kai Zhang, Wangmeng Zuo, Shuhang Gu, and Lei Zhang. Learning deep cnn denoiser prior for image restoration. In CVPR, 2017. 8
  36. 36.Kai Zhang, Wangmeng Zuo, and Lei Zhang. Learning a single convolutional super-resolution network for multiple degradations. In CVPR, 2018. 1, 2, 6, 7, 8
  37. 37.Lei Zhang and Xiaolin Wu. An edge-guided image interpolation algorithm via directional filtering and data fusion. TIP, 2006. 1, 2
  38. 38.Yulun Zhang, Kunpeng Li, Kai Li, Lichen Wang, Bineng Zhong, and Yun Fu. Image super-resolution using very deep residual channel attention networks. In ECCV, 2018. 1, 2, 6, 7, 8
  39. 39.Yulun Zhang, Yapeng Tian, Yu Kong, Bineng Zhong, and Yun Fu. Residual dense network for image super-resolution. In CVPR, 2018. 1, 2, 3, 5, 6, 7, 8

Citation

MLA
Dai, T., et al. “Second-Order Attention Network for Single Image Super-Resolution”. 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019, pp. 11057–66, https://doi.org/10.1109/CVPR.2019.01132.
APA
Dai, T., Cai, J., Zhang, Y., Xia, S.-T., & Zhang, L. (2019). Second-Order Attention Network for Single Image Super-Resolution. 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 11057–11066. https://doi.org/10.1109/CVPR.2019.01132
Chicago
Dai, T., J. Cai, Y. Zhang, S.-T. Xia, and L. Zhang. 2019. “Second-Order Attention Network for Single Image Super-Resolution”. 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 11057–66. https://doi.org/10.1109/CVPR.2019.01132.
Harvard
Dai, T. et al. (2019) “Second-Order Attention Network for Single Image Super-Resolution”, 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp. 11057–11066. Available at: https://doi.org/10.1109/CVPR.2019.01132.
Vancouver
1. Dai T, Cai J, Zhang Y, Xia S-T, Zhang L (2019) Second-Order Attention Network for Single Image Super-Resolution. In: 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp 11057–11066

BibTeX

@inproceedings{Dai_2019, title={Second-Order Attention Network for Single Image Super-Resolution}, url={http://dx.doi.org/10.1109/CVPR.2019.01132}, DOI={10.1109/cvpr.2019.01132}, booktitle={2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)}, publisher={IEEE}, author={Dai, Tao and Cai, Jianrui and Zhang, Yongbing and Xia, Shu-Tao and Zhang, Lei}, year={2019}, month=June, pages={11057–11066} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE