Transcending the Limit of Local Window: Advanced Super-Resolution Transformer with Adaptive Token Dictionary

Leheng ZhangYawei LiXingyu ZhouXiaorui ZhaoShuhang Gu

article2024CVPR95 citations

Proposes an adaptive token dictionary for super-resolution Transformers that overcomes the receptive-field limits of window-based attention by dynamically grouping similar image tokens across the entire image to capture long-range dependencies.

Listen

Single-image super-resolution—the process of reconstructing high-resolution images from solitary, low-resolution inputs—is critical for overcoming the physical limitations of low-cost sensors and enhancing legacy imagery. While modern vision transformer architectures achieve strong performance by capturing structural relationships, their computational cost grows quadratically with image size. To manage this burden, existing models restrict their attention mechanism to fixed, local rectangular windows. However, this artificial constraint limits the model's receptive field, preventing it from capturing long-range dependencies across the entire image and grouping unrelated image parts together simply because they are spatially adjacent.

The objective of the article is to demonstrate an advanced super-resolution transformer architecture that overcomes local window boundaries using an adaptive token dictionary. The proposed framework evaluates how external visual priors and global image content can be integrated to capture long-range, content-based similarities without creating excessive computational complexity.

To achieve this, the authors integrated three core mechanisms: a cross-attention module that matches image tokens to a learned auxiliary token dictionary containing general visual structures, an adaptive refinement strategy that dynamically updates the dictionary layer by layer using the test image's specific features, and a category-based self-attention mechanism that groups similar tokens across the entire image regardless of distance. The architecture was evaluated across five standard image benchmark datasets (Set5, Set14, BSD100, Urban100, and Manga109) across both standard and lightweight model configurations, using standard image fidelity metrics.

The evaluation produced four primary findings. First, the proposed full model consistently surpassed leading state-of-the-art architectures, delivering an improvement of 0.25 to 0.28 dB on the Urban100 benchmark across multiple magnification factors with a comparable parameter footprint (approximately 20 million parameters). Second, the lightweight version outperformed competing lightweight models across all benchmarks, notably exceeding the nearest competitor by 0.46 dB on the Manga109 four-times upscaling task while maintaining a compact size of roughly 769,000 parameters. Third, ablation testing confirmed that each component is necessary: combining cross-attention and dictionary refinement provided steady gains, while category-based attention drove the largest performance surge (up to 0.19 dB). Finally, sensitivity testing revealed optimal operating thresholds at a dictionary size of 64 to 128 tokens and a sub-category size of 128; exceeding these points led to diminishing returns or slight performance drops due to over-parameterization.

These findings indicate that content-dependent feature grouping provides a superior, more computationally efficient alternative to arbitrary spatial window partitioning in computer vision. For organizations deploying imaging systems, this approach lowers computational overhead while delivering noticeably sharper edge reconstruction and cleaner fine textures. The ability to deploy high-performing lightweight models also makes advanced image restoration feasible on resource-constrained edge devices and mobile hardware.

Engineering and deployment teams working on image enhancement should consider adopting category-based token grouping and adaptive dictionaries over standard windowed attention. When implementing this architecture, teams should carefully tune dictionary capacity and group partition sizes to prevent model saturation. Further work should explore validating the model on broader real-world camera artifacts, varied degradation types beyond standard synthetic downsampling, and actual hardware inference latency across production environments.

Cover for Transcending the Limit of Local Window: Advanced Super-Resolution Transformer with Adaptive Token Dictionary

Abstract

Single Image Super-Resolution is a classic computer vision problem that involves estimating high-resolution (HR) images from low-resolution (LR) ones. Although deep neural networks (DNNs), especially Transformers for super-resolution, have seen significant advancements in recent years, challenges still remain, particularly in limited receptive field caused by window-based self-attention. To address these issues, we introduce a group of auxiliary Adaptive Token Dictionary to SR Transformer and establish an ATD-SR method. The introduced token dictionary could learn prior information from training data and adapt the learned prior to specific testing image through an adaptive refinement step. The refinement strategy could not only provide global information to all input tokens but also group image tokens into categories. Based on category partitions, we further propose a category-based self-attention mechanism designed to leverage distant but similar tokens for enhancing input features. The experimental results show that our method achieves the best performance on various single image super-resolution benchmarks.

Table of Contents

  • 1. Introduction
  • 2. Related Works
  • 3. Methodology
  • 3.1. Motivation
  • 3.2. Token Dictionary Cross-Attention
  • 3.3. Adaptive Dictionary Refinement
  • 3.4. Adaptive Category-based Attention
  • 3.5. The Overall Network Architecture
  • 4. Experiments
  • 4.1. Experimental Settings
  • 4.2. Ablation Study
  • 4.3. Comparisons with State-of-the-Art Methods
  • 4.4. Visualization Analysis
  • 5. Conclusion
  • References

Knowls

  1. Knowl 1 — Token Dictionary Cross-Attention Mechanism

    model/method

    Token Dictionary Cross-Attention (TDCA) incorporates learned external priors into image feature representation with linear computational complexity with respect to the total number of image tokens.

    Let X∈RN×dX \in \mathbb{R}^{N \times d} denote the input feature map representing NN image tokens with channel dimension dd, and let D∈RM×dD \in \mathbb{R}^{M \times d} denote an auxiliary token dictionary containing MM learnable tokens, where M≪NM \ll N. The query, key, and value matrices are projected via:

    QX=XWQ,KD=DWK,VD=DWVQ_X = X W^Q, \quad K_D = D W^K, \quad V_D = D W^V

    where WQ∈Rd×(d/r)W^Q \in \mathbb{R}^{d \times (d/r)}, WK∈Rd×(d/r)W^K \in \mathbb{R}^{d \times (d/r)}, and WV∈Rd×dW^V \in \mathbb{R}^{d \times d} are linear projection matrices, and rr is a channel reduction ratio. TDCA computes the cosine similarity between each query image token and key dictionary token, normalizes with a learnable temperature τ\tau, and aggregates the value dictionary tokens:

    A=SoftMax⁡(Sim⁡cos⁡(QX,KD)τ)A = \operatorname{SoftMax}\left(\frac{\operatorname{Sim}_{\cos}(Q_X, K_D)}{\tau}\right)

    TDCA⁡(QX,KD,VD)=AVD\operatorname{TDCA}(Q_X, K_D, V_D) = A V_D

    where Sim⁡cos⁡(QX,KD)∈RN×M\operatorname{Sim}_{\cos}(Q_X, K_D) \in \mathbb{R}^{N \times M} computes the cosine similarity matrix, and A∈RN×MA \in \mathbb{R}^{N \times M} is the resulting cross-attention map. Cosine normalization gives each dictionary atom an equal chance of being selected regardless of token magnitude.

  2. Knowl 2 — Adaptive Dictionary Refinement Strategy

    model/method

    Adaptive Dictionary Refinement (ADR) dynamically updates the token dictionary across successive layers of a transformer network to combine external dataset priors with image-specific global internal features using a reversed attention formulation.

    Let X(l)X^{(l)} and D(l)∈RM×dD^{(l)} \in \mathbb{R}^{M \times d} denote the input feature and dictionary at the ll-th layer. The initial dictionary D(1)D^{(1)} is parameterized as learnable weights. For layer l≥1l \ge 1, let A(l)∈RN×MA^{(l)} \in \mathbb{R}^{N \times M} be the TDCA attention map computed at layer ll, and let X(l+1)∈RN×dX^{(l+1)} \in \mathbb{R}^{N \times d} be the updated output feature map from layer ll. A candidate dictionary D^(l)\hat{D}^{(l)} is constructed by aggregating features across the entire image:

    D^(l)=SoftMax⁡(Norm⁡(A(l)T))X(l+1)\hat{D}^{(l)} = \operatorname{SoftMax}\left(\operatorname{Norm}\left({A^{(l)}}^T\right)\right) X^{(l+1)}

    D(l+1)=σD^(l)+(1−σ)D(l)D^{(l+1)} = \sigma \hat{D}^{(l)} + (1 - \sigma) D^{(l)}

    where Norm⁡(⋅)\operatorname{Norm}(\cdot) denotes a normalization layer that adjusts the attention map scale, and σ∈[0,1]\sigma \in [0, 1] is a learnable interpolation weight. This mechanism propagates global image-specific contexts across layer boundaries without partitioning the image into local windows.

  3. Knowl 3 — Adaptive Category-Based Multi-Head Self-Attention

    model/method

    Adaptive Category-based Multi-head Self-Attention (AC-MSA) enables long-range non-local attention by dynamically partitioning image tokens into content categories based on dictionary similarity rather than fixed spatial rectangular windows.

    Given the cross-attention matrix A∈RN×MA \in \mathbb{R}^{N \times M} from TDCA, token xjx_j is assigned to category θi\theta^i corresponding to the dictionary atom with which it has the maximum attention weight:

    θi={xj∣arg⁡max⁡k(Ajk)=i},i∈{1,…,M}\theta^i = \{ x_j \mid \arg\max_k (A_{jk}) = i \}, \quad i \in \{1, \dots, M\}

    To balance category sizes for parallel computation, all categorized tokens are concatenated and segmented into sub-categories ϕj\phi^j of a fixed size nsn_s:

    ϕ=[θ11,…,θn11,…,θnMM]\phi = \left[ \theta_1^1, \dots, \theta_{n_1}^1, \dots, \theta_{n_M}^M \right]

    ϕj=[ϕj⋅ns+1,ϕj⋅ns+2,…,ϕ(j+1)⋅ns]\phi^j = \left[ \phi_{j \cdot n_s + 1}, \phi_{j \cdot n_s + 2}, \dots, \phi_{(j+1) \cdot n_s} \right]

    The overall AC-MSA operation processes input tokens XinX_{\text{in}} as follows:

    {ϕj}=Categorize⁡(Xin)\{ \phi^j \} = \operatorname{Categorize}(X_{\text{in}})

    ϕ^j=MSA⁡(ϕjWQ,ϕjWK,ϕjWV)\hat{\phi}^j = \operatorname{MSA}\left(\phi^j W^Q, \phi^j W^K, \phi^j W^V\right)

    Xout=UnCategorize⁡({ϕ^j})X_{\text{out}} = \operatorname{UnCategorize}\left(\{ \hat{\phi}^j \}\right)

    where MSA⁡\operatorname{MSA} is standard multi-head self-attention applied independently within each sub-category chunk ϕj\phi^j, and UnCategorize⁡\operatorname{UnCategorize} maps each transformed token back to its original spatial position.

  4. Knowl 4 — Adaptive Token Dictionary (ATD) Network Architecture

    model/method

    The Adaptive Token Dictionary (ATD) network for single image super-resolution processes low-resolution inputs through three stages:

    1. Shallow Feature Extraction: A single 3×33 \times 3 convolutional layer projects the input image into deep feature space.
    2. Deep Feature Extraction: A sequence of ATD blocks. Each ATD block contains a sequence of transformer layers and an initial learnable token dictionary D(1)D^{(1)}. Within each transformer layer, three parallel attention modules process the features simultaneously:
      • Token Dictionary Cross-Attention (TDCA)
      • Adaptive Category-based Multi-head Self-Attention (AC-MSA)
      • Shifted Window-based Multi-head Self-Attention (SW-MSA)

    The outputs of the three branches are summed, processed through LayerNorm, and passed through a Feed-Forward Network (FFN). The token dictionary is iteratively refined across layers via Adaptive Dictionary Refinement (ADR). 3. HR Reconstruction: A 3×33 \times 3 convolution followed by a pixel shuffle upsampling operation generates the final high-resolution estimate.

    Two configurations are defined:

    • ATD (Standard): 6 ATD blocks, 6 transformer layers per block, 210 channels, dictionary size M=128M = 128, channel reduction ratio r=10.5r = 10.5 (reducing dimension to 20 for similarity calculation), and sub-category size ns=128n_s = 128.
    • ATD-light: 4 ATD blocks, 48 channels, dictionary size M=64M = 64, reduction ratio r=6r = 6 (reducing dimension to 8 for similarity calculation), and sub-category size ns=128n_s = 128.
  5. Knowl 5 — Ablation Study on ATD Core Components

    data/table

    Ablation experiments evaluated on the Urban100 and Manga109 benchmarks (imes4 imes 4 super-resolution, models trained for 250k iterations on DIV2K) demonstrate the cumulative benefit of Token Dictionary Cross-Attention (TDCA), Adaptive Dictionary Refinement (ADR), and Adaptive Category-based Multi-head Self-Attention (AC-MSA) over a baseline using only shifted window-based self-attention (SW-MSA):

    TDCA ADR AC-MSA Urban100 Manga109
    PSNR (dB) SSIM PSNR (dB) SSIM
    26.25 0.7907 30.66 0.9113
    ✓ 26.32 0.7929 30.76 0.9118
    ✓ ✓ 26.36 0.7931 30.79 0.9123
    ✓ ✓ ✓ 26.51 0.7975 30.98 0.9144

    Directly adding TDCA improves PSNR by 0.07 dB on Urban100 and 0.10 dB on Manga109. ADR provides an additional gain of 0.04 dB and 0.03 dB. Adding AC-MSA brings an additional improvement of 0.15 dB on Urban100 and 0.19 dB on Manga109, totaling a 0.26 dB and 0.32 dB improvement over the baseline.

  6. Knowl 6 — Ablation on Sub-Category Size and Dictionary Size

    data/table

    The table below summarizes ablation results on the sub-category size nsn_s in AC-MSA and the token dictionary size MM in TDCA on Urban100 and Manga109 (imes4 imes 4 SR):

    nsn_s Urban100 Manga109 MM Urban100 Manga109
    PSNR SSIM PSNR SSIM PSNR SSIM PSNR SSIM
    0 26.36 0.7931 30.79 0.9123 16 26.45 0.7950 30.90 0.9137
    64 26.44 0.7948 30.90 0.9131 32 26.49 0.7965 30.97 0.9142
    128 26.51 0.7975 30.98 0.9144 64 26.51 0.7975 30.98 0.9144
    192 26.55 0.7984 31.01 0.9150 96 26.51 0.7970 30.95 0.9141

    Setting ns=0n_s = 0 corresponds to removing AC-MSA. Performance improves sharply as nsn_s increases up to 128, beyond which gains taper off (+0.04+0.04 dB from 128 to 192), making ns=128n_s = 128 the chosen operating point. For dictionary size MM, performance increases from M=16M=16 to M=64M=64, while M=96M=96 exhibits slight performance degradation (SSIM drops on Urban100 and Manga109), showing that excessively large token counts harm modeling efficiency.

  7. Knowl 7 — Ablation on Category-Based Partitioning Strategy

    data/table

    Comparison of different grouping strategies for category-based attention evaluated on Set5, Urban100, and Manga109 (imes4 imes 4 SR):

    Model Set5 Urban100 Manga109
    PSNR SSIM PSNR SSIM PSNR SSIM
    w/o CA 32.30 0.8957 26.25 0.7907 30.66 0.9113
    random CA 32.38 0.8962 26.46 0.7955 30.92 0.9139
    adaptive CA 32.46 0.8973 26.51 0.7975 30.98 0.9144

    Random category partitioning ('random CA') uses an untrained random dictionary for token grouping, which outperforms the baseline without CA ('w/o CA'). The learned adaptive token dictionary ('adaptive CA') further improves PSNR by 0.08 dB on Set5, 0.05 dB on Urban100, and 0.06 dB on Manga109 over random partitioning due to more accurate semantic/content-aligned token clustering.

  8. Knowl 8 — Classical Image Super-Resolution Benchmark Performance

    data/table

    Quantitative performance comparison (PSNR [dB] / SSIM) on five standard benchmark datasets (Set5, Set14, BSD100, Urban100, Manga109) for scale factors ×2\times 2 and ×4\times 4:

    Method Scale Params Set5 Set14 BSD100 Urban100 Manga109
    PSNR SSIM PSNR SSIM PSNR SSIM PSNR SSIM PSNR SSIM
    EDSR ×2\times 2 42.6M 38.11 0.9602 33.92 0.9195 32.32 0.9013 32.93 0.9351 39.10 0.9773
    RCAN ×2\times 2 15.4M 38.27 0.9614 34.12 0.9216 32.41 0.9027 33.34 0.9384 39.44 0.9786
    SAN ×2\times 2 15.7M 38.31 0.9620 34.07 0.9213 32.42 0.9028 33.10 0.9370 39.32 0.9792
    HAN ×2\times 2 63.6M 38.27 0.9614 34.16 0.9217 32.41 0.9027 33.35 0.9385 39.46 0.9785
    IPT ×2\times 2 115M 38.37 - 34.43 - 32.48 - 33.76 - - -
    SwinIR ×2\times 2 11.8M 38.42 0.9623 34.46 0.9250 32.53 0.9041 33.81 0.9433 39.92 0.9797
    EDT ×2\times 2 11.5M 38.45 0.9624 34.57 0.9258 32.52 0.9041 33.80 0.9425 39.93 0.9800
    CAT-A ×2\times 2 16.5M 38.51 0.9626 34.78 0.9265 32.59 0.9047 34.26 0.9440 40.10 0.9805
    ART ×2\times 2 16.4M 38.56 0.9629 34.59 0.9267 32.58 0.9048 34.30 0.9452 40.24 0.9808
    HAT ×2\times 2 20.6M 38.63 0.9630 34.86 0.9274 32.62 0.9053 34.45 0.9466 40.26 0.9809
    ATD (Ours) ×2\times 2 20.1M 38.61 0.9629 34.92 0.9275 32.64 0.9054 34.73 0.9476 40.35 0.9810
    EDSR ×4\times 4 43.0M 32.46 0.8968 28.80 0.7876 27.71 0.7420 26.64 0.8033 31.02 0.9148
    RCAN ×4\times 4 15.6M 32.63 0.9002 28.87 0.7889 27.77 0.7436 26.82 0.8087 31.22 0.9173
    SAN ×4\times 4 15.9M 32.64 0.9003 28.92 0.7888 27.78 0.7436 26.79 0.8068 31.18 0.9169
    HAN ×4\times 4 64.2M 32.64 0.9002 28.90 0.7890 27.80 0.7442 26.85 0.8094 31.42 0.9177
    IPT ×4\times 4 116M 32.64 - 29.01 - 27.82 - 27.26 - - -
    SwinIR ×4\times 4 11.9M 32.92 0.9044 29.09 0.7950 27.92 0.7489 27.45 0.8254 32.03 0.9260
    EDT ×4\times 4 11.6M 32.82 0.9031 29.09 0.7939 27.91 0.7483 27.46 0.8246 32.05 0.9254
    CAT-A ×4\times 4 16.6M 33.08 0.9052 29.18 0.7960 27.99 0.7510 27.89 0.8339 32.39 0.9285
    ART ×4\times 4 16.6M 33.04 0.9051 29.16 0.7958 27.97 0.7510 27.77 0.8321 32.31 0.9283
    HAT ×4\times 4 20.8M 33.04 0.9056 29.23 0.7973 28.00 0.7517 27.97 0.8368 32.48 0.9292
    ATD (Ours) ×4\times 4 20.3M 33.14 0.9061 29.25 0.7976 28.02 0.7524 28.22 0.8414 32.65 0.9308

    ATD achieves the highest PSNR and SSIM across all datasets on ×4\times 4 SR, outperforming HAT by 0.25 dB on Urban100 and 0.17 dB on Manga109 while using slightly fewer parameters (20.3M vs 20.8M).

  9. Knowl 9 — Lightweight Image Super-Resolution Benchmark Performance

    data/table

    Quantitative performance comparison (PSNR [dB] / SSIM) on lightweight SR benchmarks (imes2 imes 2 and imes4 imes 4):

    Method Scale Params Set5 Set14 BSD100 Urban100 Manga109
    PSNR SSIM PSNR SSIM PSNR SSIM PSNR SSIM PSNR SSIM
    CARN ×2\times 2 1,592K 37.76 0.9590 33.52 0.9166 32.09 0.8978 31.92 0.9256 38.36 0.9765
    IMDN ×2\times 2 694K 38.00 0.9605 33.63 0.9177 32.19 0.8996 32.17 0.9283 38.88 0.9774
    LAPAR-A ×2\times 2 548K 38.01 0.9605 33.62 0.9183 32.19 0.8999 32.10 0.9283 38.67 0.9772
    LatticeNet ×2\times 2 756K 38.15 0.9610 33.78 0.9193 32.25 0.9005 32.43 0.9302 - -
    SwinIR-light ×2\times 2 910K 38.14 0.9611 33.86 0.9206 32.31 0.9012 32.76 0.9340 39.12 0.9783
    ELAN ×2\times 2 582K 38.17 0.9611 33.94 0.9207 32.30 0.9012 32.76 0.9340 39.11 0.9782
    SwinIR-NG ×2\times 2 1181K 38.17 0.9612 33.94 0.9205 32.31 0.9013 32.78 0.9340 39.20 0.9781
    OmniSR ×2\times 2 772K 38.22 0.9613 33.98 0.9210 32.36 0.9020 33.05 0.9363 39.28 0.9784
    ATD-light ×2\times 2 753K 38.29 0.9616 34.10 0.9217 32.39 0.9023 33.27 0.9375 39.52 0.9789
    CARN ×4\times 4 1,592K 32.13 0.8937 28.60 0.7806 27.58 0.7349 26.07 0.7837 30.47 0.9084
    IMDN ×4\times 4 715K 32.21 0.8948 28.58 0.7811 27.56 0.7353 26.04 0.7838 30.45 0.9075
    LAPAR-A ×4\times 4 659K 32.15 0.8944 28.61 0.7818 27.61 0.7366 26.14 0.7871 30.42 0.9074
    LatticeNet ×4\times 4 777K 32.30 0.8962 28.68 0.7830 27.62 0.7367 26.25 0.7873 - -
    SwinIR-light ×4\times 4 930K 32.44 0.8976 28.77 0.7858 27.69 0.7406 26.47 0.7980 30.92 0.9151
    ELAN ×4\times 4 582K 32.43 0.8975 28.78 0.7858 27.69 0.7406 26.54 0.7982 30.92 0.9150
    SwinIR-NG ×4\times 4 1201K 32.44 0.8980 28.83 0.7870 27.73 0.7418 26.61 0.8010 31.09 0.9161
    OmniSR ×4\times 4 792K 32.49 0.8988 28.78 0.7859 27.71 0.7415 26.65 0.8018 31.02 0.9151
    ATD-light ×4\times 4 769K 32.63 0.8998 28.89 0.7886 27.79 0.7440 26.97 0.8107 31.48 0.9198

    ATD-light (769K parameters for ×4\times 4) consistently outperforms OmniSR (792K parameters) across all datasets, achieving a 0.32 dB PSNR lead on Urban100 and a 0.46 dB lead on Manga109 for ×4\times 4 super-resolution.

Coverage note — None was omitted; all primary contributions (TDCA, ADR, AC-MSA, overall ATD architectures, component and hyperparameter ablations, and full benchmark comparisons on classical and lightweight SR) have been captured.

References

  1. 1.Namhyuk Ahn, Byungkon Kang, and Kyung-Ah Sohn. Fast, accurate, and lightweight super-resolution with cascading residual network, 2018. 7, 8
  2. 2.Marco Bevilacqua, Aline Roumy, Christine Guillemot, and Marie-line Alberi Morel. Low-complexity single-image super-resolution based on nonnegative neighbor embedding. In Procedings of the British Machine Vision Conference 2012, 2012. 6, 7
  3. 3.Hanting Chen, Yunhe Wang, Tianyu Guo, Chang Xu, Yiping Deng, Zhenhua Liu, Siwei Ma, Chunjing Xu, Chao Xu, and Wen Gao. Pre-trained image processing transformer, 2020. 2, 7
  4. 4.Xiangyu Chen, Xintao Wang, Jiantao Zhou, Yu Qiao, and Chao Dong. Activating more pixels in image super-resolution transformer. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 22367–22377, 2023. 1, 2, 7, 8
  5. 5.Zheng Chen, Yulun Zhang, Jinjin Gu, Yongbing Zhang, Linghe Kong, and Xin Yuan. Cross aggregation transformer for image restoration. In NeurIPS, 2022. 2, 7
  6. 6.Haram Choi, Jeongmin Lee, and Jihoon Yang. N-gram in swin transformers for efficient lightweight image super-resolution, 2022. 7, 8
  7. 7.Xiangxiang Chu, Zhi Tian, Yuqing Wang, Bo Zhang, Haibing Ren, Xiaolin Wei, Huaxia Xia, and Chunhua Shen. Twins: Re-visiting the design of spatial attention in vision transformers, 2021. 2
  8. 8.Marcos V Conde, Ui-Jin Choi, Maxime Burchi, and Radu Timofte. Swin2sr: Swinv2 transformer for compressed image super-resolution and restoration. In Computer Vision–ECCV 2022 Workshops: Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part II, pages 669–687. Springer, 2023. 3
  9. 9.Tao Dai, Jianrui Cai, Yongbing Zhang, Shu-Tao Xia, and Lei Zhang. Second-order attention network for single image super-resolution. In 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020. 1, 2, 7
  10. 10.Chao Dong, Chen Change Loy, Kaiming He, and Xiaoou Tang. Image super-resolution using deep convolutional networks. IEEE Transactions on Pattern Analysis and Machine Intelligence, page 295–307, 2015. 1, 2
  11. 11.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition at scale, 2020. 2
  12. 12.Shuhang Gu, Wangmeng Zuo, Qi Xie, Deyu Meng, Xiangchu Feng, and Lei Zhang. Convolutional sparse coding for image super-resolution. In Proceedings of the IEEE International Conference on Computer Vision, pages 1823–1831, 2015. 3
  13. 13.Shuhang Gu, Shi Guo, Wangmeng Zuo, Yunjin Chen, Radu Timofte, Luc Van Gool, and Lei Zhang. Learned dynamic guidance for depth image reconstruction. IEEE Transactions on Pattern Analysis and Machine Intelligence, 42(10):2437–2452, 2019. 2
  14. 14.He He and Wan-Chi Siu. Single image super-resolution using gaussian process regression. In CVPR 2011, pages 449–456. IEEE, 2011. 1
  15. 15.Jia-Bin Huang, Abhishek Singh, and Narendra Ahuja. Single image super-resolution from transformed self-exemplars. In 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2015. 6, 7
  16. 16.Zheng Hui, Xinbo Gao, Yunchu Yang, and Xiumei Wang. Lightweight image super-resolution with information multi-distillation network. In Proceedings of the 27th ACM International Conference on Multimedia, 2019. 7, 8
  17. 17.Jiwon Kim, Jung Kwon Lee, and Kyoung Mu Lee. Accurate image super-resolution using very deep convolutional networks. In 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016. 1, 2
  18. 18.Jiwon Kim, Jung Kwon Lee, and Kyoung Mu Lee. Deeply-recursive convolutional network for image super-resolution. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 1637–1645, 2016. 2
  19. 19.Wenbo Li, Kun Zhou, Lu Qi, Nianjuan Jiang, Jiangbo Lu, and Jiaya Jia. Lapar: Linearly-assembled pixel-adaptive regression network for single image super-resolution and beyond, 2020. 7, 8
  20. 20.Wenbo Li, Xin Lu, Jiangbo Lu, Xiangyu Zhang, and Jiaya Jia. On efficient transformer and image pre-training for low-level vision. arXiv preprint arXiv:2112.10175, 2021. 2, 3, 7
  21. 21.Yawei Li, Yuchen Fan, Xiaoyu Xiang, Denis Demandolx, Rakesh Ranjan, Radu Timofte, and Luc Van Gool. Efficient and explicit modelling of image hierarchies for image restoration. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2023. 1, 2, 3
  22. 22.Jingyun Liang, Jiezhang Cao, Guolei Sun, Kai Zhang, Luc Van Gool, and Radu Timofte. Swinir: Image restoration using swin transformer, 2021. 1, 2, 3, 6, 7, 8
  23. 23.Bee Lim, Sanghyun Son, Heewon Kim, Seungjun Nah, and Kyoung Mu Lee. Enhanced deep residual networks for single image super-resolution. In 2017 IEEE Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), 2017. 1, 2, 7
  24. 24.Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows, 2021. 2, 3, 5, 6
  25. 25.Ze Liu, Han Hu, Yutong Lin, Zhuliang Yao, Zhenda Xie, Yixuan Wei, Jia Ning, Yue Cao, Zheng Zhang, Li Dong, et al. Swin transformer v2: Scaling up capacity and resolution. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 12009–12019, 2022. 3
  26. 26.Xiaotong Luo, Yuan Xie, Yulun Zhang, Yanyun Qu, Cuihua Li, and Yun Fu. LatticeNet: Towards Lightweight Image Super-Resolution with Lattice Block, page 272–289. 2020. 7, 8
  27. 27.Julien Mairal, Jean Ponce, Guillermo Sapiro, Andrew Zisserman, and Francis Bach. Supervised dictionary learning. Advances in neural information processing systems, 21, 2008. 3
  28. 28.Julien Mairal, Francis Bach, Jean Ponce, and Guillermo Sapiro. Online learning for matrix factorization and sparse coding. Journal of Machine Learning Research, 11(1), 2010. 3
  29. 29.D. Martin, C. Fowlkes, D. Tal, and J. Malik. A database of human segmented natural images and its application to evaluating segmentation algorithms and measuring ecological statistics. In Proceedings Eighth IEEE International Conference on Computer Vision. ICCV 2001, 2002. 7
  30. 30.Yusuke Matsui, Kota Ito, Yuji Aramaki, Azuma Fujimoto, Toru Ogawa, Toshihiko Yamasaki, and Kiyoharu Aizawa. Sketch-based manga retrieval using manga109 dataset. Multimedia Tools and Applications, page 21811–21838, 2016. 6, 7
  31. 31.Yiqun Mei, Yuchen Fan, Yuqian Zhou, Lichao Huang, Thomas S Huang, and Humphrey Shi. Image super-resolution with cross-scale non-local attention and exhaustive self-exemplars mining. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2020. 2
  32. 32.Yiqun Mei, Yuchen Fan, and Yuqian Zhou. Image super-resolution with non-local sparse attention. In 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2021. 2, 5
  33. 33.Ben Niu, Weilei Wen, Wenqi Ren, Xiangde Zhang, Lianping Yang, Shuzhen Wang, Kaihao Zhang, Xiaochun Cao, and Haifeng Shen. Single Image Super-Resolution via a Holistic Attention Network, page 191–207. 2020. 2, 7
  34. 34.Radu Timofte, Eirikur Agustsson, Luc Van Gool, Ming-Hsuan Yang, Lei Zhang, Bee Lim, Sanghyun Son, Heewon Kim, Seungjun Nah, Kyoung Mu Lee, et al. Ntire 2017 challenge on single image super-resolution: Methods and results. In 2017 IEEE Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), 2017. 6
  35. 35.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in neural information processing systems, 30, 2017. 2
  36. 36.Hang Wang, Xuanhong Chen, Bingbing Ni, Yutian Liu, and Liu jinfan. Omni aggregation networks for lightweight image super-resolution. In Conference on Computer Vision and Pattern Recognition, 2023. 2, 7, 8
  37. 37.Wenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan, Kaitao Song, Ding Liang, Tong Lu, Ping Luo, and Ling Shao. Pyramid vision transformer: A versatile backbone for dense prediction without convolutions. In 2021 IEEE/CVF International Conference on Computer Vision (ICCV), 2022. 2
  38. 38.Xiaolong Wang, Ross Girshick, Abhinav Gupta, and Kaiming He. Non-local neural networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 7794–7803, 2018. 2
  39. 39.Jianchao Yang, John Wright, Thomas S Huang, and Yi Ma. Image super-resolution via sparse representation. IEEE transactions on image processing, 19(11):2861–2873, 2010. 1, 3
  40. 40.Roman Zeyde, Michael Elad, and Matan Protter. On Single Image Scale-Up Using Sparse-Representations, page 711–730. 2012. 3, 7
  41. 41.Jiale Zhang, Yulun Zhang, Jinjin Gu, Yongbing Zhang, Linghe Kong, and Xin Yuan. Accurate image restoration with attention retractable transformer. In ICLR, 2023. 2, 7
  42. 42.Xindong Zhang, Hui Zeng, Shi Guo, and Lei Zhang. Efficient long-range attention network for image super-resolution. In European Conference on Computer Vision, 2022. 2, 7, 8
  43. 43.Yulun Zhang, Kunpeng Li, Kai Li, Lichen Wang, Bineng Zhong, and Yun Fu. Image Super-Resolution Using Very Deep Residual Channel Attention Networks, page 294–310. 2018. 1, 2, 7
  44. 44.Yulun Zhang, Yapeng Tian, Yu Kong, Bineng Zhong, and Yun Fu. Residual dense network for image super-resolution, 2018. 2

Citation

MLA
Zhang, L., et al. “Transcending the Limit of Local Window: Advanced Super-Resolution Transformer with Adaptive Token Dictionary”. arXiv, 2024, http://arxiv.org/abs/2401.08209v2.
APA
Zhang, L., Li, Y., Zhou, X., Zhao, X., & Gu, S. (2024). Transcending the Limit of Local Window: Advanced Super-Resolution Transformer with Adaptive Token Dictionary. arXiv. http://arxiv.org/abs/2401.08209v2
Chicago
Zhang, L., Y. Li, X. Zhou, X. Zhao, and S. Gu. 2024. “Transcending the Limit of Local Window: Advanced Super-Resolution Transformer with Adaptive Token Dictionary”. arXiv. http://arxiv.org/abs/2401.08209v2.
Harvard
Zhang, L. et al. (2024) “Transcending the Limit of Local Window: Advanced Super-Resolution Transformer with Adaptive Token Dictionary”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2401.08209v2.
Vancouver
1. Zhang L, Li Y, Zhou X, Zhao X, Gu S (2024) Transcending the Limit of Local Window: Advanced Super-Resolution Transformer with Adaptive Token Dictionary. arXiv

BibTeX

@article{zhang2024transcending,
  title = {Transcending the Limit of Local Window: Advanced Super-Resolution Transformer with Adaptive Token Dictionary},
  author = {Zhang, Leheng and Li, Yawei and Zhou, Xingyu and Zhao, Xiaorui and Gu, Shuhang},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2401.08209v2},
  eprint = {2401.08209}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE