Activating More Pixels in Image Super-Resolution Transformer

Xiangyu ChenXintao WangJiantao ZhouYu QiaoChao Dong

article2023CVPR1,146 citations

Proposes a Hybrid Attention Transformer that combines channel and window self-attention with overlapping cross-attention to expand the spatial range of activated pixels, outperforming existing super-resolution methods by over 1 dB.

Listen

Single-image super-resolution, which reconstructs high-resolution images from low-resolution inputs, is crucial for applications ranging from satellite imagery to digital entertainment. While modern Transformer-based artificial intelligence architectures have achieved strong results in image reconstruction, diagnostic evaluations show that they fail to use available information across the entire image. Existing models rely on narrow, localized areas of the input and suffer from visual blocking artifacts, leaving significant performance potential untapped.

The article evaluates a new deep-learning model designed to activate a much broader spatial range of input pixels for image reconstruction. The primary objective is to demonstrate that combining global channel attention with local self-attention, alongside improved cross-window data sharing and large-scale task-specific pre-training, significantly advances state-of-the-art super-resolution quality.

The authors designed the Hybrid Attention Transformer (HAT) architecture and conducted extensive benchmark experiments across five standard datasets, including Urban100 and Manga109. The approach integrates channel attention blocks to gather global image statistics, overlapping cross-attention to remove window boundaries, and an expanded window processing size. To maximize performance, the authors also implemented a simplified pre-training regimen using the ImageNet dataset focused exclusively on the target super-resolution task, followed by fine-tuning on domain-specific datasets.

The findings show substantial, measurable quality improvements over existing leading models. First, HAT activates input pixels across nearly the entire image, resolving texture errors and eliminating intermediate blocking artifacts. Second, the standard HAT architecture outperforms prior state-of-the-art models by 0.3 dB to 1.2 dB in Peak Signal-to-Noise Ratio (PSNR), yielding noticeably sharper repetitive structures and text. Third, the same-task pre-training approach delivers superior performance gains compared to complex multi-task pre-training strategies. Fourth, scaling the architecture into a larger variant (HAT-L) further expands the performance ceiling, while a lightweight variant (HAT-S) matches competitive computational footprints while still outperforming previous models.

These results demonstrate that Transformer performance in low-level vision tasks depends heavily on maximizing the spatial range of utilized input data rather than relying purely on localized attention mechanisms. For technical and operational decision-makers, this translates to superior visual fidelity and fewer reconstruction errors. Furthermore, the simplified same-task pre-training pipeline reduces architectural complexity compared to multi-task restoration models while maximizing data efficiency.

Organizations developing or deploying automated image enhancement systems should adopt hybrid attention architectures that combine channel and self-attention, while avoiding complex multi-task pre-training pipelines in favor of large-scale same-task pre-training. Depending on deployment constraints, teams can choose between the lightweight HAT-S variant for compute-constrained settings or HAT-L for maximum visual quality. Implementation plans should include pilot testing to calibrate fine-tuning learning rates and window overlap ratios to match target hardware and domain datasets.

While the findings demonstrate high confidence through comprehensive benchmarks and diagnostic attribution tools, the computational demands of large-scale pre-training and scaled variants remain significant. Decision-makers should account for increased training iterations and storage requirements when deploying these models at scale.

No sufficiently relevant recommendations were found.

Cover for Activating More Pixels in Image Super-Resolution Transformer

Abstract

Transformer-based methods have shown impressive performance in low-level vision tasks, such as image super-resolution. However, we find that these networks can only utilize a limited spatial range of input information through attribution analysis. This implies that the potential of Transformer is still not fully exploited in existing networks. In order to activate more input pixels for better reconstruction, we propose a novel Hybrid Attention Transformer (HAT). It combines both channel attention and window-based self-attention schemes, thus making use of their complementary advantages of being able to utilize global statistics and strong local fitting capability. Moreover, to better aggregate the cross-window information, we introduce an overlapping cross-attention module to enhance the interaction between neighboring window features. In the training stage, we additionally adopt a same-task pre-training strategy to exploit the potential of the model for further improvement. Extensive experiments show the effectiveness of the proposed modules, and we further scale up the model to demonstrate that the performance of this task can be greatly improved. Our overall method significantly outperforms the state-of-the-art methods by more than 1dB.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Deep Networks for Image SR
  • 2.2 Vision Transformer
  • 3 Methodology
  • 3.1 Motivation
  • 3.2 Network Architecture
  • 3.2.1 The Overall Structure
  • 3.2.2 Hybrid Attention Block (HAB)
  • 3.2.3 Overlapping Cross-Attention Block (OCAB)
  • 3.3 The Same-task Pre-training
  • 4 Experiments
  • 4.1 Experimental Setup
  • 4.2 Effects of different window sizes
  • 4.3 Ablation Study
  • 4.4 Comparison with State-of-the-Art Methods
  • 4.5 Study on the pre-training strategy
  • 5 Conclusion
  • References
  • A Training Details
  • B Analysis of Model Complexity
  • C More Visual Comparisons with LAM

Knowls

  1. Knowl 1 — HAT combines window self-attention with a channel-attention convolution branch

    model/method

    The Hybrid Attention Transformer (HAT) for single-image super-resolution has shallow feature extraction, deep feature extraction, and image reconstruction stages. A convolution first maps the low-resolution input to features with CC channels. Deep features are extracted by residual hybrid attention groups (RHAGs), followed by a 3×33\times3 convolution and a global residual connection from the shallow features; pixel shuffle then upsamples the fused features. The network is trained with an L1L_1 reconstruction loss.

    Each hybrid attention block (HAB) places a channel-attention block (CAB) in parallel with window-based multi-head self-attention, after the first layer normalization. For an input feature X∈RH×W×CX\in\mathbb{R}^{H\times W\times C}, the HAB computes

    X_N=\operatorname{LN}(X),\qquad X_M=(S)\operatorname{W\!-\!MSA}(X_N)+\alpha\operatorname{CAB}(X_N)+X, Y=MLP⁡(LN⁡(XM))+XM.Y=\operatorname{MLP}(\operatorname{LN}(X_M))+X_M.

    Here H,W,CH,W,C are spatial dimensions and channel count, LN⁡\operatorname{LN} is layer normalization, MLP⁡\operatorname{MLP} is a multilayer perceptron, and (S)\operatorname{W\!-\!MSA} denotes window self-attention with shifted windows applied at intervals. The scalar α\alpha scales the CAB branch. CAB uses two convolutions with GELU activation, compressing channels from CC to C/βC/\beta and then restoring CC, followed by channel attention to adaptively rescale channels. In the reported HAT configuration, α=0.01\alpha=0.01 and β=3\beta=3.

    For attention, each non-overlapping M×MM\times M window is treated as pixel tokens. Given query, key, and value matrices Q,K,VQ,K,V obtained by linear projections of a window feature, attention is SoftMax⁡(QKT/d+B)V\operatorname{SoftMax}(QK^{\mathsf T}/\sqrt d+B)V, where dd is the query/key dimension and BB is a relative-position bias. HAT uses M=16M=16 and a shifted-window offset of half the window size.

  2. Knowl 2 — Overlapping cross-attention connects neighboring window features

    model/method

    HAT adds an overlapping cross-attention block (OCAB) to directly exchange information across neighboring windows. For each input feature, the query features are divided into non-overlapping windows of size M×MM\times M, while the key and value features are divided into larger overlapping windows of size Mo×MoM_o\times M_o. The overlap is controlled by γ\gamma:

    Mo=(1+γ)M.M_o=(1+\gamma)M.

    The overlapping partition uses kernel size MoM_o and stride MM, with zero padding of γM/2\gamma M/2 to keep window dimensions consistent. Each query window attends to the keys and values in its larger corresponding window, using scaled dot-product attention with relative-position bias. Thus, query tokens remain within the original window while their candidate context includes pixels from adjacent windows. The OCAB also contains layer normalization, an MLP, and residual connections, and is placed in each RHAG after its HABs. The reported default is M=16M=16 and γ=0.5\gamma=0.5, giving Mo=24M_o=24.

  3. Knowl 3 — Same-task pre-training uses large-scale data before task-specific fine-tuning

    model/method

    HAT's same-task pre-training strategy trains the model on a large dataset for the exact super-resolution scale and degradation targeted at deployment, then fine-tunes it on the task-specific training set. For example, a model intended for ×4\times4 super-resolution is first trained for ×4\times4 super-resolution on ImageNet and then fine-tuned on DF2K, the union of DIV2K and Flickr2K. This differs from pre-training across different restoration tasks or across multiple degradation levels. The authors report that adequate pre-training iterations and a relatively small fine-tuning learning rate are important to obtain the benefit; they attribute the latter to avoiding overfitting to the smaller target dataset.

  4. Knowl 4 — Attribution analysis indicates that HAT uses a broader input region

    empirical result

    The paper uses local attribution maps (LAM) to examine which low-resolution input pixels contribute to reconstruction of a selected patch. Its diffusion index (DI) summarizes the spatial range of involved pixels, with a higher value indicating a broader range. The reported DI values are 4.02 for EDSR, 26.42 for RCAN, 19.52 for SwinIR, and 32.32 for HAT. Thus, in this analysis SwinIR's attributed range is narrower than RCAN's, while HAT has the broadest range among these four models. The authors also observe blocking artifacts in SwinIR intermediate features and report that these artifacts are substantially alleviated in HAT.

  5. Knowl 5 — Larger self-attention windows improve SwinIR super-resolution results

    data/table

    To isolate the effect of window size from HAT's new modules, the authors compare SwinIR models using 8×88\times8 and 16×1616\times16 self-attention windows. The larger window improves PSNR on each of the five benchmarks, with the largest reported gain on Urban100. The authors' visual comparison also shows a wider attributed input region for the 16×1616\times16 window. These results motivate using a window size of 16 in HAT.

    Window size Set5 Set14 BSD100 Urban100 Manga109
    8×88\times8 32.88 29.09 27.92 27.45 32.03
    16×1616\times16 32.97 29.12 27.95 27.81 32.15

    Entries are PSNR in dB for super-resolution; the comparison is conducted directly on SwinIR.

  6. Knowl 6 — CAB and OCAB each improve Urban100 reconstruction, with a further gain when combined

    empirical result

    An ablation on ×4\times4 super-resolution reports PSNR on Urban100 for a baseline without either proposed block and for variants adding OCAB, CAB, or both. The baseline scores 27.81 dB. Adding OCAB alone or CAB alone gives 27.91 dB in each case, and combining both gives 27.97 dB, a 0.16 dB increase over the baseline. The authors' attribution visualizations also show a broader range of involved pixels with the modules, particularly when CAB is included.

  7. Knowl 7 — CAB weighting and OCAB overlap have empirically preferred settings

    empirical result

    Urban100 ×4\times4 PSNR ablations identify the settings used in HAT. Within CAB, adding channel attention changes PSNR/SSIM from 27.92 dB/0.8362 without channel attention to 27.97 dB/0.8367 with it. The CAB scaling factor α\alpha gives 27.81 dB at 00, 27.86 dB at 11, 27.90 dB at 0.10.1, and 27.97 dB at 0.010.01; the authors select 0.010.01, noting that a small weight helps combine the CAB and self-attention branches. For OCAB, the tested overlap ratios γ=0,0.25,0.5,0.75\gamma=0,0.25,0.5,0.75 give 27.85, 27.81, 27.91, and 27.86 dB, respectively. The selected ratio is 0.50.5; the tested alternatives do not improve performance as much.

  8. Knowl 8 — HAT achieves strong Urban100 PSNR across scales, with further gains from pre-training and model scaling

    data/table

    The benchmark evaluation uses Set5, Set14, BSD100, Urban100, and Manga109, reporting PSNR and SSIM on the luminance channel. The table below gives Urban100 PSNR for representative models from the comparison: SwinIR, ImageNet-pre-trained EDT, HAT without pre-training, ImageNet-pre-trained HAT, and the larger ImageNet-pre-trained HAT-L. The HAT-L values are the best of these entries at every scale. Relative to pre-trained EDT, pre-trained HAT improves Urban100 PSNR by 0.54, 0.63, and 0.62 dB at ×2\times2, ×3\times3, and ×4\times4, respectively; HAT-L improves it by 0.82, 0.85, and 0.85 dB. The paper reports HAT's advantage across the benchmark datasets and also reports that HAT-S, a smaller variant with computation similar to SwinIR, outperforms SwinIR.

    Scale SwinIR EDT†^\dagger HAT HAT†^\dagger HAT-L†^\dagger
    ×2\times2 33.81 34.27 34.45 34.81 35.09
    ×3\times3 29.75 30.07 30.23 30.70 30.92
    ×4\times4 27.45 27.75 27.97 28.37 28.60

    Values are PSNR in dB on Urban100. The dagger denotes ImageNet pre-training followed by fine-tuning on DF2K. HAT-L doubles HAT's number of RHAGs, from 6 to 12.

  9. Knowl 9 — Same-task pre-training outperforms multi-related-task pre-training for HAT

    data/table

    For HAT on ×4\times4 super-resolution, the paper compares multi-related-task pre-training with same-task pre-training under the same training setting. Both approaches pre-train on the full ImageNet dataset and fine-tune on DF2K. Same-task pre-training gives higher PSNR at both the pre-training and fine-tuning stages on all three reported datasets, with the clearest fine-tuning difference on Urban100: 28.28 dB versus 28.21 dB.

    Strategy Stage Set5 Set14 Urban100
    Multi-related-task Pre-training 32.94 29.17 28.05
    Multi-related-task Fine-tuning 33.06 29.33 28.21
    Same-task Pre-training 33.02 29.20 28.11
    Same-task Fine-tuning 33.07 29.34 28.28

    Entries are PSNR in dB.

  10. Knowl 10 — Pre-training gains increase with model capacity in the reported super-resolution comparisons

    empirical result

    On ×4\times4 super-resolution, the paper compares four networks trained with and without the proposed same-task pre-training: SRResNet (1.5M parameters), RRDBNet (16.7M), SwinIR (11.9M), and HAT (20.8M). The figure reports PSNR gains from pre-training of +0.03+0.03, +0.12+0.12, +0.14+0.14, and +0.21+0.21 dB on Set14, respectively, and +0.09+0.09, +0.30+0.30, +0.38+0.38, and +0.40+0.40 dB on Urban100, respectively. All four models improve. In these comparisons, the larger network within each architecture family gains more, and HAT has the largest reported gain on both datasets. The authors interpret the larger Transformer gains, including SwinIR's gain relative to RRDBNet despite fewer parameters, as evidence that Transformer models benefit strongly from additional task-matched data.

Coverage note — The qualitative image comparisons are omitted because they illustrate, rather than add separate quantitative findings to, the reported attribution and benchmark results.

References

  1. 1.Hangbo Bao, Li Dong, and Furu Wei. Beit: Bert pre-training of image transformers. arXiv preprint arXiv:2106.08254, 2021.
  2. 2.Marco Bevilacqua, Aline Roumy, Christine Guillemot, and Marie Line Alberi-Morel. Low-complexity single-image super-resolution based on nonnegative neighbor embedding. 2012.
  3. 3.Hu Cao, Yueyue Wang, Joy Chen, Dongsheng Jiang, Xiaopeng Zhang, Qi Tian, and Manning Wang. Swin-unet: Unet-like pure transformer for medical image segmentation, 2021.
  4. 4.Jiezhang Cao, Yawei Li, Kai Zhang, and Luc Van Gool. Video super-resolution transformer, 2021.
  5. 5.Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. End-to-end object detection with transformers. In European conference on computer vision, pages 213–229. Springer, 2020.
  6. 6.Hanting Chen, Yunhe Wang, Tianyu Guo, Chang Xu, Yiping Deng, Zhenhua Liu, Siwei Ma, Chunjing Xu, Chao Xu, and Wen Gao. Pre-trained image processing transformer. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12299–12310, 2021.
  7. 7.Xiangxiang Chu, Zhi Tian, Yuqing Wang, Bo Zhang, Haibing Ren, Xiaolin Wei, Huaxia Xia, and Chunhua Shen. Twins: Revisiting the design of spatial attention in vision transformers. Advances in Neural Information Processing Systems, 34, 2021.
  8. 8.Tao Dai, Jianrui Cai, Yongbing Zhang, Shu-Tao Xia, and Lei Zhang. Second-order attention network for single image super-resolution. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 11065–11074, 2019.
  9. 9.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition, pages 248–255, 2009.
  10. 10.Chao Dong, Chen Change Loy, Kaiming He, and Xiaoou Tang. Learning a deep convolutional network for image super-resolution. In European conference on computer vision, pages 184–199. Springer, 2014.
  11. 11.Chao Dong, Chen Change Loy, Kaiming He, and Xiaoou Tang. Image super-resolution using deep convolutional networks. IEEE transactions on pattern analysis and machine intelligence, 38(2):295–307, 2015.
  12. 12.Chao Dong, Chen Change Loy, and Xiaoou Tang. Accelerating the super-resolution convolutional neural network. In European conference on computer vision, pages 391–407. Springer, 2016.
  13. 13.Xiaoyi Dong, Jianmin Bao, Dongdong Chen, Weiming Zhang, Nenghai Yu, Lu Yuan, Dong Chen, and Baining Guo. Cswin transformer: A general vision transformer backbone with cross-shaped windows. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12124–12134, 2022.
  14. 14.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale, 2020.
  15. 15.Jinjin Gu and Chao Dong. Interpreting super-resolution networks with local attribution maps. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9199–9208, 2021.
  16. 16.Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollar, and Ross Girshick. Masked autoencoders are scalable vision learners. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 16000–16009, 2022.
  17. 17.Dan Hendrycks and Kevin Gimpel. Gaussian error linear units (gelus), 2016.
  18. 18.Gao Huang, Yulin Wang, Kangchen Lv, Haojun Jiang, Wenhui Huang, Pengfei Qi, and Shiji Song. Glance and focus networks for dynamic visual recognition, 2022.
  19. 19.Jia-Bin Huang, Abhishek Singh, and Narendra Ahuja. Single image super-resolution from transformed self-exemplars. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 5197–5206, 2015.
  20. 20.Zilong Huang, Youcheng Ben, Guozhong Luo, Pei Cheng, Gang Yu, and Bin Fu. Shuffle transformer: Rethinking spatial shuffle for vision transformer, 2021.
  21. 21.Jiwon Kim, Jung Kwon Lee, and Kyoung Mu Lee. Accurate image super-resolution using very deep convolutional networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 1646–1654, 2016.
  22. 22.Jiwon Kim, Jung Kwon Lee, and Kyoung Mu Lee. Deeply-recursive convolutional network for image super-resolution. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 1637–1645, 2016.
  23. 23.Xiangtao Kong, Xina Liu, Jinjin Gu, Yu Qiao, and Chao Dong. Reflash dropout in image super-resolution. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6002–6012, 2022.
  24. 24.Xiangtao Kong, Hengyuan Zhao, Yu Qiao, and Chao Dong. Classsr: A general framework to accelerate super-resolution networks by data characteristic. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 12016–12025, June 2021.
  25. 25.Christian Ledig, Lucas Theis, Ferenc Huszar, Jose Caballero, Andrew Cunningham, Alejandro Acosta, Andrew Aitken, Alykhan Tejani, Johannes Totz, Zehan Wang, et al. Photorealistic single image super-resolution using a generative adversarial network. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 4681–4690, 2017.
  26. 26.Kunchang Li, Yali Wang, Junhao Zhang, Peng Gao, Guanglu Song, Yu Liu, Hongsheng Li, and Yu Qiao. Uniformer: Unifying convolution and self-attention for visual recognition, 2022.
  27. 27.Wenbo Li, Xin Lu, Jiangbo Lu, Xiangyu Zhang, and Jiaya Jia. On efficient transformer and image pre-training for low-level vision, 2021.
  28. 28.Yawei Li, Kai Zhang, Jiezhang Cao, Radu Timofte, and Luc Van Gool. Localvit: Bringing locality to vision transformers, 2021.
  29. 29.Zheyuan Li, Yingqi Liu, Xiangyu Chen, Haoming Cai, Jinjin Gu, Yu Qiao, and Chao Dong. Blueprint separable residual network for efficient image super-resolution. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, pages 833–843, June 2022.
  30. 30.Jingyun Liang, Jiezhang Cao, Yuchen Fan, Kai Zhang, Rakesh Ranjan, Yawei Li, Radu Timofte, and Luc Van Gool. Vrt: A video restoration transformer, 2022.
  31. 31.Jingyun Liang, Jiezhang Cao, Guolei Sun, Kai Zhang, Luc Van Gool, and Radu Timofte. Swinir: Image restoration using swin transformer. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 1833–1844, 2021.
  32. 32.Bee Lim, Sanghyun Son, Heewon Kim, Seungjun Nah, and Kyoung Mu Lee. Enhanced deep residual networks for single image super-resolution. In Proceedings of the IEEE conference on computer vision and pattern recognition workshops, pages 136–144, 2017.
  33. 33.Bee Lim, Sanghyun Son, Heewon Kim, Seungjun Nah, and Kyoung Mu Lee. Enhanced deep residual networks for single image super-resolution. In Proceedings of the IEEE conference on computer vision and pattern recognition workshops, pages 136–144, 2017.
  34. 34.Zudi Lin, Prateek Garg, Atmadeep Banerjee, Salma Abdel Magid, Deqing Sun, Yulun Zhang, Luc Van Gool, Donglai Wei, and Hanspeter Pfister. Revisiting rcan: Improved training for image super-resolution, 2022.
  35. 35.Ding Liu, Bihan Wen, Yuchen Fan, Chen Change Loy, and Thomas S Huang. Non-local recurrent network for image restoration. Advances in neural information processing systems, 31, 2018.
  36. 36.Li Liu, Wanli Ouyang, Xiaogang Wang, Paul Fieguth, Jie Chen, Xinwang Liu, and Matti Pietikainen. Deep learning for generic object detection: A survey. International journal of computer vision, 128(2):261–318, 2020.
  37. 37.Yihao Liu, Anran Liu, Jinjin Gu, Zhipeng Zhang, Wenhao Wu, Yu Qiao, and Chao Dong. Discovering” semantics” in super-resolution networks, 2021.
  38. 38.Yihao Liu, Hengyuan Zhao, Jinjin Gu, Yu Qiao, and Chao Dong. Evaluating the generalization ability of super-resolution networks. arXiv preprint arXiv:2205.07019, 2022.
  39. 39.Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 10012–10022, 2021.
  40. 40.David Martin, Charless Fowlkes, Doron Tal, and Jitendra Malik. A database of human segmented natural images and its application to evaluating segmentation algorithms and measuring ecological statistics. In Proceedings Eighth IEEE International Conference on Computer Vision. ICCV 2001, volume 2, pages 416–423. IEEE, 2001.
  41. 41.Yusuke Matsui, Kota Ito, Yuji Aramaki, Azuma Fujimoto, Toru Ogawa, Toshihiko Yamasaki, and Kiyoharu Aizawa. Sketch-based manga retrieval using manga109 dataset. Multimedia Tools and Applications, 76(20):21811–21838, 2017.
  42. 42.Yiqun Mei, Yuchen Fan, and Yuqian Zhou. Image super-resolution with non-local sparse attention. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3517–3526, 2021.
  43. 43.Ben Niu, Weilei Wen, Wenqi Ren, Xiangde Zhang, Lianping Yang, Shuzhen Wang, Kaihao Zhang, Xiaochun Cao, and Haifeng Shen. Single image super-resolution via a holistic attention network. In European conference on computer vision, pages 191–207. Springer, 2020.
  44. 44.Krushi Patel, Andres M Bur, Fengjun Li, and Guanghui Wang. Aggregating global features into local vision transformer, 2022.
  45. 45.Maithra Raghu, Thomas Unterthiner, Simon Kornblith, Chiyuan Zhang, and Alexey Dosovitskiy. Do vision transformers see like convolutional neural networks? Advances in Neural Information Processing Systems, 34, 2021.
  46. 46.Prajit Ramachandran, Niki Parmar, Ashish Vaswani, Irwan Bello, Anselm Levskaya, and Jon Shlens. Studying standalone self-attention in vision models. 2019.
  47. 47.Wenzhe Shi, Jose Caballero, Ferenc Huszar, Johannes Totz, Andrew P Aitken, Rob Bishop, Daniel Rueckert, and Zehan Wang. Real-time single image and video super-resolution using an efficient sub-pixel convolutional neural network. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 1874–1883, 2016.
  48. 48.Ying Tai, Jian Yang, and Xiaoming Liu. Image super-resolution via deep recursive residual network. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 3147–3155, 2017.
  49. 49.Radu Timofte, Eirikur Agustsson, Luc Van Gool, Ming-Hsuan Yang, and Lei Zhang. Ntire 2017 challenge on single image super-resolution: Methods and results. In Proceedings of the IEEE conference on computer vision and pattern recognition workshops, pages 114–125, 2017.
  50. 50.Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Herve Jegou. Training data-efficient image transformers & distillation through attention. In International Conference on Machine Learning, pages 10347–10357. PMLR, 2021.
  51. 51.Zhengzhong Tu, Hossein Talebi, Han Zhang, Feng Yang, Peyman Milanfar, Alan Bovik, and Yinxiao Li. Maxim: Multi-axis mlp for image processing. CVPR, 2022.
  52. 52.Ashish Vaswani, Prajit Ramachandran, Aravind Srinivas, Niki Parmar, Blake Hechtman, and Jonathon Shlens. Scaling local self-attention for parameter efficient visual backbones. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12894–12904, 2021.
  53. 53.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in neural information processing systems, 30, 2017.
  54. 54.Wenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan, Kaitao Song, Ding Liang, Tong Lu, Ping Luo, and Ling Shao. Pyramid vision transformer: A versatile backbone for dense prediction without convolutions. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 568–578, 2021.
  55. 55.Xintao Wang, Liangbin Xie, Chao Dong, and Ying Shan. Real-esrgan: Training real-world blind super-resolution with pure synthetic data. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 1905–1914, 2021.
  56. 56.Xintao Wang, Ke Yu, Shixiang Wu, Jinjin Gu, Yihao Liu, Chao Dong, Yu Qiao, and Chen Change Loy. Esrgan: Enhanced super-resolution generative adversarial networks. In Proceedings of the European conference on computer vision (ECCV) workshops, pages 0–0, 2018.
  57. 57.Zhendong Wang, Xiaodong Cun, Jianmin Bao, Wengang Zhou, Jianzhuang Liu, and Houqiang Li. Uformer: A general u-shaped transformer for image restoration. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 17683–17693, 2022.
  58. 58.Bichen Wu, Chenfeng Xu, Xiaoliang Dai, Alvin Wan, Peizhao Zhang, Zhicheng Yan, Masayoshi Tomizuka, Joseph Gonzalez, Kurt Keutzer, and Peter Vajda. Visual transformers: Token-based image representation and processing for computer vision, 2020.
  59. 59.Haiping Wu, Bin Xiao, Noel Codella, Mengchen Liu, Xiyang Dai, Lu Yuan, and Lei Zhang. Cvt: Introducing convolutions to vision transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 22–31, 2021.
  60. 60.Sitong Wu, Tianyi Wu, Haoru Tan, and Guodong Guo. Pale transformer: A general vision transformer backbone with pale-shaped attention. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 36, pages 2731–2739, 2022.
  61. 61.Tete Xiao, Piotr Dollar, Mannat Singh, Eric Mintun, Trevor Darrell, and Ross Girshick. Early convolutions help transformers see better. Advances in Neural Information Processing Systems, 34, 2021.
  62. 62.Liangbin Xie, Xintao Wang, Chao Dong, Zhongang Qi, and Ying Shan. Finding discriminative filters for specific degradations in blind super-resolution. Advances in Neural Information Processing Systems, 34, 2021.
  63. 63.Kun Yuan, Shaopeng Guo, Ziwei Liu, Aojun Zhou, Fengwei Yu, and Wei Wu. Incorporating convolution designs into visual transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 579–588, 2021.
  64. 64.Yuhui Yuan, Rao Fu, Lang Huang, Weihong Lin, Chao Zhang, Xilin Chen, and Jingdong Wang. Hrformer: High-resolution vision transformer for dense predict. Advances in Neural Information Processing Systems, 34:7281–7293, 2021.
  65. 65.Syed Waqas Zamir, Aditya Arora, Salman Khan, Munawar Hayat, Fahad Shahbaz Khan, and Ming-Hsuan Yang. Restormer: Efficient transformer for high-resolution image restoration. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5728–5739, 2022.
  66. 66.Roman Zeyde, Michael Elad, and Matan Protter. On single image scale-up using sparse-representations. In International conference on curves and surfaces, pages 711–730. Springer, 2010.
  67. 67.Wenlong Zhang, Yihao Liu, Chao Dong, and Yu Qiao. Ranksrgan: Generative adversarial networks with ranker for image super-resolution. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 3096–3105, 2019.
  68. 68.Yulun Zhang, Kunpeng Li, Kai Li, Lichen Wang, Bineng Zhong, and Yun Fu. Image super-resolution using very deep residual channel attention networks. In Proceedings of the European conference on computer vision (ECCV), pages 286–301, 2018.
  69. 69.Yulun Zhang, Kunpeng Li, Kai Li, Bineng Zhong, and Yun Fu. Residual non-local attention networks for image restoration, 2019.
  70. 70.Yulun Zhang, Yapeng Tian, Yu Kong, Bineng Zhong, and Yun Fu. Residual dense network for image super-resolution. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2472–2481, 2018.
  71. 71.Yucheng Zhao, Guangting Wang, Chuanxin Tang, Chong Luo, Wenjun Zeng, and Zheng-Jun Zha. A battle of network structures: An empirical study of cnn, transformer, and mlp, 2021.
  72. 72.Shangchen Zhou, Jiawei Zhang, Wangmeng Zuo, and Chen Change Loy. Cross-scale internal graph neural network for image super-resolution. Advances in neural information processing systems, 33:3499–3509, 2020.

Citation

MLA
Chen, X., et al. “Activating More Pixels in Image Super-Resolution Transformer”. arXiv, 2022, http://arxiv.org/abs/2205.04437v3.
APA
Chen, X., Wang, X., Zhou, J., Qiao, Y., & Dong, C. (2022). Activating More Pixels in Image Super-Resolution Transformer. arXiv. http://arxiv.org/abs/2205.04437v3
Chicago
Chen, X., X. Wang, J. Zhou, Y. Qiao, and C. Dong. 2022. “Activating More Pixels in Image Super-Resolution Transformer”. arXiv. http://arxiv.org/abs/2205.04437v3.
Harvard
Chen, X. et al. (2022) “Activating More Pixels in Image Super-Resolution Transformer”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2205.04437v3.
Vancouver
1. Chen X, Wang X, Zhou J, Qiao Y, Dong C (2022) Activating More Pixels in Image Super-Resolution Transformer. arXiv

BibTeX

@article{chen2022activating,
  title = {Activating More Pixels in Image Super-Resolution Transformer},
  author = {Chen, Xiangyu and Wang, Xintao and Zhou, Jiantao and Qiao, Yu and Dong, Chao},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2205.04437v3},
  eprint = {2205.04437}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE