Built independently by an author, for readers. Read the story and support ChapterPal

keyword

low-level vision tasks

Low-level vision tasks are computer vision and image processing operations that analyze, manipulate, or reconstruct visual information directly at the pixel or local feature level without requiring high-level semantic understanding of a scene. Unlike high-level vision tasks such as object detection or scene classification that interpret what an image depicts, low-level tasks focus on fundamental visual characteristics like color, intensity, edges, and textures. Common examples include image restoration and enhancement procedures such as super-resolution, denoising, deblurring, deraining, and inpainting. The output of these operations is typically a modified, filtered, or restored image rather than abstract labels or descriptions, serving as fundamental steps for improving visual quality or preparing visual data for downstream processing.

2 items

Activating More Pixels in Image Super-Resolution Transformer

Activating More Pixels in Image Super-Resolution Transformer

Xiangyu Chen, Xintao Wang, Jiantao Zhou, Yu Qiao, Chao Dong

OrganizationsChinese Academy of SciencesShanghai Artificial Intelligence LaboratoryShenzhen Institute of Advanced Technology, Chinese Academy of SciencesTencentUniversity of Macau

Why you should read this

Proposes a Hybrid Attention Transformer that combines channel and window self-attention with overlapping cross-attention to expand the spatial range of activated pixels, outperforming existing super-resolution methods by over 1 dB.

Transformer-based methods have shown impressive performance in low-level vision tasks, such as image super-resolution. However, we find that these networks can only utilize a limited spatial range of input information through attribution analysis. This implies that the potential of Transformer is still not fully exploited in existing networks. In order to activate more input pixels for better reconstruction, we propose a novel Hybrid Attention Transformer (HAT). It combines both channel attention and window-based self-attention schemes, thus making use of their complementary advantages of being able to utilize global statistics and strong local fitting capability. Moreover, to better aggregate the cross-window information, we introduce an overlapping cross-attention module to enhance the interaction between neighboring window features. In the training stage, we additionally adopt a same-task pre-training strategy to exploit the potential of the model for further improvement. Extensive experiments show the effectiveness of the proposed modules, and we further scale up the model to demonstrate that the performance of this task can be greatly improved. Our overall method significantly outperforms the state-of-the-art methods by more than 1dB.

Added

2026-10-04

Pre-Trained Image Processing Transformer

Pre-Trained Image Processing Transformer

Hanting Chen, Yunhe Wang, Tianyu Guo, Chang Xu, Yiping Deng, Zhenhua Liu, Siwei Ma, Chunjing Xu, Chao Xu, Wen Gao

OrganizationsHuaweiPeking UniversityPeng Cheng LaboratoryUniversity of Sydney

Why you should read this

Introduces the Image Processing Transformer (IPT), a unified pre-trained architecture trained on large-scale corrupted datasets with contrastive learning to outperform specialized models across multiple low-level vision tasks such as super-resolution, denoising, and deraining.

As the computing power of modern hardware is increasing strongly, pre-trained deep learning models (e.g., BERT, GPT-3) learned on large-scale datasets have shown their effectiveness over conventional methods. The big progress is mainly contributed to the representation ability of transformer and its variant architectures. In this paper, we study the low-level computer vision task (e.g., denoising, super-resolution and deraining) and develop a new pre-trained model, namely, image processing transformer (IPT). To maximally excavate the capability of transformer, we present to utilize the well-known ImageNet benchmark for generating a large amount of corrupted image pairs. The IPT model is trained on these images with multi-heads and multi-tails. In addition, the contrastive learning is introduced for well adapting to different image processing tasks. The pre-trained model can therefore efficiently employed on desired task after fine-tuning. With only one pre-trained model, IPT outperforms the current state-of-the-art methods on various low-level benchmarks. Code is available at this https URL and this https URL

Added

2026-09-15