Built independently by an author, for readers. Read the story and support ChapterPal

keyword

Pre-Trained Image Processing Transformer

A Pre-Trained Image Processing Transformer is a deep learning framework based on the transformer architecture that is pre-trained on large-scale datasets to address low-level computer vision and image restoration tasks. Rather than relying solely on task-specific convolutional networks, it leverages self-attention mechanisms alongside specialized input and output modules to handle multiple image restoration objectives, such as denoising, super-resolution, and deraining. Pre-training on extensive sets of corrupted image pairs allows the model to learn generalizable visual representations and spatial relationships, enabling it to adapt efficiently to various downstream image enhancement tasks through fine-tuning.

1 item

Pre-Trained Image Processing Transformer

Pre-Trained Image Processing Transformer

Hanting Chen, Yunhe Wang, Tianyu Guo, Chang Xu, Yiping Deng, Zhenhua Liu, Siwei Ma, Chunjing Xu, Chao Xu, Wen Gao

OrganizationsHuaweiPeking UniversityPeng Cheng LaboratoryUniversity of Sydney

Why you should read this

Introduces the Image Processing Transformer (IPT), a unified pre-trained architecture trained on large-scale corrupted datasets with contrastive learning to outperform specialized models across multiple low-level vision tasks such as super-resolution, denoising, and deraining.

As the computing power of modern hardware is increasing strongly, pre-trained deep learning models (e.g., BERT, GPT-3) learned on large-scale datasets have shown their effectiveness over conventional methods. The big progress is mainly contributed to the representation ability of transformer and its variant architectures. In this paper, we study the low-level computer vision task (e.g., denoising, super-resolution and deraining) and develop a new pre-trained model, namely, image processing transformer (IPT). To maximally excavate the capability of transformer, we present to utilize the well-known ImageNet benchmark for generating a large amount of corrupted image pairs. The IPT model is trained on these images with multi-heads and multi-tails. In addition, the contrastive learning is introduced for well adapting to different image processing tasks. The pre-trained model can therefore efficiently employed on desired task after fine-tuning. With only one pre-trained model, IPT outperforms the current state-of-the-art methods on various low-level benchmarks. Code is available at this https URL and this https URL

Added

2026-09-15