Cross Aggregation Transformer for Image Restoration
Zheng ChenYulun ZhangJinjin GuYongbing ZhangLinghe KongXin Yuan
Proposes the Cross Aggregation Transformer for image restoration, utilizing parallel horizontal and vertical rectangle-window self-attention alongside an axial-shift operation to expand receptive fields and capture long-range dependencies with linear computational complexity.
Image restoration is a foundational visual computing problem that involves recovering high-quality images from degraded, low-quality inputs. While modern deep learning architectures—specifically Vision Transformers—excel at capturing global context, their high computational complexity has traditionally forced models to use small, square processing windows. These constrained windows prevent the model from establishing broad image context and understanding directional structures like parallel lines or repetitive urban patterns, which limits overall restoration fidelity.
To overcome these constraints, the article introduces the Cross Aggregation Transformer, a novel neural network architecture designed to process high-resolution images efficiently. The primary objective is to demonstrate that aggregating features across perpendicular rectangular windows and coupling global attention with local convolutional operations significantly improves image restoration quality across multiple tasks without excessive computational cost.
To evaluate this architecture, the authors conducted extensive experiments across three primary domains: image super-resolution, JPEG compression artifact removal, and real-world photographic denoising. The evaluation tested the model against prominent standard benchmark datasets, including Urban100, Classic5, LIVE1, SIDD, and DND. The model design incorporates two primary structural configurations—a regular rectangle-window variant and an axial full-stripe variant—and compares their performance against prior state-of-the-art methods in terms of distortion metrics, structural similarity, and computational efficiency.
The findings show that the proposed architecture consistently outperforms existing convolutional and transformer-based methods. In image super-resolution, the model achieved substantial gains over leading competitors, improving signal-to-noise metrics on challenging urban scenes by up to 0.45 dB. In JPEG artifact reduction and real-world image denoising, the framework consistently achieved superior or highly competitive visual sharpness and reconstruction accuracy while using fewer parameters than prominent alternatives. Furthermore, the tests revealed that the added local feature module increased computational load by less than 0.32% while delivering consistent measurable quality gains.
These results demonstrate that rectangular attention windows and global-local feature coupling solve the trade-off between computational efficiency and large-scale image context. For organizations deploying imaging systems, this approach enables higher-fidelity visual reconstruction at manageable computational and parameter costs, reducing the hardware footprint needed to process high-resolution degraded imagery.
Organizations evaluating image restoration pipelines should consider adopting rectangular-window attention architectures for tasks where directional structures and fine details are critical. Teams facing strict computational constraints can deploy the regular rectangular-window variant to match baseline resource limits while still gaining fidelity, whereas applications prioritizing maximum output quality should adopt the axial stripe configuration. Further testing and validation on proprietary operational datasets are recommended before full deployment.
The article notes certain design boundaries, showing that using overly narrow axial stripes can degrade performance by capturing insufficient context or introducing noise. Additionally, while the results provide high confidence across standard academic benchmarks, the authors did not report multi-run statistical error bars, meaning performance under diverse production conditions should be verified empirically.
- Paper: Swin Transformer: Hierarchical Vision Transformer using Shifted Windows, Ze Liu et al. (2021). It introduces shifted local-window self-attention, establishing the foundational windowing mechanism whose limited inter-window interaction CAT explicitly sets out to overcome.
- Paper: SwinIR: Image Restoration Using Swin Transformer, Jingyun Liang et al. (2021). It provides the baseline paradigm of applying window-based Swin Transformers to image restoration that CAT directly builds upon and enhances.
- Paper: CSWin Transformer: A General Vision Transformer Backbone with Cross-Shaped Windows, Xiaoyi Dong et al. (2021). It introduces cross-shaped horizontal and vertical stripe window attention, inspiring the rectangular window self-attention design used in CAT.
- Paper: Uformer: A General U-Shaped Transformer for Image Restoration, Zhendong Wang et al. (2021). It establishes a benchmark architecture for window-based transformer restoration and integrating local context modules into transformer blocks.
- Paper: CvT: Introducing Convolutions to Vision Transformers, Haiping Wu et al. (2021). It pioneered the strategy of injecting convolutional inductive biases into vision transformers to couple local and global features.
- Paper: Pre-Trained Image Processing Transformer, Hanting Chen et al. (2020). It pioneered the application of vision transformer architectures to low-level image restoration tasks.
- Paper: TransNeXt: Robust Foveal Visual Perception for Vision Transformers, Dai Shi (2024). It further addresses the depth degradation and visual artifact limitations of localized window-based vision transformers through biomimetic aggregated attention.
- Paper: Efficient Multi-Scale Attention Module with Cross-Spatial Learning, Daliang Ouyang et al. (2023). It builds on multi-scale spatial and channel attention principles to achieve efficient cross-spatial feature learning across parallel branches.
