Cross Aggregation Transformer for Image Restoration

Zheng ChenYulun ZhangJinjin GuYongbing ZhangLinghe KongXin Yuan

article2022NeurIPS225 citations

Proposes the Cross Aggregation Transformer for image restoration, utilizing parallel horizontal and vertical rectangle-window self-attention alongside an axial-shift operation to expand receptive fields and capture long-range dependencies with linear computational complexity.

Listen

Image restoration is a foundational visual computing problem that involves recovering high-quality images from degraded, low-quality inputs. While modern deep learning architectures—specifically Vision Transformers—excel at capturing global context, their high computational complexity has traditionally forced models to use small, square processing windows. These constrained windows prevent the model from establishing broad image context and understanding directional structures like parallel lines or repetitive urban patterns, which limits overall restoration fidelity.

To overcome these constraints, the article introduces the Cross Aggregation Transformer, a novel neural network architecture designed to process high-resolution images efficiently. The primary objective is to demonstrate that aggregating features across perpendicular rectangular windows and coupling global attention with local convolutional operations significantly improves image restoration quality across multiple tasks without excessive computational cost.

To evaluate this architecture, the authors conducted extensive experiments across three primary domains: image super-resolution, JPEG compression artifact removal, and real-world photographic denoising. The evaluation tested the model against prominent standard benchmark datasets, including Urban100, Classic5, LIVE1, SIDD, and DND. The model design incorporates two primary structural configurations—a regular rectangle-window variant and an axial full-stripe variant—and compares their performance against prior state-of-the-art methods in terms of distortion metrics, structural similarity, and computational efficiency.

The findings show that the proposed architecture consistently outperforms existing convolutional and transformer-based methods. In image super-resolution, the model achieved substantial gains over leading competitors, improving signal-to-noise metrics on challenging urban scenes by up to 0.45 dB. In JPEG artifact reduction and real-world image denoising, the framework consistently achieved superior or highly competitive visual sharpness and reconstruction accuracy while using fewer parameters than prominent alternatives. Furthermore, the tests revealed that the added local feature module increased computational load by less than 0.32% while delivering consistent measurable quality gains.

These results demonstrate that rectangular attention windows and global-local feature coupling solve the trade-off between computational efficiency and large-scale image context. For organizations deploying imaging systems, this approach enables higher-fidelity visual reconstruction at manageable computational and parameter costs, reducing the hardware footprint needed to process high-resolution degraded imagery.

Organizations evaluating image restoration pipelines should consider adopting rectangular-window attention architectures for tasks where directional structures and fine details are critical. Teams facing strict computational constraints can deploy the regular rectangular-window variant to match baseline resource limits while still gaining fidelity, whereas applications prioritizing maximum output quality should adopt the axial stripe configuration. Further testing and validation on proprietary operational datasets are recommended before full deployment.

The article notes certain design boundaries, showing that using overly narrow axial stripes can degrade performance by capturing insufficient context or introducing noise. Additionally, while the results provide high confidence across standard academic benchmarks, the authors did not report multi-run statistical error bars, meaning performance under diverse production conditions should be verified empirically.

Cover for Cross Aggregation Transformer for Image Restoration

Abstract

Recently, Transformer architecture has been introduced into image restoration to replace convolution neural network (CNN) with surprising results. Considering the high computational complexity of Transformer with global attention, some methods use the local square window to limit the scope of self-attention. However, these methods lack direct interaction among different windows, which limits the establishment of long-range dependencies. To address the above issue, we propose a new image restoration model, Cross Aggregation Transformer (CAT). The core of our CAT is the Rectangle-Window Self-Attention (Rwin-SA), which utilizes horizontal and vertical rectangle window attention in different heads parallelly to expand the attention area and aggregate the features cross different windows. We also introduce the Axial-Shift operation for different window interactions. Furthermore, we propose the Locality Complementary Module to complement the self-attention mechanism, which incorporates the inductive bias of CNN (e.g., translation invariance and locality) into Transformer, enabling global-local coupling. Extensive experiments demonstrate that our CAT outperforms recent state-of-the-art methods on several image restoration applications. The code and models are available at https://github.com/zhengchen1999/CAT.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Method
  • 3.1 CAT Architecture
  • 3.2 Cross Aggregation Transformer Block
  • 4 Experiments
  • 4.1 Experimental Settings
  • 4.2 Ablation Study
  • 4.3 Image Super-Resolution
  • 4.4 JPEG Compression Artifact Reduction
  • 4.5 Real Image Denoising
  • 4.6 Model Size Analyses
  • 5 Conclusion
  • References
  • Checklist

Knowls

  1. Knowl 1 — Cross Aggregation Transformer (CAT) Architecture for Image Restoration

    model/method

    Cross Aggregation Transformer (CAT) is an image restoration network structured into three sequential stages: shallow feature extraction, deep feature extraction, and image reconstruction.

    1. Shallow Feature Extraction: Given a degraded low-quality input image ILQimesRH×W×CinI_{LQ} imes \mathbb{R}^{H \times W \times C_{in}} (where HH, WW, and CinC_{in} represent height, width, and input channel count), a single 3×33 \times 3 convolution layer extracts shallow feature representations F0∈RH×W×CF_0 \in \mathbb{R}^{H \times W \times C}, where CC is the feature channel dimension.

    2. Deep Feature Extraction: The shallow features F0F_0 pass through N1N_1 Residual Groups (RGs) followed by a 3×33 \times 3 convolution layer and a global residual connection to yield deep features FDF∈RH×W×CF_{DF} \in \mathbb{R}^{H \times W \times C}. Each RG contains N2N_2 Cross Aggregation Transformer Blocks (CATBs) followed by a 3×33 \times 3 convolution layer and a local residual skip connection. The default configuration uses N1=6N_1 = 6 residual groups, N2=6N_2 = 6 CATBs per group, feature dimension C=180C = 180, and 6 attention heads.

    3. Task-Specific Reconstruction Module: To generate the high-quality restored image IHQ∈RH×W×CoutI_{HQ} \in \mathbb{R}^{H \times W \times C_{out}} from FDFF_{DF}:

    • Image Super-Resolution (SR): A sub-pixel convolution upsampling layer is bounded by pre- and post-upsampling 3×33 \times 3 convolution layers.
    • JPEG Compression Artifact Reduction: A single 3×33 \times 3 convolution maps channels from CC to CoutC_{out}, and the input image ILQI_{LQ} is added via a global residual connection (IHQ=ILQ+Conv(FDF)I_{HQ} = I_{LQ} + \text{Conv}(F_{DF})).
    • Real Image Denoising: The CATB modules are arranged inside a 4-level encoder-decoder U-Net architecture with block counts [4,6,6,8][4, 6, 6, 8] and attention head allocations [2,2,4,8][2, 2, 4, 8] across levels 1 to 4.
  2. Knowl 2 — Rectangle-Window Self-Attention (Rwin-SA)

    model/method

    Rectangle-Window Self-Attention (Rwin-SA) is a local multi-head self-attention mechanism that splits attention heads across orthogonal rectangular windows to expand the receptive field and enable cross-window feature aggregation without increasing computational cost over square windows.

    Given an input feature map X∈RH×W×CX \in \mathbb{R}^{H \times W \times C}, an even number of attention heads MM is partitioned into two equal groups of size M/2M/2. For each head in the first group, XX is split into non-overlapping horizontal rectangular windows (sh<swsh < sw, denoted H-Rwin) of dimension sh×swsh \times sw. For each head in the second group, XX is split into vertical rectangular windows (sh>swsh > sw, denoted V-Rwin) of dimension sw×shsw \times sh.

    For the ii-th rectangular window feature Xi∈R(sh×sw)×CX_i \in \mathbb{R}^{(sh \times sw) \times C} (where i=1,…,HWsh⋅swi = 1, \dots, \frac{HW}{sh \cdot sw}) in head mm, the queries, keys, and values are computed using projection matrices WQm,WKm,WVm∈RC×dW_Q^m, W_K^m, W_V^m \in \mathbb{R}^{C \times d} with head dimension d=D=C/Md = D = C/M: (Qim,Kim,Vim)=(XiWQm,XiWKm,XiWVm)(Q_i^m, K_i^m, V_i^m) = (X_i W_Q^m, X_i W_K^m, X_i W_V^m)

    The self-attention feature Yim∈R(sh×sw)×DY_i^m \in \mathbb{R}^{(sh \times sw) \times D} is formulated as: Yim=Attention(Qim,Kim,Vim)=SoftMax(Qim(Kim)Td+B)VimY_i^m = \text{Attention}(Q_i^m, K_i^m, V_i^m) = \text{SoftMax}\left(\frac{Q_i^m (K_i^m)^T}{\sqrt{d}} + B\right) V_i^m where BB is a dynamic relative position bias matrix. The window outputs are rearranged into the full spatial feature Ym∈RH×W×DY^m \in \mathbb{R}^{H \times W \times D}.

    The outputs of all MM heads (the first M/2M/2 using H-Rwin and the remaining M/2M/2 using V-Rwin) are concatenated along the channel dimension and projected by WP∈RC×CW_P \in \mathbb{R}^{C \times C}: Rwin-SA(X)=Concat(Y1,Y2,…,YM)WP\text{Rwin-SA}(X) = \text{Concat}(Y^1, Y^2, \dots, Y^M) W_P

  3. Knowl 3 — Computational Complexity of Regular-Rwin and Axial-Rwin

    theoretical result

    For an input feature map of height HH, width WW, and channel dimension CC, the computational complexity of rectangular window self-attention depends on whether fixed-size rectangular windows (regular-Rwin) or image-adaptive full-axis stripe windows (axial-Rwin) are used.

    1. Regular Rectangle-Window Attention (regular-Rwin): Uses fixed window dimensions sh×swsh \times sw for horizontal windows and sw×shsw \times sh for vertical windows. Its computational complexity is: O(regular-Rwin)=HWC×(4C+2sh×sw)\mathcal{O}(\text{regular-Rwin}) = H W C \times (4C + 2sh \times sw) For square images (H=WH = W), the asymptotic complexity is O(H2)\mathcal{O}(H^2), scaling linearly with the total pixel count HWHW.

    2. Axial Rectangle-Window Attention (axial-Rwin): Sets one side length of the rectangular window to match the image dimensions HH or WW, while the short-side length is fixed to a hyperparameter slsl (yielding horizontal windows of size sl×Wsl \times W and vertical windows of size H×slH \times sl). Its computational complexity is: O(axial-Rwin)=HWC×(4C+sl×H+sl×W)\mathcal{O}(\text{axial-Rwin}) = H W C \times (4C + sl \times H + sl \times W) For square inputs (H=WH = W), the complexity scales as O(H3)\mathcal{O}(H^3), which is higher than regular-Rwin's O(H2)\mathcal{O}(H^2) but strictly lower than full global self-attention's quadratic-with-area complexity of O(H4)\mathcal{O}(H^4).

  4. Knowl 4 — Axial-Shift Operation for Window Interaction

    model/method

    The Axial-Shift Operation is a shifted window partitioning strategy applied alternately between consecutive Cross Aggregation Transformer Blocks to enable information exchange across isolated rectangular windows.

    Axial-shift splits into two directional operations tailored to rectangular geometry:

    • H-Shift: Moves the window partition down and left by sh2\frac{sh}{2} and sw2\frac{sw}{2} pixels for the horizontal rectangle window (H-Rwin).
    • V-Shift: Moves the window partition down and left by sw2\frac{sw}{2} and sh2\frac{sh}{2} pixels for the vertical rectangle window (V-Rwin).

    In implementation, feature maps undergo a down-left cyclical shift before self-attention, and a masking mechanism is applied during self-attention computation to prevent spurious attention between non-adjacent pixels wrap-around boundaries. Axial-shift facilitates two distinct interactions:

    1. Explicit Interaction: Directly merges adjacent windows of the same orientation (H-Rwin with H-Rwin, or V-Rwin with V-Rwin) across consecutive blocks.
    2. Implicit Interaction: Fuses features between orthogonal windows (H-Rwin with V-Rwin) via parallel multi-head execution and linear channel projection.
  5. Knowl 5 — Locality Complementary Module (LCM)

    model/method

    The Locality Complementary Module (LCM) incorporates convolutional inductive bias (translation invariance and local 2D structural awareness) directly into the self-attention mechanism by applying a depth-wise convolution in parallel with multi-head self-attention on the value projection.

    Let X∈RH×W×CX \in \mathbb{R}^{H \times W \times C} be the input feature map and V∈RH×W×CV \in \mathbb{R}^{H \times W \times C} be the unpartitioned value feature projected linearly from XX. The LCM applies a 3×33 \times 3 depth-wise convolution Conv(⋅)\text{Conv}(\cdot) to VV in parallel with the multi-head rectangle-window self-attention features Y1,…,YM∈RH×W×DY^1, \dots, Y^M \in \mathbb{R}^{H \times W \times D} (where D=C/MD = C/M and MM is the number of heads): Rwin-SA(X)=(Concat(Y1,Y2,…,YM)+Conv(V))WP\text{Rwin-SA}(X) = \left(\text{Concat}(Y^1, Y^2, \dots, Y^M) + \text{Conv}(V)\right) W_P where WP∈RC×CW_P \in \mathbb{R}^{C \times C} is the output projection matrix.

    Operating directly on VV rather than XX places the learnable static weights of the convolution in the same feature domain as the content-dependent dynamic attention weights, allowing the block to adaptively couple local structural features (edges, corners) with long-range contextual dependencies.

  6. Knowl 6 — Cross Aggregation Transformer Block (CATB)

    model/method

    A Cross Aggregation Transformer Block (CATB) consists of a Rectangle-Window Self-Attention (Rwin-SA) module equipped with the Locality Complementary Module (LCM), followed by a Multi-Layer Perceptron (MLP), with Layer Normalization (LN) and residual connections applied at each stage.

    For an input feature Xin∈RH×W×CX_{in} \in \mathbb{R}^{H \times W \times C}, the forward pass through the CATB is defined by: X′=Rwin-SA(LN(Xin))+XinX' = \text{Rwin-SA}(\text{LN}(X_{in})) + X_{in} Xout=MLP(LN(X′))+X′X_{out} = \text{MLP}(\text{LN}(X')) + X' where LN(⋅)\text{LN}(\cdot) denotes Layer Normalization, and MLP(⋅)\text{MLP}(\cdot) contains two linear projection layers with an intermediate GELU non-linearity and a channel expansion ratio of 4. Consecutive CATBs alternate between standard window partitioning and axial-shifted window partitioning.

  7. Knowl 7 — Ablation Analysis of Window Geometries, Axial Shifts, and Locality Coupling

    empirical result

    Ablation experiments conducted on single image super-resolution (imes2 imes 2) trained on DIV2K and Flickr2K (150K iterations, 64×6464 \times 64 patch size) and evaluated on Urban100 with FLOPs measured at input size 3×128×1283 \times 128 \times 128 demonstrate the individual contributions of Rwin-SA, axial-shift, and LCM:

    1. Window Shape and Shift Mechanism:
    • Square window (8×88 \times 8) without shift achieves 32.50 dB PSNR / 0.9325 SSIM at 281.8G FLOPs.
    • Square window (8×88 \times 8) with Swin shift achieves 32.75 dB PSNR / 0.9347 SSIM at 281.8G FLOPs.
    • Regular rectangle window (4×164 \times 16) without axial shift achieves 32.66 dB PSNR / 0.9334 SSIM at 281.8G FLOPs (+0.16 dB over square without shift at identical FLOPs).
    • Regular rectangle window (4×164 \times 16) with axial-shift achieves 32.91 dB PSNR / 0.9360 SSIM at 281.8G FLOPs (+0.16 dB over Swin square shift).
    1. Locality Complementary Module (LCM):
    • Adding LCM to CAT-R (regular-Rwin) increases PSNR from 32.91 dB to 32.98 dB (+0.07 dB) while increasing FLOPs from 281.8G to 282.7G (+0.32%).
    • Adding LCM to CAT-A (axial-Rwin) increases PSNR from 33.01 dB to 33.11 dB (+0.10 dB) while increasing FLOPs from 349.7G to 350.7G (+0.26%).
    1. Axial-Rwin Stripe Width (slsl):
    • sl=[2,2,2,2,2,2]sl = [2, 2, 2, 2, 2, 2] yields 32.97 dB PSNR (323.5G FLOPs).
    • sl=[2,2,2,4,4,4]sl = [2, 2, 2, 4, 4, 4] yields 33.11 dB PSNR (350.7G FLOPs).
    • sl=[4,4,4,4,4,4]sl = [4, 4, 4, 4, 4, 4] yields 33.20 dB PSNR (377.9G FLOPs).
  8. Knowl 8 — Quantitative Super-Resolution Benchmark Comparisons

    data/table

    Cross Aggregation Transformer variants CAT-R (regular rectangle window, [sh,sw]=[4,16][sh, sw] = [4, 16]) and CAT-A (axial rectangle window, sl=[2,2,2,4,4,4]sl = [2, 2, 2, 4, 4, 4]) outperform prior CNN and Transformer models across standard SR benchmarks (Set5, Set14, B100, Urban100, Manga109). Models marked with '+' utilize self-ensemble testing.

    Method Scale Set5 PSNR/SSIM Set14 PSNR/SSIM B100 PSNR/SSIM Urban100 PSNR/SSIM Manga109 PSNR/SSIM
    EDSR ×2\times 2 38.11 / 0.9602 33.92 / 0.9195 32.32 / 0.9013 32.93 / 0.9351 39.10 / 0.9773
    RCAN ×2\times 2 38.27 / 0.9614 34.12 / 0.9216 32.41 / 0.9027 33.34 / 0.9384 39.44 / 0.9786
    SAN ×2\times 2 38.31 / 0.9620 34.07 / 0.9213 32.42 / 0.9028 33.10 / 0.9370 39.32 / 0.9792
    IGNN ×2\times 2 38.24 / 0.9613 34.07 / 0.9217 32.41 / 0.9025 33.23 / 0.9383 39.35 / 0.9786
    HAN ×2\times 2 38.27 / 0.9614 34.16 / 0.9217 32.41 / 0.9027 33.35 / 0.9385 39.46 / 0.9785
    CSNLN ×2\times 2 38.28 / 0.9616 34.12 / 0.9223 32.40 / 0.9024 33.25 / 0.9386 39.37 / 0.9785
    NLSA ×2\times 2 38.34 / 0.9618 34.08 / 0.9231 32.43 / 0.9027 33.42 / 0.9394 39.59 / 0.9789
    IPT ×2\times 2 38.37 / - 34.43 / - 32.48 / - 33.76 / - - / -
    SwinIR ×2\times 2 38.42 / 0.9623 34.46 / 0.9250 32.53 / 0.9041 33.81 / 0.9427 39.92 / 0.9797
    CAT-R ×2\times 2 38.48 / 0.9625 34.53 / 0.9251 32.56 / 0.9045 34.08 / 0.9443 40.09 / 0.9804
    CAT-A ×2\times 2 38.51 / 0.9626 34.78 / 0.9265 32.59 / 0.9047 34.26 / 0.9440 40.10 / 0.9805
    CAT-R+ ×2\times 2 38.52 / 0.9627 34.59 / 0.9257 32.58 / 0.9047 34.19 / 0.9450 40.18 / 0.9805
    CAT-A+ ×2\times 2 38.55 / 0.9628 34.81 / 0.9267 32.60 / 0.9048 34.34 / 0.9445 40.18 / 0.9806
    SwinIR ×3\times 3 34.97 / 0.9318 30.93 / 0.8534 29.46 / 0.8145 29.75 / 0.8826 35.12 / 0.9537
    CAT-R ×3\times 3 34.99 / 0.9320 31.00 / 0.8539 29.49 / 0.8154 29.91 / 0.8848 35.29 / 0.9542
    CAT-A ×3\times 3 35.06 / 0.9326 31.04 / 0.8538 29.52 / 0.8160 30.12 / 0.8862 35.38 / 0.9546
    CAT-R+ ×3\times 3 35.07 / 0.9324 31.06 / 0.8544 29.52 / 0.8159 30.05 / 0.8864 35.44 / 0.9548
    CAT-A+ ×3\times 3 35.10 / 0.9327 31.09 / 0.8545 29.55 / 0.8164 30.21 / 0.8872 35.48 / 0.9550
    SwinIR ×4\times 4 32.92 / 0.9044 29.09 / 0.7950 27.92 / 0.7489 27.45 / 0.8254 32.03 / 0.9260
    CAT-R ×4\times 4 32.89 / 0.9044 29.13 / 0.7955 27.95 / 0.7500 27.62 / 0.8292 32.16 / 0.9269
    CAT-A ×4\times 4 33.08 / 0.9052 29.18 / 0.7960 27.99 / 0.7510 27.89 / 0.8339 32.39 / 0.9285
    CAT-R+ ×4\times 4 32.98 / 0.9049 29.18 / 0.7963 27.98 / 0.7506 27.73 / 0.8310 32.35 / 0.9280
    CAT-A+ ×4\times 4 33.14 / 0.9059 29.23 / 0.7968 28.01 / 0.7516 27.99 / 0.8356 32.52 / 0.9293

    The largest gains occur on Urban100, where CAT-A outperforms SwinIR by 0.45 dB (imes2 imes 2), 0.37 dB (imes3 imes 3), and 0.44 dB (imes4 imes 4), reflecting the advantage of rectangular and axial window attention on structured, directional textures.

  9. Knowl 9 — Performance on JPEG Artifact Reduction and Real Image Denoising

    data/table

    CAT achieves state-of-the-art results on JPEG compression artifact reduction and competitive results on real-world raw image denoising.

    1. JPEG Compression Artifact Reduction (Y-channel PSNR/SSIM):
    Dataset qq RNAN RDN DRUNet SwinIR CAT CAT+
    LIVE1 10 29.63 / 0.8239 29.67 / 0.8247 29.79 / 0.8278 29.86 / 0.8287 29.89 / 0.8295 29.92 / 0.8299
    LIVE1 20 32.03 / 0.8877 32.07 / 0.8882 32.17 / 0.8899 32.25 / 0.8909 32.30 / 0.8913 32.32 / 0.8915
    LIVE1 30 33.45 / 0.9149 33.51 / 0.9153 33.59 / 0.9166 33.69 / 0.9174 33.73 / 0.9177 33.75 / 0.9179
    LIVE1 40 34.47 / 0.9299 34.51 / 0.9302 34.58 / 0.9312 34.67 / 0.9317 34.72 / 0.9320 34.74 / 0.9322
    Classic5 10 29.96 / 0.8178 30.00 / 0.8188 30.16 / 0.8234 30.27 / 0.8249 30.26 / 0.8250 30.30 / 0.8257
    Classic5 20 32.11 / 0.8693 32.15 / 0.8699 32.39 / 0.8734 32.52 / 0.8748 32.57 / 0.8754 32.60 / 0.8756
    Classic5 30 33.38 / 0.8924 33.43 / 0.8930 33.59 / 0.8949 33.73 / 0.8961 33.77 / 0.8964 33.80 / 0.8966
    Classic5 40 34.27 / 0.9061 34.27 / 0.9061 34.41 / 0.9075 34.52 / 0.9082 34.58 / 0.9087 34.60 / 0.9088
    1. Real Image Denoising (SIDD and DND benchmarks):
    Dataset DANet+ CycleISP MIRNet MPRNet Uformer Restormer CAT CAT+
    Params (M) 9.15 2.83 31.79 15.74 50.88 26.11 25.77 25.77
    SIDD PSNR 39.47 39.52 39.72 39.71 39.89 40.02 40.01 40.05
    SIDD SSIM 0.9570 0.9571 0.9586 0.9586 0.9594 0.9603 0.9600 0.9602
    DND PSNR 39.58 39.56 39.88 39.82 39.98 40.03 40.05 40.08
    DND SSIM 0.9545 0.9564 0.9563 0.9540 0.9554 0.9564 0.9561 0.9563

    On DND, CAT attains 40.05 dB (40.08 dB with ensemble), surpassing Restormer (40.03 dB) while requiring fewer parameters (25.77M vs. 26.11M).

  10. Knowl 10 — Model Complexity and Efficiency Comparison on Image Super-Resolution

    data/table

    Computational complexity (FLOPs measured on 3×512×5123 \times 512 \times 512 output size), parameter counts, and reconstruction performance (PSNR on Urban100 ×4\times 4) show that CAT models achieve favorable trade-offs between computational budget and restoration accuracy.

    Method Urban100 PSNR (dB) FLOPs (G) Parameters (M)
    EDSR 26.64 823.3 43.09
    RCAN 26.82 261.0 15.59
    HAN 26.85 269.1 16.07
    CSNLN 27.22 84,155.2 6.57
    SwinIR 27.45 215.3 11.90
    CAT-R (ours) 27.62 292.7 16.60
    CAT-A (ours) 27.89 360.7 16.60
    CAT-R-2 (ours) 27.59 216.3 11.93

    When constrained to the parameter and FLOP profile of SwinIR (11.93M vs 11.90M params, 216.3G vs 215.3G FLOPs), the lightweight variant CAT-R-2 achieves 27.59 dB, outperforming SwinIR (27.45 dB) by 0.14 dB.

Coverage note — None was omitted; all main architectural components, mathematical formulations, theoretical complexity bounds, ablation experiments, and empirical benchmark comparisons across Super-Resolution, JPEG artifact reduction, and real image denoising were captured.

References

  1. 1.Abdelrahman Abdelhamed, Stephen Lin, and Michael S Brown. A high-quality denoising dataset for smartphone cameras. In CVPR, 2018. 6
  2. 2.Pablo Arbelaez, Michael Maire, Charless Fowlkes, and Jitendra Malik. Contour detection and hierarchical image segmentation. TPAMI, 2010. 6
  3. 3.Marco Bevilacqua, Aline Roumy, Christine Guillemot, and Marie Line Alberi-Morel. Low-complexity single-image super-resolution based on nonnegative neighbor embedding. In BMVC, 2012. 6
  4. 4.Pierre Charbonnier, Laure Blanc-Feraud, Gilles Aubert, and Michel Barlaud. Two deterministic half-quadratic regularization algorithms for computed imaging. In ICIP, 1994. 6
  5. 5.Hanting Chen, Yunhe Wang, Tianyu Guo, Chang Xu, Yiping Deng, Zhenhua Liu, Siwei Ma, Chunjing Xu, Chao Xu, and Wen Gao. Pre-trained image processing transformer. In CVPR, 2021. 1, 3, 4, 7, 8
  6. 6.Haoyu Chen, Jinjin Gu, and Zhi Zhang. Attention in attention network for image super-resolution. arXiv preprint arXiv:2104.09497, 2021. 2
  7. 7.Xiangxiang Chu, Zhi Tian, Yuqing Wang, Bo Zhang, Haibing Ren, Xiaolin Wei, Huaxia Xia, and Chunhua Shen. Twins: Revisiting the design of spatial attention in vision transformers. In NeurIPS, 2021. 3
  8. 8.Xiangxiang Chu, Zhi Tian, Bo Zhang, Xinlong Wang, Xiaolin Wei, Huaxia Xia, and Chunhua Shen. Conditional positional encodings for vision transformers. arXiv preprint arXiv:2102.10882, 2021. 5
  9. 9.Tao Dai, Jianrui Cai, Yongbing Zhang, Shu-Tao Xia, and Lei Zhang. Second-order attention network for single image super-resolution. In CVPR, 2019. 7, 8
  10. 10.Chao Dong, Yubin Deng, Chen Change Loy, and Xiaoou Tang. Compression artifacts reduction by a deep convolutional network. In ICCV, 2015. 1, 2
  11. 11.Chao Dong, Chen Change Loy, Kaiming He, and Xiaoou Tang. Learning a deep convolutional network for image super-resolution. In ECCV, 2014. 1, 2
  12. 12.Xiaoyi Dong, Jianmin Bao, Dongdong Chen, Weiming Zhang, Nenghai Yu, Lu Yuan, Dong Chen, and Baining Guo. Cswin transformer: A general vision transformer backbone with cross-shaped windows. In CVPR, 2022. 1, 2, 5
  13. 13.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. In ICLR, 2021. 1, 2, 3
  14. 14.Alessandro Foi, Vladimir Katkovnik, and Karen Egiazarian. Pointwise shape-adaptive dct for high-quality denoising and deblocking of grayscale and color images. TIP, 2007. 6
  15. 15.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In CVPR, 2016. 2
  16. 16.Gao Huang, Zhuang Liu, Kilian Q Weinberger, and Laurens van der Maaten. Densely connected convolutional networks. In CVPR, 2017. 2
  17. 17.Jia-Bin Huang, Abhishek Singh, and Narendra Ahuja. Single image super-resolution from transformed self-exemplars. In CVPR, 2015. 4, 6, 7
  18. 18.Diederik Kingma and Jimmy Ba. Adam: A method for stochastic optimization. In ICLR, 2015. 6
  19. 19.Xiangtao Kong, Xina Liu, Jinjin Gu, Yu Qiao, and Chao Dong. Reflash dropout in image super-resolution. In CVPR, 2022. 2
  20. 20.Zheyuan Li, Yingqi Liu, Xiangyu Chen, Haoming Cai, Jinjin Gu, Yu Qiao, and Chao Dong. Blueprint separable residual network for efficient image super-resolution. In CVPR, 2022. 2
  21. 21.Jingyun Liang, Jiezhang Cao, Guolei Sun, Kai Zhang, Luc Van Gool, and Radu Timofte. Swinir: Image restoration using swin transformer. In ICCVW, 2021. 1, 3, 4, 6, 7, 8, 9, 10
  22. 22.Bee Lim, Sanghyun Son, Heewon Kim, Seungjun Nah, and Kyoung Mu Lee. Enhanced deep residual networks for single image super-resolution. In CVPRW, 2017. 6, 7, 8, 9, 10
  23. 23.Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. In ICCV, 2021. 2, 5, 7
  24. 24.Ilya Loshchilov and Frank Hutter. Sgdr: Stochastic gradient descent with warm restarts. In ICLR, 2017. 6
  25. 25.Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. In ICLR, 2019. 6
  26. 26.Kede Ma, Zhengfang Duanmu, Qingbo Wu, Zhou Wang, Hongwei Yong, Hongliang Li, and Lei Zhang. Waterloo exploration database: New challenges for image quality assessment models. TIP, 2016. 6
  27. 27.David Martin, Charless Fowlkes, Doron Tal, and Jitendra Malik. A database of human segmented natural images and its application to evaluating segmentation algorithms and measuring ecological statistics. In ICCV, 2001. 6
  28. 28.Yusuke Matsui, Kota Ito, Yuji Aramaki, Azuma Fujimoto, Toru Ogawa, Toshihiko Yamasaki, and Kiyoharu Aizawa. Sketch-based manga retrieval using manga109 dataset. Multimedia Tools and Applications, 2017. 6
  29. 29.Yiqun Mei, Yuchen Fan, and Yuqian Zhou. Image super-resolution with non-local sparse attention. In CVPR, 2021. 7, 8
  30. 30.Yiqun Mei, Yuchen Fan, Yuqian Zhou, Lichao Huang, Thomas S Huang, and Humphrey Shi. Image super-resolution with cross-scale non-local attention and exhaustive self-exemplars mining. In CVPR, 2020. 7, 8, 10
  31. 31.Ben Niu, Weilei Wen, Wenqi Ren, Xiangde Zhang, Lianping Yang, Shuzhen Wang, Kaihao Zhang, Xiaochun Cao, and Haifeng Shen. Single image super-resolution via a holistic attention network. In ECCV, 2020. 2, 7, 8, 10
  32. 32.Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer. Automatic differentiation in pytorch. 2017. 6
  33. 33.Tobias Plotz and Stefan Roth. Benchmarking denoising algorithms with real photographs. In CVPR, 2017. 6
  34. 34.Olaf Ronneberger, Philipp Fischer, and Thomas Brox. U-net: Convolutional networks for biomedical image segmentation. In MICCAI, 2015. 2, 3
  35. 35.Hamid R Sheikh, Muhammad F Sabir, and Alan C Bovik. A statistical evaluation of recent full reference image quality assessment algorithms. TIP, 2006. 6
  36. 36.Wenzhe Shi, Jose Caballero, Ferenc Huszár, Johannes Totz, Andrew P Aitken, Rob Bishop, Daniel Rueckert, and Zehan Wang. Real-time single image and video super-resolution using an efficient sub-pixel convolutional neural network. In CVPR, 2016. 3
  37. 37.Radu Timofte, Eirikur Agustsson, Luc Van Gool, Ming-Hsuan Yang, Lei Zhang, Bee Lim, Sanghyun Son, Heewon Kim, Seungjun Nah, Kyoung Mu Lee, et al. Ntire 2017 challenge on single image super-resolution: Methods and results. In CVPRW, 2017. 6, 7
  38. 38.Wenxiao Wang, Lu Yao, Long Chen, Binbin Lin, Deng Cai, Xiaofei He, and Wei Liu. Crossformer: A versatile vision transformer hinging on cross-scale attention. In ICLR, 2022. 4
  39. 39.Zhendong Wang, Xiaodong Cun, Jianmin Bao, Wengang Zhou, Jianzhuang Liu, and Houqiang Li. Uformer: A general u-shaped transformer for image restoration. In CVPR, 2022. 3, 9
  40. 40.Zhou Wang, Alan C Bovik, Hamid R Sheikh, and Eero P Simoncelli. Image quality assessment: from error visibility to structural similarity. TIP, 2004. 6
  41. 41.Jianwei Yang, Chunyuan Li, Pengchuan Zhang, Xiyang Dai, Bin Xiao, Lu Yuan, and Jianfeng Gao. Focal self-attention for local-global interactions in vision transformers. In NeurIPS, 2021. 1, 2
  42. 42.Kun Yuan, Shaopeng Guo, Ziwei Liu, Aojun Zhou, Fengwei Yu, and Wei Wu. Incorporating convolution designs into visual transformers. In CVPR, 2021. 5
  43. 43.Zongsheng Yue, Qian Zhao, Lei Zhang, and Deyu Meng. Dual adversarial network: Toward real-world noise removal and noise generation. In ECCV, 2020. 9
  44. 44.Syed Waqas Zamir, Aditya Arora, Salman Khan, Munawar Hayat, Fahad Shahbaz Khan, and Ming-Hsuan Yang. Restormer: Efficient transformer for high-resolution image restoration. In CVPR, 2022. 1, 3, 4, 6, 9
  45. 45.Syed Waqas Zamir, Aditya Arora, Salman Khan, Munawar Hayat, Fahad Shahbaz Khan, Ming-Hsuan Yang, and Ling Shao. Cycleisp: Real image restoration via improved data synthesis. In CVPR, 2020. 9
  46. 46.Syed Waqas Zamir, Aditya Arora, Salman Khan, Munawar Hayat, Fahad Shahbaz Khan, Ming-Hsuan Yang, and Ling Shao. Learning enriched features for real image restoration and enhancement. In ECCV, 2020. 9
  47. 47.Syed Waqas Zamir, Aditya Arora, Salman Khan, Munawar Hayat, Fahad Shahbaz Khan, Ming-Hsuan Yang, and Ling Shao. Multi-stage progressive image restoration. In CVPR, 2021. 9
  48. 48.Roman Zeyde, Michael Elad, and Matan Protter. On single image scale-up using sparse-representations. In Proc. 7th Int. Conf. Curves Surf., 2010. 6
  49. 49.Kai Zhang, Yawei Li, Wangmeng Zuo, Lei Zhang, Luc Van Gool, and Radu Timofte. Plug-and-play image restoration with deep denoiser prior. TPAMI, 2021. 8, 9
  50. 50.Kai Zhang, Wangmeng Zuo, Yunjin Chen, Deyu Meng, and Lei Zhang. Beyond a gaussian denoiser: Residual learning of deep cnn for image denoising. TIP, 2017. 1
  51. 51.Yulun Zhang, Kunpeng Li, Kai Li, Lichen Wang, Bineng Zhong, and Yun Fu. Image super-resolution using very deep residual channel attention networks. In ECCV, 2018. 2, 3, 6, 7, 8, 10
  52. 52.Yulun Zhang, Kunpeng Li, Kai Li, Bineng Zhong, and Yun Fu. Residual non-local attention networks for image restoration. In ICLR, 2019. 2, 6, 8, 9
  53. 53.Yulun Zhang, Yapeng Tian, Yu Kong, Bineng Zhong, and Yun Fu. Residual dense network for image restoration. TPAMI, 2020. 8, 9
  54. 54.Shangchen Zhou, Jiawei Zhang, Wangmeng Zuo, and Chen Change Loy. Cross-scale internal graph neural network for image super-resolution. In NeurIPS, 2020. 7, 8

Citation

MLA
Chen, Z., et al. “Cross Aggregation Transformer for Image Restoration”. Advances in Neural Information Processing Systems, vol. 35, 2022, pp. 25478–90, https://proceedings.neurips.cc/paper_files/paper/2022/file/a37fea8e67f907311826bc1ba2654d97-Paper-Conference.pdf.
APA
Chen, Z., Zhang, Y., Gu, J., zhang, . yongbing ., Kong, L., & Yuan, X. (2022). Cross Aggregation Transformer for Image Restoration. Advances in Neural Information Processing Systems, 35, 25478–25490. https://proceedings.neurips.cc/paper_files/paper/2022/file/a37fea8e67f907311826bc1ba2654d97-Paper-Conference.pdf
Chicago
Chen, Z., Y. Zhang, J. Gu, . yongbing . zhang, L. Kong, and X. Yuan. 2022. “Cross Aggregation Transformer for Image Restoration”. Advances in Neural Information Processing Systems 35: 25478–90. https://proceedings.neurips.cc/paper_files/paper/2022/file/a37fea8e67f907311826bc1ba2654d97-Paper-Conference.pdf.
Harvard
Chen, Z. et al. (2022) “Cross Aggregation Transformer for Image Restoration”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 25478–25490. Available at: https://proceedings.neurips.cc/paper_files/paper/2022/file/a37fea8e67f907311826bc1ba2654d97-Paper-Conference.pdf.
Vancouver
1. Chen Z, Zhang Y, Gu J, zhang yongbing, Kong L, Yuan X (2022) Cross Aggregation Transformer for Image Restoration. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 25478–25490

BibTeX

@inproceedings{chen2022cross,
  title = {Cross Aggregation Transformer for Image Restoration},
  author = {Chen, Zheng and Zhang, Yulun and Gu, Jinjin and zhang, yongbing and Kong, Linghe and Yuan, Xin},
  year = {2022},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {35},
  pages = {25478-25490},
  url = {https://proceedings.neurips.cc/paper_files/paper/2022/file/a37fea8e67f907311826bc1ba2654d97-Paper-Conference.pdf}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors