Pan-Sharpening with Customized Transformer and Invertible Neural Network

Man ZhouJie HuangYanchi FangXueyang FuAiping Liu

article2022AAAI144 citations

Proposes a pan-sharpening framework that combines a customized cross-modal transformer for long-range dependency modeling with an invertible neural network to achieve information-lossless fusion of panchromatic and multispectral satellite images.

Listen

Satellite remote sensing systems rely on pan-sharpening to merge high-resolution panchromatic images with low-resolution multispectral data, generating detailed, high-resolution color imagery critical for mapping, environmental monitoring, and security applications. Conventional methods and standard convolutional neural networks often struggle to capture long-range spatial relationships or preserve information during feature fusion. The article aims to evaluate a novel fusion architecture that integrates a customized attention-based transformer to model long-range dependencies with a densely connected, information-lossless invertible neural network to enhance feature fusion while reducing computational complexity.

The evaluated approach processes satellite imagery through a dual-branch extraction framework that separates local convolutional processing from long-range cross-modality attention, followed by an invertible network module with affine coupling layers to combine the features. The original experiments evaluated this architecture across multiple satellite datasets, including WorldView-2, WorldView-3, and GaoFen-2, assessing performance against ten existing baseline methods using standard image quality and reconstruction metrics.

Critically, the article is accompanied by a formal retraction issued by the authors and publisher in July 2025. Verification tests conducted in June 2025 uncovered significant experimental flaws: a subset of the test samples had been inadvertently mixed into the training dataset. This data leakage caused artificially inflated performance metrics that cannot be reproduced under correct experimental conditions. While the original results claimed the model outperformed state-of-the-art methods with fewer computational operations and lower parameter counts, these conclusions are invalid due to the corrupted evaluation setup.

From a risk and decision-making standpoint, this work cannot be relied upon for practical deployment, benchmarking, or architectural baseline selection. Attempting to implement or deploy this method based on the reported metrics creates substantial operational and technical risks, as real-world performance will fail to match the published benchmarks. Organizations and researchers should not adopt this proposed architecture until the authors or independent teams re-evaluate the model on properly partitioned datasets without data overlap. Stakeholders must treat the reported results as unverified and await reproducible validation under strict data-splitting protocols.

Cover for Pan-Sharpening with Customized Transformer and Invertible Neural Network

Abstract

In remote sensing imaging systems, pan-sharpening is an important technique to obtain high-resolution multispectral images from a high-resolution panchromatic image and its corresponding low-resolution multispectral image. Owing to the powerful learning capability of convolution neural network (CNN), CNN-based methods have dominated this field. However, due to the limitation of the convolution operator, long-range spatial features are often not accurately obtained, thus limiting the overall performance. To this end, we propose a novel and effective method by exploiting a customized transformer architecture and information-lossless invertible neural module for long-range dependencies modeling and effective feature fusion in this paper. Specifically, the customized transformer formulates the PAN and MS features as queries and keys to encourage joint feature learning across two modalities while the designed invertible neural module enables effective feature fusion to generate the expected pan-sharpened results. To the best of our knowledge, this is the first attempt to introduce transformer and invertible neural network into pan-sharpening field. Extensive experiments over different kinds of satellite datasets demonstrate that our method outperforms state-of-the-art algorithms both visually and quantitatively with fewer parameters and flops. Further, the ablation experiments also prove the effectiveness of the proposed customized long-range transformer and effective invertible neural feature fusion module for pan-sharpening.

Table of Contents

  • Invertible Neural Module for Feature Fusion
  • References

Knowls

  1. Knowl 1 — Dual-Branch Pan-Sharpening Network with Transformer and Invertible Neural Network

    model/method

    The pan-sharpening network fuses a single-channel high-resolution panchromatic (PAN) image P∈R1×H×WP \in \mathbb{R}^{1 \times H \times W} and a low-resolution multispectral (MS) image M∈RC×H4×W4M \in \mathbb{R}^{C \times \frac{H}{4} \times \frac{W}{4}} to reconstruct a high-resolution multispectral (HR-MS) image H∈RC×H×WH \in \mathbb{R}^{C \times H \times W}.

    The architecture operates in three consecutive stages:

    1. Initial Feature Extraction: The low-resolution MS image is upsampled by a factor of 4 via bicubic interpolation to match the PAN spatial dimensions, denoted as M^=(M)↑4∈RC×H×W\hat{M} = (M)\uparrow_4 \in \mathbb{R}^{C \times H \times W}. Two separate 3×33 \times 3 convolutional layers map M^\hat{M} and PP into modality-specific shallow feature maps M0∈R8×H×WM_0 \in \mathbb{R}^{8 \times H \times W} and P0∈R8×H×WP_0 \in \mathbb{R}^{8 \times H \times W}, respectively.

    2. Two-Stream Local and Long-Range Feature Extraction: Shallow features are fed in parallel into a local feature branch consisting of 3×33 \times 3 convolutional layers to yield local feature map L0L_0, and a customized cross-modal transformer branch to yield long-range feature map G0G_0.

    3. Invertible Feature Fusion and Reconstruction: Local features L0L_0 and long-range features G0G_0 are fused through a stack of k=3k = 3 densely connected invertible affine coupling blocks. The outputs across all coupling blocks {L0,G0,…,Lk,Gk}\{L_0, G_0, \dots, L_k, G_k\} are concatenated and passed through a Residual Channel Attention Block (RCAB) to predict high-frequency residual detail, which is added to the upsampled MS image:

    H=(M)↑4+RCAB([L0,G0,…,Lk,Gk])H = (M)\uparrow_4 + \text{RCAB}([L_0, G_0, \dots, L_k, G_k])

    where [… ][\dots] denotes concatenation along the channel dimension.

  2. Knowl 2 — Customized Cross-Modal Transformer for Long-Range Dependency Modeling

    model/method

    To capture long-range complementary spatial-spectral interactions between panchromatic (PAN) and multispectral (MS) modalities, the shallow feature representations M0∈R8×H×WM_0 \in \mathbb{R}^{8 \times H \times W} and P0∈R8×H×WP_0 \in \mathbb{R}^{8 \times H \times W} are partitioned into 16×1616 \times 16 pixel patches {M1,…,Mn}\{M^1, \dots, M^n\} and {P1,…,Pn}\{P^1, \dots, P^n\}.

    Convolutional layers project these patches into Query (QQ), Key (KK), and two Value components (V1,V2V_1, V_2):

    Q=Conv([M1,…,Mn])Q = \text{Conv}([M^1, \dots, M^n])

    K=Conv([P1,…,Pn])K = \text{Conv}([P^1, \dots, P^n])

    V1=Conv([M1,…,Mn])V_1 = \text{Conv}([M^1, \dots, M^n])

    V2=Conv([P1,…,Pn])V_2 = \text{Conv}([P^1, \dots, P^n])

    where [… ][\dots] denotes channel concatenation. Query and key tensors are unfolded into patch vectors qiq_i and kjk_j (i,j∈[1,H×W]i, j \in [1, H \times W]). The cross-modal patch relevance ri,jr_{i,j} and relevance matrix RR are calculated via normalized inner products:

    ri,j=⟨qi∥qi∥2,kj∥kj∥2⟩,R=QTKr_{i,j} = \left\langle \frac{q_i}{\|q_i\|_2}, \frac{k_j}{\|k_j\|_2} \right\rangle, \quad R = Q^T K

    A hard-attention index map hi=arg⁡max⁡j(ri,j)h_i = \arg\max_j (r_{i,j}) selects the most relevant patches from V1V_1 and V2V_2:

    ti1=vhi1,ti2=vhi2t^1_i = v^1_{h_i}, \quad t^2_i = v^2_{h_i}

    yielding position-indexed aligned features T1T_1 and T2T_2. A soft-attention map S=softmax(R)S = \text{softmax}(R) modulates the aligned features, and the enhanced long-range representation G0G_0 is constructed as:

    G01=P0+Conv([P0,T1])⊙SG^1_0 = P_0 + \text{Conv}([P_0, T_1]) \odot S

    G02=M0+Conv([M0,T2])⊙SG^2_0 = M_0 + \text{Conv}([M_0, T_2]) \odot S

    G0=Conv([G01,G02])G_0 = \text{Conv}([G^1_0, G^2_0])

    where ⊙\odot represents the Hadamard (element-wise) product.

  3. Knowl 3 — Densely-Connected Invertible Neural Network Feature Fusion Module

    model/method

    To preserve information during multi-modal feature fusion without loss, a densely connected invertible neural network (INN) composed of k=3k = 3 affine coupling layers is used. The input naturally splits into local features L0L_0 and long-range features G0G_0.

    For each coupling layer i∈{1,…,k}i \in \{1, \dots, k\}, the forward transformation is defined as:

    Li=Li−1+ϕ(Gi−1)L_i = L_{i-1} + \phi(G_{i-1})

    Gi=Gi−1⊙exp⁡(ρ(Li))+η(Li)G_i = G_{i-1} \odot \exp(\rho(L_i)) + \eta(L_i)

    where ⊙\odot denotes the Hadamard product, exp⁡(⋅)\exp(\cdot) is the element-wise exponential function, and ρ(⋅)\rho(\cdot), η(⋅)\eta(\cdot), and ϕ(⋅)\phi(\cdot) are non-linear neural network mapping functions. The transformation ϕ(⋅)\phi(\cdot) provides additive updates from long-range to local features, while ρ(⋅)\rho(\cdot) and η(⋅)\eta(\cdot) provide scale and translation updates to the long-range branch.

    Dense skip connections feed intermediate features from all units {L0,G0,L1,G1,…,Lk,Gk}\{L_0, G_0, L_1, G_1, \dots, L_k, G_k\} via channel concatenation into the final reconstruction layer.

  4. Knowl 4 — Half Instance Normalization Transformation Block

    model/method

    The transformation operations ρ(⋅)\rho(\cdot), η(⋅)\eta(\cdot), and ϕ(⋅)\phi(\cdot) in the invertible affine coupling layers are each parameterized using two cascaded Half Instance Normalization (HIN) blocks.

    Given an input feature map Fin∈RCin×H×WF_{in} \in \mathbb{R}^{C_{in} \times H \times W}:

    1. A 3×33 \times 3 convolution projects FinF_{in} into intermediate features Fmid∈R16×H×WF_{mid} \in \mathbb{R}^{16 \times H \times W}:

    Fmid=Conv3×3(Fin)F_{mid} = \text{Conv}_{3\times 3}(F_{in})

    1. FmidF_{mid} is split equally along channels into Fmid1,Fmid2∈R8×H×WF_{mid1}, F_{mid2} \in \mathbb{R}^{8 \times H \times W}:

    (Fmid1,Fmid2)=split(Fmid)(F_{mid1}, F_{mid2}) = \text{split}(F_{mid})

    1. Instance Normalization (IN) is applied to Fmid1F_{mid1} while leaving Fmid2F_{mid2} untouched to preserve contextual information:

    Fres=concat(IN(Fmid1),Fmid2)F_{res} = \text{concat}(\text{IN}(F_{mid1}), F_{mid2})

    1. FresF_{res} passes through a 3×33 \times 3 convolution and two LeakyReLU layers, and is added to a residual shortcut of FinF_{in} via 1×11 \times 1 convolution:

    Fout=Fres+FinF_{out} = F_{res} + F_{in}

  5. Knowl 5 — Mean Absolute Error Optimization Objective

    equation

    The pan-sharpening model parameters are optimized end-to-end using the mean absolute error (L1L_1 loss) between the predicted high-resolution multispectral image and the ground truth:

    L=∑i=1K∥Hi−Hgt,i∥1\mathcal{L} = \sum_{i=1}^K \|H_i - H_{gt,i}\|_1

    where KK denotes the total number of training image pairs, Hi∈RC×H×WH_i \in \mathbb{R}^{C \times H \times W} is the pan-sharpened output image for sample ii, Hgt,i∈RC×H×WH_{gt,i} \in \mathbb{R}^{C \times H \times W} is the corresponding ground-truth high-resolution multispectral image, and ∥⋅∥1\|\cdot\|_1 is the element-wise ℓ1\ell_1 norm.

  6. Knowl 6 — Pan-Sharpening Experimental Setup and Wald Protocol

    experimental setup

    Experiments are conducted on three satellite benchmarks: WorldView-II, WorldView-III, and GaoFen2.

    • Wald Protocol for Paired Data: Due to the absence of full-resolution ground truth, the Wald protocol downsamples original MS images H∈RM×N×CH \in \mathbb{R}^{M \times N \times C} and PAN images P∈RrM×rN×1P \in \mathbb{R}^{rM \times rN \times 1} by scale factor r=4r = 4. The downsampled versions L∈RM/r×N/r×CL \in \mathbb{R}^{M/r \times N/r \times C} and p∈RM×N×1p \in \mathbb{R}^{M \times N \times 1} serve as model inputs, with HH serving as ground truth.
    • Patch Size and Preprocessing: MS images are cropped into 32×3232 \times 32 patches and PAN images into 128×128128 \times 128 patches, with pixel intensities normalized to the range [0,1][0, 1].
    • Training Details: Implemented in PyTorch and trained on a single NVIDIA GeForce GTX 2080Ti GPU with Adam optimizer for 1000 epochs, batch size 4, and an initial learning rate of 8×10−48 \times 10^{-4} decayed by a factor of 0.5 at epoch 200.
    • Image Quality Assessment Metrics: Peak Signal-to-Noise Ratio (PSNR, dB, ↑\uparrow), Structural Similarity Index (SSIM, ↑\uparrow), Spectral Angle Mapper (SAM, ↓\downarrow), and Relative Dimensionless Global Error in Synthesis (ERGAS, ↓\downarrow).
  7. Knowl 7 — Pan-Sharpening Performance Across Satellite Datasets

    data/table

    The pan-sharpening method was evaluated against traditional approaches (SFIM, Brovey, GS, IHS, GFPCA) and deep learning models (PNN, PANNET, MSDCNN, SRPPNN, GPPNN) across WorldView-II, GaoFen2, and WorldView-III datasets under the Wald protocol.

    Method WorldView II GaoFen2 WorldView III
    PSNR↑\uparrow SSIM↑\uparrow SAM↓\downarrow ERGAS↓\downarrow PSNR↑\uparrow SSIM↑\uparrow SAM↓\downarrow EGAS↓\downarrow PSNR↑\uparrow SSIM↑\uparrow SAM↓\downarrow EGAS↓\downarrow
    SFIM 34.1297 0.8975 0.0439 2.3449 36.9060 0.8882 0.0318 1.7398 21.8212 0.5457 0.1208 8.9730
    Brovey 35.8646 0.9216 0.0403 1.8238 37.7974 0.9026 0.0218 1.3720 22.5060 0.5466 0.1159 8.2331
    GS 35.6376 0.9176 0.0423 1.8774 37.2260 0.9034 0.0309 1.6736 22.5608 0.5470 0.1217 8.2433
    IHS 35.2962 0.9027 0.0461 2.0278 38.1754 0.9100 0.0243 1.5336 22.5579 0.5354 0.1266 8.3616
    GFPCA 34.5581 0.9038 0.0488 2.1411 37.9443 0.9204 0.0314 1.5604 22.3344 0.4826 0.1294 8.3964
    PNN 40.7550 0.9624 0.0259 1.0646 43.1208 0.9704 0.0172 0.8528 29.9418 0.9121 0.0824 3.3206
    PANNET 40.8176 0.9626 0.0257 1.0557 43.0659 0.9685 0.0178 0.8577 29.6840 0.9072 0.0851 3.4263
    MSDCNN 41.3355 0.9664 0.0242 0.9940 45.6874 0.9827 0.0135 0.6389 30.3038 0.9184 0.0782 3.1884
    SRPPNN 41.4538 0.9679 0.0233 0.9899 47.1998 0.9877 0.0106 0.5586 30.4346 0.9202 0.0770 3.1553
    GPPNN 41.1622 0.9684 0.0244 1.0315 44.2145 0.9815 0.0137 0.7361 30.1785 0.9175 0.0776 3.2593
    Ours 41.6903 0.9704 0.0227 0.9514 47.3528 0.9893 0.0102 0.5479 30.5365 0.9225 0.0747 3.0997

    The proposed method achieved the best quantitative performance across all datasets, attaining the highest PSNR and SSIM, and lowest SAM and ERGAS.

  8. Knowl 8 — Model Complexity and Computational Efficiency Comparison

    data/table

    Model complexity was measured by parameter count (millions, M) and computational operations (FLOPs ×1010\times 10^{10}, G) using input tensors of size 1×4×32×321 \times 4 \times 32 \times 32 for the multispectral input and 1×1×128×1281 \times 1 \times 128 \times 128 for the panchromatic input:

    Metric PNN PANNET MSDCNN SRPPNN GPPNN Ours
    Params (M) 0.689 0.688 2.390 17.114 1.198 0.706
    FLOPs (G) 1.1289 1.1275 3.9158 21.1059 1.3967 1.3907

    The proposed model achieves high reconstruction quality with 0.706 M0.706\text{ M} parameters and 1.3907 G1.3907\text{ G} FLOPs, offering a competitive trade-off compared to high-complexity models like SRPPNN (17.114 M17.114\text{ M} parameters, 21.1059 G21.1059\text{ G} FLOPs).

  9. Knowl 9 — Ablation Study on Transformer and Invertible Neural Network Modules

    data/table

    To evaluate the individual contributions of the customized Transformer module and the densely connected Invertible Neural Network (INN) fusion module, ablation experiments were performed on three satellite datasets. Configuration (I) removed the Transformer branch while expanding feature channels in the local branch to match parameter count. Configuration (II) replaced the invertible fusion module with standard densely connected convolutional layers of matching parameter size.

    Configuration WorldView II GaoFen2 WorldView III
    PSNR↑\uparrow SSIM↑\uparrow SAM↓\downarrow ERGAS↓\downarrow PSNR↑\uparrow SSIM↑\uparrow SAM↓\downarrow EGAS↓\downarrow PSNR↑\uparrow SSIM↑\uparrow SAM↓\downarrow EGAS↓\downarrow
    (I) w/o Transformer 41.1932 0.9684 0.0238 1.0059 46.7804 0.9879 0.0110 0.5835 30.2943 0.9193 0.0784 3.1882
    (II) w/o Invertible Fusion 41.2232 0.9683 0.0238 1.0049 46.9368 0.9880 0.0106 0.5728 30.2588 0.9181 0.0785 3.2053
    Full Model (Ours) 41.6903 0.9704 0.0227 0.9514 47.3528 0.9893 0.0102 0.5479 30.5365 0.9225 0.0747 3.0997

    Removing the Transformer module caused PSNR to drop by 0.4971 dB0.4971\text{ dB} on WorldView-II, while removing the Invertible Fusion module caused a drop of 0.4671 dB0.4671\text{ dB}, confirming that both components contribute significantly to the overall fusion performance.

  10. Knowl 10 — Retraction Note on Training-Test Set Data Leakage

    limitation

    The paper was formally retracted by the authors and AAAI Press in July 2025. Verification tests conducted in June 2025 revealed that a subset of test samples had been inadvertently included in the training dataset during experimental preparation. This data leakage caused artificially inflated performance metrics across all reported benchmarks that could not be reproduced when using correctly partitioned, mutually disjoint datasets.

Coverage note — None was omitted; all contributed network architecture components, equations, experimental setups, quantitative benchmarks, computational complexity comparisons, ablation studies, and the formal retraction limitation note have been extracted as knowls.

References

  1. 1.Aiazzi, B.; Alparone, L.; Baronti, S.; Garzelli, A.; and Selva, M. 2003. An MTF-based spectral distortion minimizing model for pan-sharpening of very high resolution multispectral images of urban areas. In Joint Workshop on Remote Sensing and Data Fusion over Urban Areas.
  2. 2.Aiazzi, B. S., B.; and Selva, M. 2007. Improving Component Substitution Pansharpening Through Multivariate Regression of MS +Pan Data. IEEE Transactions on Geoscience and Remote Sensing, 45(10): 3230–3239.
  3. 3.Alparone, L.; Wald, L.; Chanussot, J.; Thomas, C.; Gamba, P.; and Bruce, L. M. 2007. Comparison of Pansharpening Algorithms: Outcome of the 2006 GRS-S Data Fusion Contest. IEEE Transactions on Geoscience and Remote Sensing, 45(10): 3012–3021.
  4. 4.Benzenati, T.; Kallel, A.; and Kessentini, Y. 2021. Two Stages Pan-Sharpening Details Injection Approach Based on Very Deep Residual Networks. IEEE Transactions on Geoscience and Remote Sensing, 59(6): 4984–4992.
  5. 5.Cai, J.; and Huang, B. 2021. Super-Resolution-Guided Progressive Pansharpening Based on a Deep Convolutional Neural Network. IEEE Transactions on Geoscience and Remote Sensing, 59(6): 5206–5220.
  6. 6.Carion, N.; Massa, F.; Synnaeve, G.; Usunier, N.; Kirillov, A.; and Zagoruyko, S. 2020. End-to-End Object Detection with Transformers. In European Conference on Computer Vision.
  7. 7.Chen, L.; Lu, X.; Zhang, J.; Chu, X.; and Chen, C. 2021. HINet: Half Instance Normalization Network for Image Restoration. arXiv:2105.06086.
  8. 8.Choi, Y. K., J.; and Kim, Y. 2011. A New Adaptive Component-Substitution-Based Satellite Image Fusion by Using Partial Replacement. IEEE Transactions on Geoscience and Remote Sensing, 49(1): p.295–309.
  9. 9.Dinh, L.; Krueger, D.; and Bengio, Y. 2015. NICE: Non-linear Independent Components Estimation. International Conference on Learning Representations.
  10. 10.Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; and Houlsby, N. 2020. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.
  11. 11.Garzelli, A.; Nencini, F.; and Capobianco, L. 2008. Optimal MMSE Pan Sharpening of Very High Resolution Multispectral Images. IEEE Transactions on Geoscience and Remote Sensing, 46(1): 228–236.
  12. 12.Ghahremani; Morteza; Ghassemian; and Hassan. 2016. A Compressed-Sensing-Based Pan-Sharpening Method for Spectral Distortion Reduction. IEEE Transactions on Geoscience and Remote Sensing.
  13. 13.Gillespie, A. R.; Kahle, A. B.; and Walker, R. E. 1987. Color enhancement of highly correlated images. II. Channel ratio and ”chromaticity” transformation techniques - ScienceDirect. Remote Sensing of Environment, 22(3): 343–365.
  14. 14.H. Wang, H. A. A. Y., Y. Zhu; and Chen, L.-C. 2021. Max-deeplab: End-to-end panoptic segmentation with mask transformers. In IEEE Conference on Computer Vision and Pattern Recognition.
  15. 15.Haydn, R.; Dalke, G. W.; Henkel, J.; and Bare, J. E. 1982. Application of the IHS color transform to the processing of multisensor data and image enhancement. National Academy of Sciences of the United States of America, 79(13): 571–577.
  16. 16.Hu, J.; Hu, P.; Kang, X.; Zhang, H.; and Fan, S. 2021. Pan-Sharpening via Multiscale Dynamic Convolutional Neural Network. IEEE Transactions on Geoscience and Remote Sensing, 59(3): 2231–2244.
  17. 17.J. R. H. Yuhas, A. F. G.; and Boardman, J. M. 1992. Discrimination among semi-arid landscape endmembers using the spectral angle mapper (SAM) algorithm. Proc. Summaries Annu. JPL Airborne Geosci. Workshop, 147–149.
  18. 18.Kang, L. S., X.; and Benediktsson, J. A. 2014. Pansharpening With Matting Model. IEEE Transactions on Geoscience and Remote Sensing, 52(8): 5088–5099.
  19. 19.Kaplan, N. H.; and Erer, I. 2012. Bilateral pyramid based pansharpening of multispectral satellite images. In Geoscience and Remote Sensing Symposium.
  20. 20.Laben, C.; and Brower, B. 2000. Process for Enhancing the Spatial Resolution of Multispectral Imagery Using Pan-Sharpening. US Patent 6011875A.
  21. 21.Laurent Dinh, J. S.-D.; and Bengio., S. 2017. Density estimation using real NVP. ICLR.
  22. 22.Li, Y.; Zhang, K.; Cao, J.; Timofte, R.; and Gool, L. V. 2021. LocalViT: Bringing Locality to Vision Transformers. CoRR, abs/2104.05707.
  23. 23.Liao, W.; Xin, H.; Coillie, F. V.; Thoonen, G.; and Philips, W. 2017. Two-stage fusion of thermal hyperspectral and visible RGB image by PCA and guided filter. In Workshop on Hyperspectral Image and Signal Processing: Evolution in Remote Sensing.
  24. 24.Liu., J. G. 2000. Smoothing filter-based intensity modulation: A spectral preserve image fusion technique for improving spatial details. International Journal of Remote Sensing, 21(18): 3461–3472.
  25. 25.Liu, Q.; Zhou, H.; Xu, Q.; Liu, X.; and Wang, Y. 2020. PS-GAN: A Generative Adversarial Network for Remote Sensing Image Pan-Sharpening. IEEE Transactions on Geoscience and Remote Sensing, 1–16.
  26. 26.Liu, Y.; Qin, Z.; Anwar, S.; Ji, P.; Kim, D.; Caldwell, S.; and Gedeon, T. 2021. Invertible Denoising Network: A Light Solution for Real Noise Removal. In IEEE Conference on Computer Vision and Pattern Recognition, 13365–13374.
  27. 27.Lu, S.-P.; Wang, R.; Zhong, T.; and Rosin, P. L. 2021. Large-Capacity Image Steganography Based on Invertible Neural Networks. In IEEE Conference on Computer Vision and Pattern Recognition, 10816–10825.
  28. 28.Masi, G.; Cozzolino, D.; Verdoliva, L.; and Scarpa, G. 2016. Pansharpening by convolutional neural networks. Remote Sensing, 8(7): 594.
  29. 29.Paschalidou, D.; Katharopoulos, A.; Geiger, A.; and Fidler, S. 2021. Neural Parts: Learning Expressive 3D Shape Abstractions With Invertible Neural Networks. In IEEE Conference on Computer Vision and Pattern Recognition, 3204–3215.
  30. 30.Peng, J.; Liu, L.; Wang, J.; Zhang, E.; Zhu, X.; Zhang, Y.; Feng, J.; and Jiao, L. 2021. PSMD-Net: A Novel Pan-Sharpening Method Based on a Multiscale Dense Network. IEEE Transactions on Geoscience and Remote Sensing, 59(6): 4957–4971.
  31. 31.Shah, V. P.; Younan, N. H.; and King, R. L. 2008. An Efficient Pan-Sharpening Method via a Combined Adaptive PCA Approach and Contourlets. IEEE Transactions on Geoscience and Remote Sensing, 46(5): 1323–1335.
  32. 32.Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, L.; and Polosukhin, I. 2017. Attention is All you Need. In Advances in Neural Information Processing Systems, volume 30.
  33. 33.Wang, D.; Bai, Y.; Wu, C.; Li, Y.; Shang, C.; and Shen, Q. 2021a. Convolutional LSTM-Based Hierarchical Feature Fusion for Multispectral Pan-Sharpening. IEEE Transactions on Geoscience and Remote Sensing, 1–16.
  34. 34.Wang, J.; Shao, Z.; Huang, X.; Lu, T.; and Zhang, R. 2021b. A Dual-Path Fusion Network for Pan-Sharpening. IEEE Transactions on Geoscience and Remote Sensing, 1–14.
  35. 35.Wang, Z.; Cun, X.; Bao, J.; and Liu, J. 2021c. Uformer: A General U-Shaped Transformer for Image Restoration. CoRR, abs/2106.03106.
  36. 36.Wang, Z.; Cun, X.; Bao, J.; and Liu, J. 2021d. Uformer: A General U-Shaped Transformer for Image Restoration.
  37. 37.Xing, Y.; Qian, Z.; and Chen, Q. 2021. Invertible Image Signal Processing. In IEEE Conference on Computer Vision and Pattern Recognition, 6287–6296.
  38. 38.Xu, H.; Ma, J.; Shao, Z.; Zhang, H.; Jiang, J.; and Guo, X. 2021a. SDPNet: A Deep Network for Pan-Sharpening With Enhanced Information Representation. IEEE Transactions on Geoscience and Remote Sensing, 59(5): 4120–4134.
  39. 39.Xu, S.; Zhang, J.; Zhao, Z.; Sun, K.; Liu, J.; and Zhang, C. 2021b. Deep Gradient Projection Networks for Pan-sharpening. In IEEE Conference on Computer Vision and Pattern Recognition, 1366–1375.
  40. 40.Yang, F.; Yang, H.; Fu, J.; Lu, H.; and Guo, B. 2020. Learning Texture Transformer Network for Image Super-Resolution. In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).
  41. 41.Yang, J.; Fu, X.; Hu, Y.; Huang, Y.; Ding, X.; and Paisley, J. 2017. PanNet: A deep network architecture for pan-sharpening. In IEEE International Conference on Computer Vision, 5449–5457.
  42. 42.Yokoya, N.; Member, S.; IEEE; Yairi, T.; and Iwasaki, A. 2012. Coupled Nonnegative Matrix Factorization Unmixing for Hyperspectral and Multispectral Data Fusion. IEEE Transactions on Geoscience and Remote Sensing, 50(2): 528–537.
  43. 43.Yuan, K.; Guo, S.; Liu, Z.; Zhou, A.; Yu, F.; and Wu, W. 2021. Incorporating Convolution Designs Into Visual Transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 579–588.
  44. 44.Yuan, Q.; Wei, Y.; Meng, X.; Shen, H.; and Zhang, L. 2018. A Multiscale and Multidepth Convolutional Neural Network for Remote Sensing Imagery Pan-Sharpening. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 11(3): 978–989.
  45. 45.Zhang, S.; Zhang, C.; Kang, N.; and Li, Z. 2021. iVPF: Numerical Invertible Volume Preserving Flow for Efficient Lossless Compression. In IEEE Conference on Computer Vision and Pattern Recognition, 620–629.
  46. 46.Zhang, Y.; Li, K.; Li, K.; Wang, L.; Zhong, B.; and Fu, Y. 2018. Image super-resolution using very deep residual channel attention networks. In European Conference on Computer Vision, 286–301.
  47. 47.Zhu, X.; Su, W.; Lu, L.; Li, B.; Wang, X.; and Dai, J. 2020. Deformable DETR: Deformable Transformers for End-to-End Object Detection. In International Conference on Learning Representations.

Citation

MLA
Zhou, M., et al. “RETRACTED: Pan-Sharpening with Customized Transformer and Invertible Neural Network”. Proceedings of the AAAI Conference on Artificial Intelligence, vol. 36, no. 3, 2022, pp. 3553–61, https://doi.org/10.1609/AAAI.V36I3.20267.
APA
Zhou, M., Huang, J., Fang, Y., Fu, X., & Liu, A. (2022). RETRACTED: Pan-Sharpening with Customized Transformer and Invertible Neural Network. Proceedings of the AAAI Conference on Artificial Intelligence, 36(3), 3553–3561. https://doi.org/10.1609/AAAI.V36I3.20267
Chicago
Zhou, M., J. Huang, Y. Fang, X. Fu, and A. Liu. 2022. “RETRACTED: Pan-Sharpening with Customized Transformer and Invertible Neural Network”. Proceedings of the AAAI Conference on Artificial Intelligence 36 (3): 3553–61. https://doi.org/10.1609/AAAI.V36I3.20267.
Harvard
Zhou, M. et al. (2022) “RETRACTED: Pan-Sharpening with Customized Transformer and Invertible Neural Network”, Proceedings of the AAAI Conference on Artificial Intelligence, 36(3), pp. 3553–3561. Available at: https://doi.org/10.1609/AAAI.V36I3.20267.
Vancouver
1. Zhou M, Huang J, Fang Y, Fu X, Liu A (2022) RETRACTED: Pan-Sharpening with Customized Transformer and Invertible Neural Network. Proceedings of the AAAI Conference on Artificial Intelligence 36:3553–3561

BibTeX

@article{Zhou_2022, title={RETRACTED: Pan-Sharpening with Customized Transformer and Invertible Neural Network}, volume={36}, ISSN={2159-5399}, url={http://dx.doi.org/10.1609/AAAI.V36I3.20267}, DOI={10.1609/aaai.v36i3.20267}, number={3}, journal={Proceedings of the AAAI Conference on Artificial Intelligence}, publisher={Association for the Advancement of Artificial Intelligence (AAAI)}, author={Zhou, Man and Huang, Jie and Fang, Yanchi and Fu, Xueyang and Liu, Aiping}, year={2022}, month=June, pages={3553–3561} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF