HyperTransformer: A Textural and Spectral Feature Fusion Transformer for Pansharpening

Wele Gedara Chaminda BandaraVishal M. Patel

article2022CVPR152 citations

Proposes a transformer-based pansharpening framework that uses multi-head soft-attention to formulate hyperspectral and panchromatic representations as queries and keys, effectively transferring high-resolution textures while minimizing spatial and spectral distortions across multiple scales.

Listen

Hyperspectral imaging captures detailed chemical and material information across numerous narrow spectral bands, which is essential for remote sensing tasks such as environmental monitoring, disaster response, and object detection. However, physical hardware constraints typically force satellites and aircraft to capture either high-spectral-detail images at low spatial resolution or sharp panchromatic images with no detailed spectral information. Image fusion techniques, known as pansharpening, combine these two image types to synthesize high-resolution hyperspectral data. Existing methods often rely on simple image concatenation or standard convolutional networks that fail to establish long-range dependencies, resulting in significant spatial blurring and spectral distortion.

The article demonstrates that using an attention-based transformer architecture, named HyperTransformer, enables precise transfer of high-resolution textural details from panchromatic images into low-resolution hyperspectral images while preserving spectral fidelity. The main objective was to formulate and evaluate a dedicated feature-fusion model that explicitly captures cross-feature relationships and multi-scale visual details during the sharpening process.

To achieve this, the authors designed a deep learning architecture consisting of separate feature extractors for panchromatic and hyperspectral inputs, a multi-head soft-attention mechanism, and a multi-scale fusion framework. The attention module utilizes hyperspectral features as queries and panchromatic features as keys and values to match and transfer relevant structural textures. The model was trained using a combination of standard reconstruction loss, synthesized color perceptual loss, and transfer perceptual loss. Evaluation was conducted using standard down-sampling protocols across three standard remote sensing datasets—Pavia Center, Botswana, and Chikusei—and benchmarked against seventeen classical and state-of-the-art deep learning methods using six standard quality metrics.

The findings show that HyperTransformer consistently and significantly outperforms all existing methods across every evaluated benchmark. On the Pavia Center dataset, the proposed method reduced spectral angle distortion by approximately 27% and root mean square error by approximately 33% compared to previous top-performing techniques, while improving overall signal-to-noise ratio by roughly 13%. On the Botswana and Chikusei datasets, spectral errors dropped by roughly 19% and 13%, and spatial errors decreased by about 12% and 14%, respectively. Ablation analyses confirmed that utilizing sixteen attention heads and injecting textural details across three spatial scales (one-time, two-time, and four-time resolutions) provided optimal feature alignment and performance.

These results demonstrate that attention-guided feature fusion resolves the fundamental trade-off between spatial sharpness and spectral accuracy that has limited earlier automated fusion pipelines. In operational remote sensing workflows, adopting this architecture can significantly improve downstream analytics, such as automated land cover classification and target detection, without requiring cost-prohibitive hardware upgrades on airborne or satellite platforms.

Organizations developing remote sensing and earth observation processing pipelines should adopt multi-scale attention mechanisms for data fusion. Decision-makers are encouraged to test pre-trained HyperTransformer weights on pilot operational data to benchmark computational throughput against quality gains. Future research should prioritize enhancing performance in the ultraviolet and infrared spectral extremes, where pansharpening error remains relatively higher due to the limited wavelength coverage of standard panchromatic sensors.

arXiv: 2203.02503wgcban/HyperTransformer
Cover for HyperTransformer: A Textural and Spectral Feature Fusion Transformer for Pansharpening

Abstract

Pansharpening aims to fuse a registered high-resolution panchromatic image (PAN) with a low-resolution hyperspectral image (LR-HSI) to generate an enhanced HSI with high spectral and spatial resolution. Existing pansharpening approaches neglect using an attention mechanism to transfer HR texture features from PAN to LR-HSI features, resulting in spatial and spectral distortions. In this paper, we present a novel attention mechanism for pansharpening called HyperTransformer, in which features of LR-HSI and PAN are formulated as queries and keys in a transformer, respectively. HyperTransformer consists of three main modules, namely two separate feature extractors for PAN and HSI, a multi-head feature soft-attention module, and a spatial-spectral feature fusion module. Such a network improves both spatial and spectral quality measures of the pansharpened HSI by learning cross-feature space dependencies and long-range details of PAN and LR-HSI. Furthermore, HyperTransformer can be utilized across multiple spatial scales at the backbone for obtaining improved performance. Extensive experiments conducted on three widely used datasets demonstrate that HyperTransformer achieves significant improvement over the state-of-the-art methods on both spatial and spectral quality measures. Implementation code and pre-trained weights can be accessed at https://github.com/wgcban/HyperTransformer.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Methodology
  • 3.1. Feature Extractors for PAN and LR-HSI
  • 3.2. HR Texture Transfer through Multi-Head Feature Soft-Attention (MHFSA)
  • 3.3. Textural-Spectral Feature Fusion (TSFF)
  • 3.4. Multi-Scale Feature Fusion (MSFF)
  • 3.5. Loss Functions
  • 4. Experiments
  • 4.1. Datasets and Performance Metrics
  • 4.2. Results and Discussion
  • 4.3. Ablation Studies
  • 5. Limitations and Future Work
  • 6. Conclusion
  • 7. Acknowledgment
  • References

Knowls

  1. Knowl 1 — HyperTransformer Framework for Textural and Spectral Feature Fusion

    model/method

    The HyperTransformer framework is designed for hyperspectral (HS) pansharpening by transferring high-resolution (HR) textural features from a panchromatic (PAN) image to low-resolution hyperspectral (LR-HSI) features via a cross-attention mechanism.

    Let p∈R1×H×W\mathbf{p} \in \mathbb{R}^{1 \times H \times W} denote the HR PAN image, p↓↑∈R1×H×W\mathbf{p}\downarrow\uparrow \in \mathbb{R}^{1 \times H \times W} denote the PAN image down-sampled and up-sampled by a scale factor of 4 using bicubic interpolation, and y↑∈RC×H×W\mathbf{y}\uparrow \in \mathbb{R}^{C \times H \times W} denote the LR-HSI up-sampled by a factor of 4 via bicubic interpolation (where CC is the number of spectral bands, and H,WH, W are spatial height and width). Down-and-up sampling of the PAN image ensures domain and resolution consistency with y↑\mathbf{y}\uparrow.

    The Query (QQ), Key (KK), and Value (VV) representations for the attention mechanism are extracted using two dedicated VGG-like feature extractors, fFE-HSI(⋅)f_{\text{FE-HSI}}(\cdot) for HSI and fFE-PAN(⋅)f_{\text{FE-PAN}}(\cdot) for PAN:

    Q=fFE-HSI(y↑)Q = f_{\text{FE-HSI}}(\mathbf{y}\uparrow) K=fFE-PAN(p↓↑)K = f_{\text{FE-PAN}}(\mathbf{p}\downarrow\uparrow) V=fFE-PAN(p)V = f_{\text{FE-PAN}}(\mathbf{p})

    Here, Q∈Rfq×h×wQ \in \mathbb{R}^{f_q \times h \times w}, K∈Rfk×h×wK \in \mathbb{R}^{f_k \times h \times w}, and V∈Rfv×h×wV \in \mathbb{R}^{f_v \times h \times w}, where fq=fk=fvf_q = f_k = f_v denote feature channel dimensions and h,wh, w denote the spatial dimensions of the feature maps.

  2. Knowl 2 — Multi-Head Feature Soft-Attention and Feature Cross-Correlation Embedding

    model/method

    The Multi-Head Feature Soft-Attention (MHFSA) module transfers high-resolution texture details from PAN features to LR-HSI features by measuring cross-feature space dependencies.

    1. Global Feature Descriptor Generation: The queries Q∈Rfq×w×hQ \in \mathbb{R}^{f_q \times w \times h}, keys K∈Rfk×w×hK \in \mathbb{R}^{f_k \times w \times h}, and values V∈Rfv×w×hV \in \mathbb{R}^{f_v \times w \times h} are first flattened into 2D matrices q∈Rfq×whq \in \mathbb{R}^{f_q \times wh}, k∈Rfk×whk \in \mathbb{R}^{f_k \times wh}, and v∈Rfv×whv \in \mathbb{R}^{f_v \times wh}. Using NN separate linear (fully-connected) layers, each feature map is projected into NN global descriptor heads with dimensionality reduction ratio β\beta:

    q(j,i,:)=flinear-qi(q(j,:))\mathbf{q}(j, i, :) = f_{\text{linear-q}}^i(q(j, :)) k(j,i,:)=flinear-ki(k(j,:))\mathbf{k}(j, i, :) = f_{\text{linear-k}}^i(k(j, :)) v(j,i,:)=flinear-vi(v(j,:))\mathbf{v}(j, i, :) = f_{\text{linear-v}}^i(v(j, :))

    where q,k,v∈Rf,N,βwh\mathbf{q}, \mathbf{k}, \mathbf{v} \in \mathbb{R}^{f, N, \beta wh}, i∈{1,…,N}i \in \{1, \dots, N\} indexes the head, and jj indexes the feature channel.

    1. Feature Cross-Correlation Embedding (FCCE): Permuting the first two dimensions yields q′∈RN×fq×βwh\mathbf{q}' \in \mathbb{R}^{N \times f_q \times \beta wh}, k′∈RN×fk×βwh\mathbf{k}' \in \mathbb{R}^{N \times f_k \times \beta wh}, and v′∈RN×fv×βwh\mathbf{v}' \in \mathbb{R}^{N \times f_v \times \beta wh}. The feature cross-correlation matrix C∈RN×fq×fk\mathbb{C} \in \mathbb{R}^{N \times f_q \times f_k} across all NN heads is computed via batch matrix multiplication:

    C=MatMul((q′−mean(q′)),(k′−mean(k′))T)\mathbb{C} = \text{MatMul}\left((\mathbf{q}' - \text{mean}(\mathbf{q}')), (\mathbf{k}' - \text{mean}(\mathbf{k}'))^T\right)

    Row-wise Softmax normalizes the correlations:

    C~=Softmax(C,dim=1)\tilde{\mathbb{C}} = \text{Softmax}(\mathbb{C}, \text{dim}=1)

    1. Multi-Head Soft-Attention & Linear Fusion: The attention weights modulate the value descriptors:

    t=MatMul(C~,v′)∈RN×fq×βwht = \text{MatMul}(\tilde{\mathbb{C}}, \mathbf{v}') \in \mathbb{R}^{N \times f_q \times \beta wh}

    The tensor tt is permuted back to Rfq×N×βwh\mathbb{R}^{f_q \times N \times \beta wh} and passed through a linear layer and reshaping to obtain the transferred texture feature map T∈Rft×h×wT \in \mathbb{R}^{f_t \times h \times w}:

    T=Reshape(Linear(t))T = \text{Reshape}(\text{Linear}(t))

    where ft=fq=fk=fvf_t = f_q = f_k = f_v.

  3. Knowl 3 — Multi-Scale and Textural-Spectral Feature Fusion Architecture

    model/method

    The pansharpening framework integrates features from the attention module into the backbone network at multiple spatial scales using Textural-Spectral Feature Fusion (TSFF) modules.

    Textural-Spectral Feature Fusion (TSFF): At each scale, the texturally enhanced feature map TT produced by the Multi-Head Feature Soft-Attention module is concatenated with the intermediate spectral feature map FF from the backbone network. A 3×33 \times 3 convolutional layer followed by Batch Normalization (BN) fuses the channels:

    F~=BatchNorm(Conv3×3(Concat(T,F)))\tilde{F} = \text{BatchNorm}(\text{Conv}_{3 \times 3}(\text{Concat}(T, F)))

    where F~\tilde{F} serves as the updated feature representation fed forward in the backbone.

    Multi-Scale Feature Fusion (MSFF): HyperTransformer modules are deployed at three separate spatial resolutions across the backbone network:

    • The native LR-HSI spatial scale (denoted by ×1↑\times 1\uparrow).
    • An intermediate 2×2\times up-sampled scale (denoted by ×2↑\times 2\uparrow).
    • The full HR target scale (denoted by ×4↑\times 4\uparrow).

    Injecting texture features across all three spatial scales allows the network to learn multi-scale cross-feature space dependencies and long-range structural details simultaneously.

  4. Knowl 4 — Compound Training Loss for Hyperspectral Pansharpening

    equation

    The overall training objective Loverall\mathcal{L}_{\text{overall}} is a weighted linear combination of an L1L_1 reconstruction loss Lrec\mathcal{L}_{\text{rec}}, a synthesized VGG perceptual loss Lvgg-per\mathcal{L}_{\text{vgg-per}}, and a transfer perceptual loss Lt-per\mathcal{L}_{\text{t-per}}:

    Loverall=λrecLrec+λvgg-perLvgg-per+λt-perLt-per\mathcal{L}_{\text{overall}} = \lambda_{\text{rec}} \mathcal{L}_{\text{rec}} + \lambda_{\text{vgg-per}} \mathcal{L}_{\text{vgg-per}} + \lambda_{\text{t-per}} \mathcal{L}_{\text{t-per}}

    with regularization hyperparameters λrec=1.0\lambda_{\text{rec}} = 1.0, λvgg-per=0.1\lambda_{\text{vgg-per}} = 0.1, and λt-per=0.05\lambda_{\text{t-per}} = 0.05.

    1. Reconstruction Loss: Defined as the mean absolute pixel error between predicted HR-HSI x∈RC×H×W\mathbf{x} \in \mathbb{R}^{C \times H \times W} and reference target HR-HSI xref\mathbf{x}_{\text{ref}}:

    Lrec=1CWH∥xref−x∥1\mathcal{L}_{\text{rec}} = \frac{1}{CWH} \|\mathbf{x}_{\text{ref}} - \mathbf{x}\|_1

    1. Synthesized VGG Perceptual Loss: RGB representations xrefrgb\mathbf{x}_{\text{ref}}^{\text{rgb}} and xrgb\mathbf{x}^{\text{rgb}} are first synthesized from the reference and predicted HSIs using Gaussian-approximated spectral response curves for the red, green, and blue bands. The L2L_2 distance between their features in layer ii of a pre-trained VGG-19 network is computed:

    Lvgg-per=1CiWiHi∥fivgg(xrefrgb)−fivgg(xrgb)∥2\mathcal{L}_{\text{vgg-per}} = \frac{1}{C_i W_i H_i} \|f_i^{\text{vgg}}(\mathbf{x}_{\text{ref}}^{\text{rgb}}) - f_i^{\text{vgg}}(\mathbf{x}^{\text{rgb}})\|_2

    where (Ci,Hi,Wi)(C_i, H_i, W_i) are the channel, height, and width dimensions of layer ii.

    1. Transfer Perceptual Loss: Constrains the predicted HR-HSI feature representations from the HSI feature extractor fFE-HSI(x)sf_{\text{FE-HSI}}(\mathbf{x})^s to match the transferred texture representations TsT^s from HyperTransformer across spatial scales s∈{1,2,4}s \in \{1, 2, 4\}:

    Lt-per=1CsWsHs∥fFE-HSI(x)s−Ts∥2\mathcal{L}_{\text{t-per}} = \frac{1}{C_s W_s H_s} \|f_{\text{FE-HSI}}(\mathbf{x})^s - T^s\|_2

    where (Cs,Ws,Hs)(C_s, W_s, H_s) is the feature dimension at scale ss.

  5. Knowl 5 — Pansharpening Evaluation Protocol and Dataset Benchmarking

    experimental setup

    The pansharpening framework is evaluated on three benchmark datasets:

    • Pavia Center: 102 spectral bands, reference patch dimension 102×160×160102 \times 160 \times 160. RGB composite synthesized using spectral bands 60 (R), 30 (G), and 10 (B).
    • Botswana: 145 spectral bands, reference patch dimension 145×120×120145 \times 120 \times 120. RGB composite synthesized using spectral bands 61 (R), 35 (G), and 10 (B).
    • Chikusei: 128 spectral bands, reference patch dimension 128×256×256128 \times 256 \times 256. RGB composite synthesized using spectral bands 29 (R), 20 (G), and 12 (B).

    Data Degradation (Wald's Protocol): LR-HSIs are generated from reference HSIs using an 8×88 \times 8 Gaussian filter followed by a spatial down-sampling factor of 4. Approximately 80% of cubic patches are randomly selected for training and the remaining 20% for testing.

    Evaluation Metrics: Fusion quality is measured by:

    • Cross-Correlation (CC, ideal = 1.0)
    • Spectral Angle Mapper (SAM, ideal = 0.0)
    • Root Mean Square Error (RMSE ×10−2\times 10^{-2}, ideal = 0.0)
    • Erreur Relative Globale Adimensionnelle de Synthèse (ERGAS, ideal = 0.0)
    • Peak Signal-to-Noise Ratio (PSNR in dB, ideal = ∞\infty)
  6. Knowl 6 — Quantitative Evaluation of HyperTransformer Across Hyperspectral Datasets

    data/table

    Quantitative comparison of HyperTransformer against classical and deep learning pansharpening methods across Pavia Center, Botswana, and Chikusei datasets under Wald's protocol (4×4\times spatial scaling):

    Method Pavia Center Botswana Chikusei
    CC SAM RMSE ERGAS PSNR CC SAM RMSE ERGAS PSNR CC SAM RMSE ERGAS PSNR
    PCA 0.845 8.92 3.45 6.64 31.26 0.943 2.38 1.98 2.22 40.03 0.887 6.99 2.47 7.71 30.98
    GFPCA 0.902 8.31 3.98 7.44 29.09 0.901 2.66 2.45 2.75 37.83 0.883 4.76 1.98 7.00 30.96
    BF 0.918 9.60 3.44 6.63 30.22 0.931 2.47 1.88 2.34 40.01 0.903 5.15 1.94 6.62 37.89
    BFS 0.925 8.10 3.05 6.00 31.09 0.932 2.39 1.85 2.32 40.15 0.917 4.69 1.72 6.39 37.99
    SFIM 0.946 6.76 2.55 5.43 32.61 0.932 3.44 2.81 2.25 39.58 0.928 3.79 1.43 6.43 39.55
    GS 0.961 6.62 2.55 4.95 32.93 0.946 2.34 1.93 2.17 40.14 0.733 5.64 2.96 8.17 35.13
    GSA 0.950 7.15 2.34 4.70 33.52 0.955 2.04 1.59 1.85 41.89 0.943 3.52 1.42 4.30 41.38
    MGH 0.955 6.81 2.25 4.77 33.97 0.960 2.07 1.54 1.69 42.43 0.929 3.82 1.45 6.40 39.85
    CNMF 0.960 6.64 2.20 4.39 34.14 0.942 2.61 1.73 2.10 40.98 0.900 4.72 1.91 5.75 39.65
    MG 0.956 6.55 2.20 4.45 34.12 0.960 2.02 1.51 1.68 42.47 0.938 3.81 1.52 4.41 41.05
    HySure 0.966 6.13 1.80 3.77 35.91 0.956 2.15 1.46 1.77 42.30 0.960 2.98 1.13 3.69 43.14
    HyperPNN 0.967 6.09 1.67 3.82 36.70 0.970 1.67 1.15 1.44 44.45 0.946 3.97 1.11 4.77 41.57
    PanNet 0.968 6.36 1.83 3.89 35.61 0.926 2.17 1.53 2.82 40.41 0.956 3.79 0.88 5.32 41.90
    DARN 0.969 6.43 1.56 3.95 37.30 0.973 1.58 1.09 1.35 44.42 0.953 3.60 1.05 4.44 42.24
    HyperKite 0.980 5.61 1.29 2.85 38.65 0.979 1.46 1.01 1.21 45.53 0.974 2.85 1.03 3.62 43.53
    SIPSA 0.948 5.27 2.38 4.52 33.65 0.901 2.34 2.20 2.54 38.55 0.947 2.87 1.06 5.09 41.02
    GPPNN 0.963 6.52 1.91 4.05 35.36 0.962 1.90 1.36 1.65 43.01 0.970 2.75 0.66 4.24 44.07
    HyperTransformer (Ours) 0.989 3.85 0.87 2.01 43.80 0.982 1.18 0.89 1.04 46.97 0.980 2.40 0.57 3.12 45.87

    RMSE values are scaled by ×10−2\times 10^{-2}. HyperTransformer achieves top performance across all five evaluation metrics on all three datasets, outperforming the closest competing state-of-the-art method.

  7. Knowl 7 — Ablation Analysis of Attention Mechanism and Attention Head Count

    data/table

    Ablation study on the Pavia Center dataset evaluating the effect of the Multi-Head Feature Soft-Attention (MHFSA) module and the number of global descriptor heads NN. The baseline model (B/L) bypasses MHFSA by directly setting the transferred texture features TT to the PAN features VV:

    NN CC SAM RMSE (×10−2\times 10^{-2}) ERGAS PSNR (dB)
    B/L 0.981 4.88 1.33 2.84 38.71
    1 0.975 4.32 1.03 2.31 40.59
    2 0.976 4.19 0.95 2.18 42.52
    8 0.987 4.06 0.92 2.13 43.20
    16 0.989 3.85 0.87 2.01 43.80
    32 0.988 4.02 0.90 2.10 43.47
    64 0.987 4.04 0.91 2.12 43.19

    Incorporating MHFSA with N=16N=16 improves baseline performance by ∼1.5%\sim 1.5\% in CC, 21%21\% in SAM, 35%35\% in RMSE, 29%29\% in ERGAS, and 13%13\% (5.09 dB) in PSNR. Increasing NN beyond 16 leads to metric saturation or slight degradation, identifying N=16N=16 as the optimal number of heads.

  8. Knowl 8 — Ablation Analysis of Multi-Scale Injection and Loss Function Terms

    data/table

    Ablation studies on the Pavia Center dataset showing the effect of multi-scale injection configurations and individual loss components.

    Multi-Scale Feature Fusion Ablation: Injections tested across spatial scales ×1↑\times 1\uparrow, ×2↑\times 2\uparrow, and ×4↑\times 4\uparrow (B/L denotes no PAN feature injection at any scale):

    ×1\times 1 ×2\times 2 ×4\times 4 CC SAM RMSE (×10−2\times 10^{-2}) ERGAS PSNR (dB)
    B/L (none) 0.956 4.86 2.04 3.90 35.81
    ✓ 0.975 4.76 1.49 3.00 38.42
    ✓ 0.985 4.29 1.08 2.40 40.80
    ✓ ✓ 0.985 4.41 1.09 2.38 41.02
    ✓ 0.986 4.01 0.98 2.21 42.88
    ✓ ✓ 0.988 3.96 0.92 2.20 43.58
    ✓ ✓ 0.988 3.85 0.89 2.09 43.60
    ✓ ✓ ✓ 0.989 3.85 0.87 2.01 43.80

    Using HyperTransformers at all three spatial scales improves performance over using only the full ×4\times 4 scale by ∼0.3%\sim 0.3\% in CC, 3.7%3.7\% in SAM, 11.2%11.2\% in RMSE, 9.0%9.0\% in ERGAS, and 2.1%2.1\% (0.92 dB) in PSNR.

    Loss Function Component Ablation:

    L1\mathcal{L}_1 Lvgg-per\mathcal{L}_{\text{vgg-per}} Lt-per\mathcal{L}_{\text{t-per}} CC SAM RMSE (×10−2\times 10^{-2}) ERGAS PSNR (dB)
    ✓ 0.987 4.07 0.91 2.12 43.00
    ✓ ✓ 0.988 4.01 0.90 2.09 43.60
    ✓ ✓ ✓ 0.989 3.85 0.87 2.01 43.80

    Adding Lvgg-per\mathcal{L}_{\text{vgg-per}} to L1\mathcal{L}_1 yields notable PSNR gains (+0.60 dB), while further adding Lt-per\mathcal{L}_{\text{t-per}} improves spectral and error metrics (SAM from 4.01 to 3.85, RMSE from 0.0090 to 0.0087).

  9. Knowl 9 — Spectral Reconstruction Distortion in Ultraviolet and Infrared Bands

    limitation

    Spectral error analysis reveals relatively high Mean Absolute Error (MAE) in the reconstructed HSIs near the ultraviolet spectrum (bands 1–10) and near-infrared/infrared spectrum (bands 90–104). This localized distortion occurs because the single broadband PAN image does not contain sufficient ultraviolet and infrared spectral features to provide adequate textural and spectral guidance for those wavelength intervals during feature transfer.

Coverage note — None was omitted; all contributed models, attention mechanisms, loss formulations, datasets, benchmark results, ablation studies, and limitations are fully covered.

References

  1. 1.Bruno Aiazzi, Luciano Alparone, Stefano Baronti, and Andrea Garzelli. Context-driven fusion of high spatial and spectral resolution images based on oversampled multiresolution analysis. IEEE Transactions on geoscience and remote sensing, 40(10):2300–2312, 2002.
  2. 2.B Aiazzi, L Alparone, S Baronti, A Garzelli, and M Selva. Mtf-tailored multiscale fusion of high-resolution ms and pan imagery. Photogrammetric Engineering & Remote Sensing, 72(5):591–596, 2006.
  3. 3.Bruno Aiazzi, Stefano Baronti, and Massimo Selva. Improving component substitution pansharpening through multivariate regression of ms + pan data. IEEE Transactions on Geoscience and Remote Sensing, 45(10):3230–3239, 2007.
  4. 4.Coloma Ballester, Vicent Caselles, Laura Igual, Joan Verdera, and Bernard Rouge. A variational model for p+ ´xs image fusion. International Journal of Computer Vision, 69(1):43–58, 2006.
  5. 5.Wele Gedara Chaminda Bandara and Vishal M Patel. A transformer-based siamese network for change detection. arXiv preprint arXiv:2201.01293, 2022.
  6. 6.Wele Gedara Chaminda Bandara, Jeya Maria Jose Valanarasu, and Vishal M Patel. Spin road mapper: Extracting roads from aerial images via spatial and interaction space graph reasoning for autonomous driving. arXiv preprint arXiv:2109.07701, 2021.
  7. 7.Wele Gedara Chaminda Bandara, Jeya Maria Jose Valanarasu, and Vishal M. Patel. Hyperspectral pansharpening based on improved deep image prior and residual reconstruction. IEEE Transactions on Geoscience and Remote Sensing, 60:1–16, 2022.
  8. 8.Jose M Bioucas-Dias, Antonio Plaza, Nicolas Dobigeon, ´Mario Parente, Qian Du, Paul Gader, and Jocelyn Chanussot. Hyperspectral unmixing overview: Geometrical, statistical, and sparse regression-based approaches. IEEE journal of selected topics in applied earth observations and remote sensing, 5(2):354–379, 2012.
  9. 9.Peter J Burt and Edward H Adelson. The laplacian pyramid as a compact image code. In Readings in computer vision, pages 671–679. Elsevier, 1987.
  10. 10.Wjoseph Carper, Thomasm Lillesand, and Ralphw Kiefer. The use of intensity-hue-saturation transformations for merging spot panchromatic and multispectral image data. Photogrammetric Engineering and remote sensing, 56(4):459–467, 1990.
  11. 11.Renwei Dian, Shutao Li, Anjing Guo, and Leyuan Fang. Deep hyperspectral image sharpening. IEEE Transactions on Neural Networks and Learning Systems, 29(11):5345–5355, 2018.
  12. 12.Xueyang Fu, Zihuang Lin, Yue Huang, and Xinghao Ding. A variational pan-sharpening with local gradient constraints. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10265–10274, 2019.
  13. 13.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016.
  14. 14.Lin He, Jiawei Zhu, Jun Li, Antonio Plaza, Jocelyn Chanussot, and Bo Li. Hyperpnn: Hyperspectral pansharpening via spectrally predictive convolutional neural networks. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 12(8):3092–3100, 2019.
  15. 15.Xiyan He, Laurent Condat, Jose M Bioucas-Dias, Jocelyn ´Chanussot, and Junshi Xia. A new pansharpening method based on spatial and spectral sparsity priors. IEEE Transactions on Image Processing, 23(9):4160–4174, 2014.
  16. 16.Justin Johnson, Alexandre Alahi, and Li Fei-Fei. Perceptual losses for real-time style transfer and super-resolution. In European conference on computer vision, pages 694–711. Springer, 2016.
  17. 17.P Kwarteng and A Chavez. Extracting spectral contrast in landsat thematic mapper image data using selective principal component analysis. Photogramm. Eng. Remote Sens, 55(1):339–348, 1989.
  18. 18.P Kwarteng and A Chavez. Extracting spectral contrast in landsat thematic mapper image data using selective principal component analysis. Photogramm. Eng. Remote Sens, 55(1):339–348, 1989.
  19. 19.Craig A Laben and Bernard V Brower. Process for enhancing the spatial resolution of multispectral imagery using pansharpening, Jan. 4 2000. US Patent 6,011,875.
  20. 20.Craig A Laben and Bernard V Brower. Process for enhancing the spatial resolution of multispectral imagery using pansharpening, Jan. 4 2000. US Patent 6,011,875.
  21. 21.Florence Laporterie-Dejean, H ´ el´ ene de Boissezon, Guy Flouzat, and Marie-Jose Lef ´ evre-Fonollosa. Thematic and statistical evaluations of five panchromatic/multispectral fusion methods on simulated pleiades-hr images. Information Fusion, 6(3):193–212, 2005.
  22. 22.Jaehyup Lee, Soomin Seo, and Munchurl Kim. Sipsa-net: Shift-invariant pan sharpening with moving object alignment for satellite imagery. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021.
  23. 23.Jaehyup Lee, Soomin Seo, and Munchurl Kim. Sipsa-net: Shift-invariant pan sharpening with moving object alignment for satellite imagery. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10166–10174, 2021.
  24. 24.Jaehyup Lee, Soomin Seo, and Munchurl Kim. Sipsa-net: Shift-invariant pan sharpening with moving object alignment for satellite imagery. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10166–10174, 2021.
  25. 25.Wenzhi Liao, Frieke Van Coillie, Sidharta Gautama, Aleksandra Piurica, and Wilfried Philips. Fusion of thermal infrared hyperspectral and vis rgb data using guided filter and supervised fusion graph. 2014.
  26. 26.GA Licciardi, Alberto Villa, Muhammad Murtaza Khan, and Jocelyn Chanussot. Image fusion and spectral unmixing of hyperspectral images for spatial improvement of classification maps. In 2012 IEEE International Geoscience and Remote Sensing Symposium, pages 7290–7293. IEEE, 2012.
  27. 27.JG Liu. Smoothing filter-based intensity modulation: A spectral preserve image fusion technique for improving spatial details. International Journal of Remote Sensing, 21(18):3461–3472, 2000.
  28. 28.Laetitia Loncan, Luis B De Almeida, Jose M BioucasDias, Xavier Briottet, Jocelyn Chanussot, Nicolas Dobigeon, Sophie Fabre, Wenzhi Liao, Giorgio A Licciardi, Miguel Simoes, et al. Hyperspectral pansharpening: A review. IEEE Geoscience and remote sensing magazine, 3(3):27–46, 2015.
  29. 29.Stephane G Mallat. A theory for multiresolution signal decomposition: the wavelet representation. In Fundamental Papers in Wavelet Theory, pages 494–513. Princeton University Press, 2009.
  30. 30.Giuseppe Masi, Davide Cozzolino, Luisa Verdoliva, and Giuseppe Scarpa. Pansharpening by convolutional neural networks. Remote Sensing, 8(7):594, 2016.
  31. 31.Ali Mohammadzadeh, Ahad Tavakoli, and Mohammad J Valadan Zoej. Road extraction based on fuzzy logic and mathematical morphology from pan-sharpened ikonos images. The photogrammetric record, 21(113):44–60, 2006.
  32. 32.Frosti Palsson, Johannes R Sveinsson, and Magnus O Ulfarsson. A new pansharpening algorithm based on total variation. IEEE Geoscience and Remote Sensing Letters, 11(1):318–322, 2013.
  33. 33.Frosti Palsson, Johannes R Sveinsson, and Magnus O Ulfarsson. Multispectral and hyperspectral image fusion using a 3-d-convolutional neural network. IEEE Geoscience and Remote Sensing Letters, 14(5):639–643, 2017.
  34. 34.Antonio Plaza, Jon Atli Benediktsson, Joseph W. Boardman, Jason Brazile, Lorenzo Bruzzone, Gustavo Camps-Valls, Jocelyn Chanussot, Mathieu Fauvel, Paolo Gamba, Anthony Gualtieri, Mattia Marconcini, James C. Tilton, and Giovanna Trianni. Recent advances in techniques for hyperspectral image processing. Remote Sensing of Environment, 113:S110–S122, 2009. Imaging Spectroscopy Special Issue.
  35. 35.Miguel Simoes, Jose Bioucas-Dias, Luis B Almeida, and ´Jocelyn Chanussot. A convex formulation for hyperspectral image superresolution via subspace-based regularization. IEEE Transactions on Geoscience and Remote Sensing, 53(6):3373–3388, 2014.
  36. 36.Karen Simonyan and Andrew Zisserman. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556, 2014.
  37. 37.Carlos Souza Jr, Laurel Firestone, Luciano Moreira Silva, and Dar Roberts. Mapping forest degradation in the eastern amazon from spot 4 through spectral mixture models. Remote sensing of environment, 87(4):494–506, 2003.
  38. 38.S.G. Ungar, J.S. Pearlman, J.A. Mendenhall, and D. Reuter. Overview of the earth observing one (eo-1) mission. IEEE Transactions on Geoscience and Remote Sensing, 41(6):1149–1159, 2003.
  39. 39.Gemine Vivone, Luciano Alparone, Jocelyn Chanussot, Mauro Dalla Mura, Andrea Garzelli, Giorgio A Licciardi, Rocco Restaino, and Lucien Wald. A critical comparison among pansharpening algorithms. IEEE Transactions on Geoscience and Remote Sensing, 53(5):2565–2586, 2014.
  40. 40.Lucien Wald. Quality of high resolution synthesised images: Is there a simple criterion? In Third conference" Fusion of Earth data: merging point measurements, raster maps and remotely sensed images", pages 99–103. SEE/URISCA, 2000.
  41. 41.Jiaming Wang, Zhenfeng Shao, Xiao Huang, Tao Lu, and Ruiqian Zhang. A dual-path fusion network for pansharpening. IEEE Transactions on Geoscience and Remote Sensing, 2021.
  42. 42.Qi Wei, Jose Bioucas-Dias, Nicolas Dobigeon, and Jean- ´Yves Tourneret. Hyperspectral and multispectral image fusion based on a sparse representation. IEEE Transactions on Geoscience and Remote Sensing, 53(7):3658–3668, 2015.
  43. 43.Qi Wei, Nicolas Dobigeon, and Jean-Yves Tourneret. Bayesian fusion of multi-band images. IEEE Journal of Selected Topics in Signal Processing, 9(6):1117–1127, 2015.
  44. 44.Yancong Wei, Qiangqiang Yuan, Huanfeng Shen, and Liangpei Zhang. Boosting the accuracy of multispectral image pansharpening by learning a deep residual network. IEEE Geoscience and Remote Sensing Letters, 14(10):1795–1799, 2017.
  45. 45.Xiao Wu, Ting-Zhu Huang, Liang-Jian Deng, and Tian-Jing Zhang. Dynamic cross feature fusion for remote sensing pansharpening. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 14687–14696, 2021.
  46. 46.Shuang Xu, Jiangshe Zhang, Zixiang Zhao, Kai Sun, Junmin Liu, and Chunxia Zhang. Deep gradient projection networks for pan-sharpening. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1366–1375, 2021.
  47. 47.Fuzhi Yang, Huan Yang, Jianlong Fu, Hongtao Lu, and Baining Guo. Learning texture transformer network for image super-resolution. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 5791–5800, 2020.
  48. 48.Junfeng Yang, Xueyang Fu, Yuwen Hu, Yue Huang, Xinghao Ding, and John Paisley. Pannet: A deep network architecture for pan-sharpening. In Proceedings of the IEEE international conference on computer vision, pages 5449–5457, 2017.
  49. 49.Junfeng Yang, Xueyang Fu, Yuwen Hu, Yue Huang, Xinghao Ding, and John Paisley. Pannet: A deep network architecture for pan-sharpening. In Proceedings of the IEEE international conference on computer vision, pages 5449–5457, 2017.
  50. 50.Jing Yao, Danfeng Hong, Jocelyn Chanussot, Deyu Meng, Xiaoxiang Zhu, and Zongben Xu. Cross-attention in coupled unmixing nets for unsupervised hyperspectral superresolution. In European Conference on Computer Vision, pages 208–224. Springer, 2020.
  51. 51.Jing Yao, Danfeng Hong, Jocelyn Chanussot, Deyu Meng, Xiaoxiang Zhu, and Zongben Xu. Cross-attention in coupled unmixing nets for unsupervised hyperspectral superresolution. In European Conference on Computer Vision, pages 208–224. Springer, 2020.
  52. 52.Naoto Yokoya and Akira Iwasaki. Airborne hyperspectral data over chikusei. Space Appl. Lab., Univ. Tokyo, Tokyo, Japan, Tech. Rep. SAL-2016-05-27, 2016.
  53. 53.Naoto Yokoya and Akira Iwasaki. Airborne hyperspectral data over chikusei. Space Appl. Lab., Univ. Tokyo, Tokyo, Japan, Tech. Rep. SAL-2016-05-27, 2016.
  54. 54.Naoto Yokoya, Takehisa Yairi, and Akira Iwasaki. Coupled nonnegative matrix factorization unmixing for hyperspectral and multispectral data fusion. IEEE Transactions on Geoscience and Remote Sensing, 50(2):528–537, 2011.
  55. 55.Naoto Yokoya, Takehisa Yairi, and Akira Iwasaki. Coupled nonnegative matrix factorization unmixing for hyperspectral and multispectral data fusion. IEEE Transactions on Geoscience and Remote Sensing, 50(2):528–537, 2011.
  56. 56.Yongnian Zeng, Wei Huang, Maoguo Liu, Honghui Zhang, and Bin Zou. Fusion of satellite images in urban area: Assessing the quality of resulting images. In 2010 18th International Conference on Geoinformatics, pages 1–4. IEEE, 2010.
  57. 57.Lei Zhang, Jiangtao Nie, Wei Wei, Yanning Zhang, Shengcai Liao, and Ling Shao. Unsupervised adaptation learning for hyperspectral imagery super-resolution. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3073–3082, 2020.
  58. 58.Yuxuan Zheng, Jiaojiao Li, Yunsong Li, Jie Guo, Xianyun Wu, and Jocelyn Chanussot. Hyperspectral pansharpening using deep prior and dual attention residual network. IEEE Transactions on Geoscience and Remote Sensing, 58(11):8059–8076, 2020.
  59. 59.Man Zhou, Xueyang Fu, Jie Huang, Feng Zhao, Aiping Liu, and Rujing Wang. Effective pan-sharpening with transformer and invertible neural network. IEEE Transactions on Geoscience and Remote Sensing, 60:1–15, 2022.
  60. 60.Xiao Xiang Zhu and Richard Bamler. A sparse image fusion algorithm with application to pan-sharpening. IEEE transactions on geoscience and remote sensing, 51(5):2827–2836, 2012.
  61. 61.Xiao Xiang Zhu, Claas Grohnfeldt, and Richard Bamler. Exploiting joint sparsity for pansharpening: The j-sparsefi algorithm. IEEE Transactions on Geoscience and Remote Sensing, 54(5):2664–2681, 2015.

Citation

MLA
Bandara, W. G. C., and V. M. Patel. “HyperTransformer: A Textural and Spectral Feature Fusion Transformer for Pansharpening”. arXiv, 2022, http://arxiv.org/abs/2203.02503v3.
APA
Bandara, W. G. C., & Patel, V. M. (2022). HyperTransformer: A Textural and Spectral Feature Fusion Transformer for Pansharpening. arXiv. http://arxiv.org/abs/2203.02503v3
Chicago
Bandara, W. G. C., and V. M. Patel. 2022. “HyperTransformer: A Textural and Spectral Feature Fusion Transformer for Pansharpening”. arXiv. http://arxiv.org/abs/2203.02503v3.
Harvard
Bandara, W.G.C. and Patel, V.M. (2022) “HyperTransformer: A Textural and Spectral Feature Fusion Transformer for Pansharpening”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2203.02503v3.
Vancouver
1. Bandara WGC, Patel VM (2022) HyperTransformer: A Textural and Spectral Feature Fusion Transformer for Pansharpening. arXiv

BibTeX

@article{bandara2022hypertransformer,
  title = {HyperTransformer: A Textural and Spectral Feature Fusion Transformer for Pansharpening},
  author = {Bandara, Wele Gedara Chaminda and Patel, Vishal M.},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2203.02503v3},
  eprint = {2203.02503}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE