Optimization-Inspired Cross-Attention Transformer for Compressive Sensing

Jiechong SongChong MouShiqi WangSiwei MaJian Zhang

article2023CVPR105 citations

Proposes a lightweight deep unfolding framework that integrates cross-attention mechanisms directly into the iterative optimization steps of compressive sensing, preserving inter-stage feature information to achieve state-of-the-art image reconstruction with substantially fewer parameters.

Listen

Compressive sensing is a critical technology used across medical imaging, remote monitoring, and camera systems to capture and store signals efficiently using far fewer measurements than traditional methods. While modern deep unfolding networks have improved image reconstruction by merging classical optimization algorithms with deep neural networks, existing solutions face notable bottlenecks. They typically require massive numbers of parameters and heavy computational budgets, while suffering from feature information loss because they pass data between processing stages primarily at the image level rather than retaining rich internal representations.

To overcome these limitations, the article evaluates and demonstrates a lightweight reconstruction model called the Optimization-inspired Cross-attention Transformer Unfolding Framework. The main objective is to establish an interpretable, highly accurate deep unfolding framework that operates directly in feature space and significantly reduces memory and computational requirements.

The authors conducted extensive computational experiments to validate this approach. The system incorporates an iterative module consisting of a Dual Cross Attention sub-module—which includes an Inertia-Supplied Cross Attention block and a Projection-Guided Cross Attention block—paired with a Feed-Forward Network. The model was trained on 400 images from the standard BSD500 benchmark dataset and tested on two widely recognized evaluation datasets, Set11 and Urban100, across various compression ratios ranging from 10% to 50%.

The experimental findings show that the proposed framework consistently outperforms existing state-of-the-art models across standard image quality metrics. On the Set11 benchmark, the enhanced configuration of the framework achieved an average reconstruction quality improvement of 0.29 dB to 3.91 dB over nine competing modern methods. Visual evaluations revealed sharper edges and structural details compared to alternative approaches. Crucially, the model achieved these results while dramatically reducing operational overhead: at a 10% sampling ratio, the base framework required only 0.40 million parameters and 189.3 billion floating-point operations, compared to 16.90 million parameters and over 13,391 billion operations for leading alternatives. Additional robustness tests demonstrated that the architecture maintains stable reconstruction quality when subjected to varying levels of Gaussian noise.

These results demonstrate that high-performance image reconstruction does not require computationally heavy models. By successfully passing multi-channel feature information across iterations and integrating inertia forces into the optimization steps, the proposed framework lowers deployment costs, decreases hardware memory demands, and accelerates processing times without sacrificing image fidelity. This makes advanced compressive sensing significantly more practical for resource-constrained edge devices and real-time medical or remote imaging systems.

Stakeholders and engineering teams developing imaging pipelines should consider adopting cross-attention feature-space unfolding frameworks to optimize efficiency and reconstruction fidelity. Before deploying to production environments, organizations should conduct pilot testing on domain-specific hardware to evaluate real-time throughput. A key limitation noted is the reliance on synthetic image datasets and artificial Gaussian noise rather than open, real-world compressive sensing datasets. Future efforts should focus on validating performance on specialized real-world hardware, adapting the framework to broader image restoration inverse problems, and extending its capabilities to video processing applications.

Cover for Optimization-Inspired Cross-Attention Transformer for Compressive Sensing

Abstract

By integrating certain optimization solvers with deep neural networks, deep unfolding network (DUN) with good interpretability and high performance has attracted growing attention in compressive sensing (CS). However, existing DUNs often improve the visual quality at the price of a large number of parameters and have the problem of feature information loss during iteration. In this paper, we propose an Optimization-inspired Cross-attention Transformer (OCT) module as an iterative process, leading to a lightweight OCT-based Unfolding Framework (OCTUF) for image CS. Specifically, we design a novel Dual Cross Attention (Dual-CA) sub-module, which consists of an Inertia-Supplied Cross Attention (ISCA) block and a Projection-Guided Cross Attention (PGCA) block. ISCA block introduces multi-channel inertia forces and increases the memory effect by a cross attention mechanism between adjacent iterations. And, PGCA block achieves an enhanced information interaction, which introduces the inertia force into the gradient descent step through a cross attention block. Extensive CS experiments manifest that our OCTUF achieves superior performance compared to state-of-the-art methods while training lower complexity. Codes are available at https://github.com/songjiechong/OCTUF.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 2.1. Deep Unfolding Network
  • 2.2. Vision Transformer
  • 3. Proposed Method
  • 3.1. Overall Architecture
  • 3.2. Dual Cross Attention
  • 3.3. Loss Function
  • 4. Experiments
  • 4.1. Implementation Details
  • 4.2. Qualitative Evaluation
  • 4.3. Ablation Study
  • 4.4. Complexity Analysis
  • 4.5. Sensitivity to Noise
  • 5. Conclusion
  • References

Knowls

  1. Knowl 1 — OCTUF (Optimization-Inspired Cross-Attention Transformer Unfolding Framework)

    model/method

    The Optimization-inspired Cross-attention Transformer Unfolding Framework (OCTUF) is a deep unfolding network for image compressive sensing (CS) that operates entirely in the feature domain. Inspired by the inertial proximal algorithm for nonconvex optimization (iPiano), OCTUF unrolls an iterative optimization process over KK stages. For an input measurement y=Φx∈RMy = \Phi x \in \mathbb{R}^M, where x∈RNx \in \mathbb{R}^N is the original signal and Φ∈RM×N\Phi \in \mathbb{R}^{M \times N} is the measurement matrix (M≪NM \ll N):

    1. Initialization: The initial reconstruction x(0)=Φ⊤yx^{(0)} = \Phi^\top y is mapped to feature space via a 3×33 \times 3 convolution Conv0(⋅)\text{Conv}_0(\cdot) to produce X(0)∈RH×W×CX^{(0)} \in \mathbb{R}^{H \times W \times C}, where HH and WW denote spatial dimensions and CC denotes the feature channel depth.

    2. Iterative Stage: For each iteration k∈{1,2,…,K}k \in \{1, 2, \dots, K\}, the Optimization-inspired Cross-attention Transformer (OCT) module executes:

    S(k)=HDual-CA(X(k−1),Z(k−2))S^{(k)} = H_{\text{Dual-CA}}(X^{(k-1)}, Z^{(k-2)})

    X(k)=HFFN(S(k))X^{(k)} = H_{\text{FFN}}(S^{(k)})

    where S(k),X(k)∈RH×W×CS^{(k)}, X^{(k)} \in \mathbb{R}^{H \times W \times C} represent intermediate and refined feature tensors. The auxiliary feature tensor Z(k−2)∈RH×W×(C−1)Z^{(k-2)} \in \mathbb{R}^{H \times W \times (C-1)} consists of the last C−1C-1 channels clipped from the previous iteration's feature map X(k−2)X^{(k-2)} (for k=1k=1, Z(k−2)Z^{(k-2)} is omitted). HDual-CA(⋅)H_{\text{Dual-CA}}(\cdot) is the Dual Cross Attention sub-module representing the projection step, and HFFN(⋅)H_{\text{FFN}}(\cdot) is the Feed-Forward Network sub-module representing the proximal mapping (denoising) step.

    1. Reconstruction: The final restored image x^\hat{x} is extracted by taking the first channel from the terminal stage output X(K)∈RH×W×CX^{(K)} \in \mathbb{R}^{H \times W \times C}.
  2. Knowl 2 — Dual Cross Attention (Dual-CA) Sub-Module

    model/method

    The Dual Cross Attention (Dual-CA) sub-module HDual-CAH_{\text{Dual-CA}} performs the gradient projection step of the unfolded iPiano optimization in feature space. Given the prior stage feature X(k−1)∈RH×W×CX^{(k-1)} \in \mathbb{R}^{H \times W \times C}, Dual-CA splits it into two distinct chunks:

    1. r(k−1)∈RH×W×1r^{(k-1)} \in \mathbb{R}^{H \times W \times 1} (the first channel), representing the spatial image-level proxy used as the input to the gradient descent term.
    2. Z(k−1)∈RH×W×(C−1)Z^{(k-1)} \in \mathbb{R}^{H \times W \times (C-1)} (the remaining C−1C-1 channels), representing the multi-channel inertia state.

    Dual-CA sequentially executes an Inertia-Supplied Cross Attention (ISCA) block followed by a Projection-Guided Cross Attention (PGCA) block:

    S(k)=HPGCA(F(r(k−1)),HISCA(Z(k−1),Z(k−2)))S^{(k)} = H_{\text{PGCA}}\left(\mathcal{F}\left(r^{(k-1)}\right), H_{\text{ISCA}}\left(Z^{(k-1)}, Z^{(k-2)}\right)\right)

    where F(r(k−1))\mathcal{F}(r^{(k-1)}) represents the data-fidelity gradient descent operation on the image proxy, HISCAH_{\text{ISCA}} aggregates inter-iteration inertial memory from states Z(k−1)Z^{(k-1)} and Z(k−2)Z^{(k-2)}, and HPGCAH_{\text{PGCA}} fuses the gradient-updated spatial component with the updated inertial feature representation into S(k)∈RH×W×CS^{(k)} \in \mathbb{R}^{H \times W \times C}.

  3. Knowl 3 — Channel-Wise Cross Attention (CA) Block

    model/method

    The Cross Attention (CA) block GCA(V,K,Q)G_{\text{CA}}(V, K, Q) aggregates context across different components along the channel dimension with linear computational complexity with respect to spatial dimensions H×WH \times W. Given inputs V,K,QV, K, Q where V,K∈RH×W×(C−1)V, K \in \mathbb{R}^{H \times W \times (C-1)} and Q∈RH×W×CqQ \in \mathbb{R}^{H \times W \times C_q} (with Cq=C−1C_q=C-1 or Cq=1C_q=1):

    1. Feature Projection and Tokenization: Each input is transformed using a 1×11 \times 1 convolution ConvV,K,Q(⋅)\text{Conv}_{V,K,Q}(\cdot) to map features to channel depth C−1C-1, a 3×33 \times 3 depth-wise convolution DconvV,K,Q(⋅)\text{Dconv}_{V,K,Q}(\cdot) to encode spatial context, and a spatial flattening/reshape operator R(⋅)\mathcal{R}(\cdot):

    V^=R(DconvV(ConvV(V)))∈RHW×(C−1)\hat{V} = \mathcal{R}(\text{Dconv}_V(\text{Conv}_V(V))) \in \mathbb{R}^{HW \times (C-1)}

    K^=R(DconvK(ConvK(K)))∈RHW×(C−1)\hat{K} = \mathcal{R}(\text{Dconv}_K(\text{Conv}_K(K))) \in \mathbb{R}^{HW \times (C-1)}

    Q^=R(DconvQ(ConvQ(Q)))∈RHW×(C−1)\hat{Q} = \mathcal{R}(\text{Dconv}_Q(\text{Conv}_Q(Q))) \in \mathbb{R}^{HW \times (C-1)}

    1. Cross-Channel Attention Map: Transposed matrix multiplication across spatial tokens produces a channel-wise correlation matrix re-weighted by softmax:

    A=Softmax(K^⊤Q^)∈R(C−1)×(C−1)A = \text{Softmax}(\hat{K}^\top \hat{Q}) \in \mathbb{R}^{(C-1) \times (C-1)}

    1. Feature Aggregation: The value tokens V^\hat{V} are aggregated via matrix multiplication with attention map AA, reshaped back to spatial dimensions, and projected through a 1×11 \times 1 convolution ConvA(⋅)\text{Conv}_A(\cdot):

    GCA(V,K,Q)=ConvA(R(V^A))∈RH×W×(C−1)G_{\text{CA}}(V, K, Q) = \text{Conv}_A(\mathcal{R}(\hat{V} A)) \in \mathbb{R}^{H \times W \times (C-1)}

  4. Knowl 4 — Inertia-Supplied Cross Attention (ISCA) Block

    model/method

    The Inertia-Supplied Cross Attention (ISCA) block replaces standard vector subtraction α(x(k−1)−x(k−2))\alpha(x^{(k-1)} - x^{(k-2)}) with an adaptive multi-channel cross-attention mechanism across adjacent iterations. Given inertial features Z(k−1)∈RH×W×(C−1)Z^{(k-1)} \in \mathbb{R}^{H \times W \times (C-1)} from iteration k−1k-1 and Z(k−2)∈RH×W×(C−1)Z^{(k-2)} \in \mathbb{R}^{H \times W \times (C-1)} from iteration k−2k-2:

    1. Query, Key, and Value inputs are constructed using Layer Normalization (LN):

    VISCA(k)=LN(Z(k−2)),KISCA(k)=LN(Z(k−2)),QISCA(k)=LN(Z(k−1))V_{\text{ISCA}}^{(k)} = \text{LN}(Z^{(k-2)}), \quad K_{\text{ISCA}}^{(k)} = \text{LN}(Z^{(k-2)}), \quad Q_{\text{ISCA}}^{(k)} = \text{LN}(Z^{(k-1)})

    1. The updated inertial feature Z^(k−1)=HISCA(Z(k−1),Z(k−2))\hat{Z}^{(k-1)} = H_{\text{ISCA}}(Z^{(k-1)}, Z^{(k-2)}) is computed by adding a residual connection to the CA block output:

    Z^(k−1)=GCA(VISCA(k),KISCA(k),QISCA(k))+Z(k−1)\hat{Z}^{(k-1)} = G_{\text{CA}}\left(V_{\text{ISCA}}^{(k)}, K_{\text{ISCA}}^{(k)}, Q_{\text{ISCA}}^{(k)}\right) + Z^{(k-1)}

    where Z^(k−1)∈RH×W×(C−1)\hat{Z}^{(k-1)} \in \mathbb{R}^{H \times W \times (C-1)} captures inter-stage memory and momentum information.

  5. Knowl 5 — Projection-Guided Cross Attention (PGCA) Block

    model/method

    The Projection-Guided Cross Attention (PGCA) block couples the gradient descent fidelity update with the inertia representation. Given the single-channel spatial representation r(k−1)∈RH×W×1r^{(k-1)} \in \mathbb{R}^{H \times W \times 1} and the inertia-updated feature Z^(k−1)∈RH×W×(C−1)\hat{Z}^{(k-1)} \in \mathbb{R}^{H \times W \times (C-1)}:

    1. Gradient Descent Step: The data-fidelity gradient step is applied directly to the image proxy:

    r^(k−1)=r(k−1)−ρ(k)Φ⊤(Φr(k−1)−y)\hat{r}^{(k-1)} = r^{(k-1)} - \rho^{(k)} \Phi^\top \left(\Phi r^{(k-1)} - y\right)

    where Φ∈RM×N\Phi \in \mathbb{R}^{M \times N} is the measurement matrix, y∈RMy \in \mathbb{R}^M is the compressed measurement, and ρ(k)\rho^{(k)} is a learnable step-size parameter.

    1. Cross-Attention Fusion: Layer normalization is applied to inputs, where r^(k−1)\hat{r}^{(k-1)} serves as query (projected to C−1C-1 channels inside GCAG_{\text{CA}}) and Z^(k−1)\hat{Z}^{(k-1)} serves as key and value:

    VPGCA(k)=LN(Z^(k−1)),KPGCA(k)=LN(Z^(k−1)),QPGCA(k)=LN(r^(k−1))V_{\text{PGCA}}^{(k)} = \text{LN}(\hat{Z}^{(k-1)}), \quad K_{\text{PGCA}}^{(k)} = \text{LN}(\hat{Z}^{(k-1)}), \quad Q_{\text{PGCA}}^{(k)} = \text{LN}(\hat{r}^{(k-1)})

    O(k)=GCA(VPGCA(k),KPGCA(k),QPGCA(k))+Z^(k−1)O^{(k)} = G_{\text{CA}}\left(V_{\text{PGCA}}^{(k)}, K_{\text{PGCA}}^{(k)}, Q_{\text{PGCA}}^{(k)}\right) + \hat{Z}^{(k-1)}

    1. Channel Reintegration: O(k)∈RH×W×(C−1)O^{(k)} \in \mathbb{R}^{H \times W \times (C-1)} and r^(k−1)∈RH×W×1\hat{r}^{(k-1)} \in \mathbb{R}^{H \times W \times 1} are concatenated along the channel dimension and fused via a 1×11 \times 1 convolution ConvO(⋅)\text{Conv}_O(\cdot) to yield the updated projection feature:

    S(k)=ConvO(Concat(O(k),r^(k−1)))∈RH×W×CS^{(k)} = \text{Conv}_O\left(\text{Concat}\left(O^{(k)}, \hat{r}^{(k-1)}\right)\right) \in \mathbb{R}^{H \times W \times C}

  6. Knowl 6 — Proximal Mapping Feed-Forward Network (FFN) Sub-Module

    model/method

    The Feed-Forward Network (FFN) sub-module HFFNH_{\text{FFN}} acts as the deep proximal mapping operator (denoiser) in feature space for solving:

    x(k)=arg⁡min⁡x12∥x−s(k)∥22+λR(x)x^{(k)} = \arg\min_x \frac{1}{2}\|x - s^{(k)}\|_2^2 + \lambda R(x)

    Given input S(k)∈RH×W×CS^{(k)} \in \mathbb{R}^{H \times W \times C}, HFFNH_{\text{FFN}} is implemented as a two-stage sequential block with a global skip connection:

    X(k)=HFFN(S(k))=S(k)+FFB2(LN(FFB1(LN(S(k)))))X^{(k)} = H_{\text{FFN}}(S^{(k)}) = S^{(k)} + \text{FFB}_2\left(\text{LN}\left(\text{FFB}_1\left(\text{LN}\left(S^{(k)}\right)\right)\right)\right)

    where LN(⋅)\text{LN}(\cdot) denotes Layer Normalization and each Feed-Forward Block (FFB) consists of the following sequence:

    1. A 1×11 \times 1 convolution expanding channels from CC to 4C4C,
    2. GELU non-linear activation,
    3. A 3×33 \times 3 depth-wise convolution on 4C4C channels,
    4. GELU non-linear activation,
    5. A 1×11 \times 1 convolution projecting channels from 4C4C back to CC.
  7. Knowl 7 — Training Loss and Optimization Details of OCTUF

    model/method

    OCTUF is trained end-to-end to reconstruct full-sampled images xj∈RNx_j \in \mathbb{R}^N from compressed measurements yj=Φxj∈RMy_j = \Phi x_j \in \mathbb{R}^M. The training minimizes the Mean Squared Error (MSE) loss function:

    L(Θ)=1NNa∑j=1Na∥xj−x^j∥22\mathcal{L}(\Theta) = \frac{1}{N N_a} \sum_{j=1}^{N_a} \|x_j - \hat{x}_j\|_2^2

    where NaN_a is the number of training image patches, NN is the pixel count per image patch (N=1024N=1024 for 32×3232 \times 32 image blocks), and x^j\hat{x}_j is the output of OCTUF given measurement yjy_j.

    The parameter set Θ={Φ,Conv0(⋅)}∪{HDual-CA(k)(⋅),HFFN(k)(⋅)}k=1K\Theta = \left\{\Phi, \text{Conv}_0(\cdot)\right\} \cup \left\{H_{\text{Dual-CA}}^{(k)}(\cdot), H_{\text{FFN}}^{(k)}(\cdot)\right\}_{k=1}^K is optimized using the Adam optimizer across 100 epochs with a batch size of 16. The learning rate starts at 5×10−45 \times 10^{-4} (for OCTUF with K=10K=10 stages) or 2×10−42 \times 10^{-4} (for OCTUF+ with K=16K=16 stages), includes 3 warm-up epochs, and decays to 5×10−55 \times 10^{-5} following a cosine annealing schedule. The learnable step size parameters ρ(k)\rho^{(k)} are initialized to 0.50.5, and feature channel depth is set to C=32C=32.

  8. Knowl 8 — Reconstruction Quality on Set11 and Urban100 Benchmarks

    data/table

    Average PSNR (dB) and SSIM reconstruction performance of OCTUF and OCTUF+ compared to representative CS methods on the Set11 and Urban100 datasets across CS sampling ratios of 10%, 25%, 30%, 40%, and 50%:

    Set11 Dataset
    Method 10% 25% 30% 40% 50% Average
    ISTA-Net+ (CVPR 2018) 26.58/0.8066 32.48/0.9242 33.81/0.9393 36.04/0.9581 38.06/0.9706 33.39/0.9197
    DPA-Net (TIP 2020) 27.66/0.8530 32.38/0.9311 33.35/0.9425 35.21/0.9580 36.80/0.9685 33.08/0.9306
    AMP-Net (TIP 2020) 29.40/0.8779 34.63/0.9481 36.03/0.9586 38.28/0.9715 40.34/0.9804 35.74/0.9473
    MAC-Net (ECCV 2020) 27.68/0.8182 32.91/0.9244 33.96/0.9372 35.94/0.9560 37.67/0.9668 33.63/0.9205
    COAST (TIP 2021) 28.74/0.8619 33.98/0.9407 35.11/0.9505 37.11/0.9646 38.94/0.9744 34.78/0.9384
    MADUN (ACM MM 2021) 29.91/0.8986 35.66/0.9601 36.94/0.9676 39.15/0.9772 40.77/0.9832 36.48/0.9573
    CASNet (TIP 2022) 30.36/0.9014 35.67/0.9591 36.92/0.9662 39.04/0.9760 40.93/0.9826 36.58/0.9571
    TransCS (TIP 2022) 29.54/0.8877 35.06/0.9548 35.62/0.9588 38.46/0.9737 40.49/0.9815 35.83/0.9513
    FSOINet (ICASSP 2022) 30.46/0.9023 35.80/0.9595 37.00/0.9665 39.14/0.9764 41.08/0.9832 36.70/0.9576
    MR-CCSNet (CVPR 2022) -/- 34.77/0.9546 -/- -/- 40.73/0.9828 -/-
    OCTUF (Ours) 30.70/0.9030 36.10/0.9604 37.21/0.9673 39.41/0.9773 41.34/0.9838 36.95/0.9584
    OCTUF+ (Ours) 30.73/0.9036 36.10/0.9607 37.32/0.9676 39.43/0.9774 41.35/0.9838 36.99/0.9586
    Urban100 Dataset
    Method 10% 25% 30% 40% 50% Average
    ISTA-Net+ (CVPR 2018) 23.61/0.7238 28.93/0.8840 30.21/0.9079 32.43/0.9377 34.43/0.9571 29.92/0.8821
    DPA-Net (TIP 2020) 24.55/0.7841 28.80/0.8944 29.47/0.9034 31.09/0.9311 32.08/0.9447 29.20/0.8915
    AMP-Net (TIP 2020) 26.04/0.8151 30.89/0.9202 32.19/0.9365 34.37/0.9578 36.33/0.9712 31.96/0.9202
    MAC-Net (ECCV 2020) 24.21/0.7445 28.79/0.8798 29.99/0.9017 31.94/0.9272 34.03/0.9513 29.79/0.8809
    COAST (TIP 2021) 25.94/0.8035 31.10/0.9168 32.23/0.9321 34.22/0.9530 35.99/0.9665 31.90/0.9144
    MADUN (ACM MM 2021) 27.13/0.8393 32.54/0.9347 33.77/0.9472 35.80/0.9633 37.75/0.9746 33.40/0.9318
    CASNet (TIP 2022) 27.46/0.8616 32.20/0.9396 33.37/0.9511 35.48/0.9669 37.45/0.9777 33.19/0.9394
    TransCS (TIP 2022) 26.72/0.8413 31.72/0.9330 31.95/0.9483 35.22/0.9648 37.20/0.9761 32.56/0.9327
    FSOINet (ICASSP 2022) 27.53/0.8627 32.62/0.9430 33.84/0.9540 35.93/0.9688 37.80/0.9777 33.54/0.9412
    OCTUF (Ours) 27.79/0.8621 32.99/0.9445 34.21/0.9555 36.25/0.9669 38.29/0.9797 33.91/0.9423
    OCTUF+ (Ours) 27.92/0.8652 33.08/0.9453 34.27/0.9559 36.31/0.9700 38.28/0.9795 33.97/0.9432

    OCTUF+ achieves the highest average PSNR on Set11 (36.99 dB) and Urban100 (33.97 dB), outperforming previous best-performing methods across all sampling ratios.

  9. Knowl 9 — Complexity and Computational Cost Comparison

    data/table

    Comparison of parameter capacity (in Millions), model memory footprint (in MB), and computational complexity (in Giga-FLOPs) for reconstructing a 256×256256 \times 256 image at a CS ratio of 10%10\%:

    Method MADUN CASNet FSOINet OCTUF OCTUF+
    Params (M) 3.14 16.90 0.64 0.40 0.58
    Size (MB) 11.9 66.3 7.8 5.2 7.5
    FLOPs (G) 419.2 13391.5 266.6 189.3 294.6

    OCTUF achieves higher reconstruction performance while using only 0.40M parameters and 189.3G FLOPs, representing a 37.5% reduction in parameters and a 29.0% reduction in FLOPs compared to FSOINet (0.64M parameters, 266.6G FLOPs), and an 87.3% parameter reduction compared to MADUN (3.14M parameters).

  10. Knowl 10 — Ablation Analysis on OCTUF Structural Components

    empirical result

    Component-wise ablation experiments on Set11 demonstrate the contribution of each module:

    1. Dual-CA and FFN Contribution (CS ratio = 50%, 10 iterations):

      • Baseline using ResBlocks with comparable parameter count: 38.25 dB PSNR / 0.9759 SSIM (0.72M parameters).
      • Adding FFN alone: 38.96 dB PSNR (+0.71 dB, 0.72M parameters).
      • Adding Dual-CA alone: 41.16 dB PSNR (+2.91 dB, 0.82M parameters).
      • Full OCTUF (Dual-CA + FFN + LayerNorm): 41.34 dB PSNR / 0.9838 SSIM (+3.09 dB over baseline, 0.82M parameters).
      • Removing all LayerNorm operations reduces performance to 41.21 dB PSNR.
    2. Dual-CA Sub-module Breakdown (CS ratio = 30%, Set11):

      • Proximal network only (FFN): 34.59 dB PSNR.
      • Unfolded gradient descent block in image domain (FFN + GDB): 35.93 dB PSNR (+1.34 dB).
      • Unfolded iteration in feature domain (FFN + GDB + Feature Domain): 36.82 dB PSNR (+0.89 dB).
      • Adding naive inertia term (direct subtraction α(x(k−1)−x(k−2))\alpha(x^{(k-1)}-x^{(k-2)})): 36.83 dB PSNR (+0.01 dB).
      • Using multi-channel ISCA without PGCA: 37.13 dB PSNR (+0.31 dB over image-level inertia).
      • Using PGCA without ISCA: 37.08 dB PSNR.
      • Full Dual-CA (ISCA + PGCA): 37.21 dB PSNR.

Coverage note — Qualitative visual comparison images (Figures 4 and 5), feature map attention heatmaps (Figure 6), and the Gaussian noise sensitivity curve (Figure 7) were omitted as their quantitative conclusions are fully represented in the benchmark tables and text.

References

  1. 1.Pablo Arbelaez, Michael Maire, Charless Fowlkes, and Jitendra Malik. Contour detection and hierarchical image segmentation. IEEE Transactions on Pattern Analysis and Machine Intelligence, 33(5):898–916, 2010. 5
  2. 2.Radu Ioan Bot¸, Erno Robert Csetnek, and Szilard Csaba Laszlo. An inertial forward–backward algorithm for the minimization of the sum of two nonconvex functions. EURO Journal on Computational Optimization, 4(1):3–25, 2016. 3
  3. 3.Yuanhao Cai, Jing Lin, Xiaowan Hu, Haoqian Wang, Xin Yuan, Yulun Zhang, Radu Timofte, and Luc Van Gool. Mask-guided spectral-wise transformer for efficient hyperspectral image reconstruction. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2022. 3
  4. 4.Yuanhao Cai, Jing Lin, Haoqian Wang, Xin Yuan, Henghui Ding, Yulun Zhang, Radu Timofte, and Luc Van Gool. Degradation-aware unfolding half-shuffle Transformer for spectral compressive imaging. In Proceedings of the International Conference on Neural Information Processing Systems (NeurIPS), 2022. 1, 3
  5. 5.Emmanuel J Candes, Justin Romberg, and Terence Tao. Robust uncertainty principles: Exact signal reconstruction from highly incomplete frequency information. IEEE Transactions on Information Theory, 52(2):489–509, 2006. 1
  6. 6.Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. End-to-end object detection with Transformers. In Proceedings of the European Conference on Computer Vision (ECCV), 2020. 2
  7. 7.Bin Chen and Jian Zhang. Content-aware scalable deep compressed sensing. IEEE Transactions on Image Processing, 31:5412–5426, 2022. 1, 2, 5, 6, 7, 8
  8. 8.Hanting Chen, Yunhe Wang, Tianyu Guo, Chang Xu, Yiping Deng, Zhenhua Liu, Siwei Ma, Chunjing Xu, Chao Xu, and Wen Gao. Pre-trained image processing transformer. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12299–12310, 2021. 2
  9. 9.Jiwei Chen, Yubao Sun, Qingshan Liu, and Rui Huang. Learning memory augmented cascading network for compressed sensing of images. In Proceedings of the European Conference on Computer Vision (ECCV), 2020. 2, 5, 7
  10. 10.Wenjun Chen, Chunling Yang, and Xin Yang. FSOINET: feature-space optimization-inspired network for image compressive sensing. In Proceedings of the International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2022. 1, 2, 5, 6, 7
  11. 11.Yunjin Chen and Thomas Pock. Trainable nonlinear reaction diffusion: A flexible framework for fast and effective image restoration. IEEE Transactions on Pattern Analysis and Machine Intelligence, 39(6):1256–1272, 2016. 2
  12. 12.Patrick L Combettes and Valerie R Wajs. Signal recovery by proximal forward-backward splitting. Multiscale Modeling & Simulation, 4(4):1168–1200, 2005. 3
  13. 13.Weisheng Dong, Peiyao Wang, Wotao Yin, Guangming Shi, Fangfang Wu, and Xiaotong Lu. Denoising prior driven deep neural network for image restoration. IEEE Transactions on Pattern Analysis and Machine Intelligence, 41(10):2305–2318, 2018. 5, 6, 7
  14. 14.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. In Proceedings of the International Conference on Learning Representations (ICLR), 2021. 2
  15. 15.Marco F Duarte, Mark A Davenport, Dharmpal Takhar, Jason N Laska, Ting Sun, Kevin F Kelly, and Richard G Baraniuk. Single-pixel imaging via compressive sampling. IEEE Signal Processing Magazine, 25(2):83–91, 2008. 1
  16. 16.Zi-En Fan, Feng Lian, and Jia-Ni Quan. Global sensing and measurements reuse for image compressed sensing. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2022. 1, 5, 7
  17. 17.Xinwei Gao, Jian Zhang, Wenbin Che, Xiaopeng Fan, and Debin Zhao. Block-based compressive sensing coding of natural images by local structural measurement matrix. In Proceedings of Data Compression Conference (DCC), 2015. 2
  18. 18.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016. 7
  19. 19.Sheena A Josselyn and Susumu Tonegawa. Memory engrams: Recalling the past and imagining the future. Science, 367(6473), 2020. 1, 6
  20. 20.Yookyung Kim, Mariappan S Nadar, and Ali Bilgin. Compressed sensing using a Gaussian scale mixtures model in wavelet domain. In Proceedings of the IEEE International Conference on Image Processing (ICIP), 2010. 2
  21. 21.Diederik P. Kingma and Jimmy Ba. Adam: A method for stochastic optimization. In Proceedings of the International Conference on Learning Representations (ICLR), 2015. 5
  22. 22.Filippos Kokkinos and Stamatios Lefkimmiatis. Deep image demosaicking using a cascade of convolutional residual denoising networks. In Proceedings of the European Conference on Computer Vision (ECCV), 2018. 2
  23. 23.Jakob Kruse, Carsten Rother, and Uwe Schmidt. Learning to push the limits of efficient FFT-based image deconvolution. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), 2017. 2
  24. 24.Kuldeep Kulkarni, Suhas Lohit, Pavan Turaga, Ronan Kerviche, and Amit Ashok. ReconNet: Non-iterative reconstruction of images from compressively sensed measurements. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016. 1, 2, 5, 6, 7, 8
  25. 25.Stamatios Lefkimmiatis. Non-local color image denoising with convolutional neural networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017. 2
  26. 26.Chengbo Li, Wotao Yin, Hong Jiang, and Yin Zhang. An efficient augmented lagrangian method with applications to total variation minimization. Computational Optimization and Applications, 56(3):507–530, 2013. 2
  27. 27.Jingyun Liang, Jiezhang Cao, Guolei Sun, Kai Zhang, Luc Van Gool, and Radu Timofte. SwinIR: Image restoration using Swin Transformer. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2021. 3
  28. 28.Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin Transformer: Hierarchical vision Transformer using shifted windows. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2021. 2, 3
  29. 29.Antoine Liutkus, David Martina, Sebastien Popoff, Gilles Chardon, Ori Katz, Geoffroy Lerosey, Sylvain Gigan, Laurent Daudet, and Igor Carron. Imaging with nature: Compressive imaging using a multiply scattering medium. Scientific Reports, 4:5552, 2014. 1
  30. 30.Ilya Loshchilov and Frank Hutter. SGDR: Stochastic gradient descent with warm restarts. In Proceedings of the International Conference on Learning Representations (ICLR), 2017. 5
  31. 31.Michael Lustig, David Donoho, and John M Pauly. Sparse MRI: The application of compressed sensing for rapid MR imaging. Magnetic Resonance in Medicine: An Official Journal of the International Society for Magnetic Resonance in Medicine, 58(6):1182–1195, 2007. 1
  32. 32.Christopher A Metzler, Arian Maleki, and Richard G Baraniuk. From denoising to compressed sensing. IEEE Transactions on Information Theory, 62(9):5117–5144, 2016. 2
  33. 33.Yurii Nesterov. Introductory lectures on convex optimization: A basic course, volume 87. Springer Science & Business Media, 2003. 3
  34. 34.Peter Ochs, Yunjin Chen, Thomas Brox, and Thomas Pock. iPiano: Inertial proximal algorithm for nonconvex optimization. SIAM Journal on Imaging Sciences, 7(2):1388–1419, 2014. 3
  35. 35.Olivier Petit, Nicolas Thome, Clement Rambour, Loic Themyr, Toby Collins, and Luc Soler. U-Net Transformer: Self and cross attention for medical image segmentation. In Proceedings of the International Workshop on Machine Learning in Medical Imaging, 2021. 2
  36. 36.Chao Ren, Xiaohai He, Chuncheng Wang, and Zhibo Zhao. Adaptive consistency prior based deep network for image denoising. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2021. 2
  37. 37.Florian Rousset, Nicolas Ducros, Andrea Farina, Gianluca Valentini, Cosimo D’Andrea, and Franc¸oise Peyrin. Adaptive basis scan by wavelet prediction for single-pixel imaging. IEEE Transactions on Computational Imaging, 3(1):36–46, 2016. 1
  38. 38.Aswin C Sankaranarayanan, Christoph Studer, and Richard G Baraniuk. CS-MUVI: Video compressive sensing for spatial-multiplexing cameras. In Proceedings of the IEEE International Conference on Computational Photography (ICCP), 2012. 1
  39. 39.Minghe Shen, Hongping Gan, Chao Ning, Yi Hua, and Tao Zhang. TransCS: A Transformer-based hybrid architecture for image compressed sensing. IEEE Transactions on Image Processing, 2022. 1, 2, 3, 5, 6, 7
  40. 40.Wuzhen Shi, Feng Jiang, Shaohui Liu, and Debin Zhao. Image compressed sensing using convolutional neural network. IEEE Transactions on Image Processing, 29:375–388, 2019. 5
  41. 41.Jiechong Song, Bin Chen, and Jian Zhang. Memory-augmented deep unfolding network for compressive sensing. In Proceedings of the ACM International Conference on Multimedia (ACM MM), 2021. 1, 2, 5, 6, 7
  42. 42.Jiechong Song, Bin Chen, and Jian Zhang. Deep memory-augmented proximal unrolling network for compressive sensing. International Journal of Computer Vision, pages 1–20, 2023. 2
  43. 43.Yueming Su and Qiusheng Lian. iPiano-Net: Nonconvex optimization inspired multi-scale reconstruction network for compressed sensing. Signal Processing: Image Communication, 89:115989, 2020. 2
  44. 44.Yubao Sun, Jiwei Chen, Qingshan Liu, Bo Liu, and Guodong Guo. Dual-path attention network for compressed sensing image reconstruction. IEEE Transactions on Image Processing, 29:9482–9495, 2020. 1, 2, 5, 6, 7
  45. 45.T. P. Szczykutowicz and G. Chen. Dual energy CT using slow kVp switching acquisition and prior image constrained compressed sensing. Physics in Medicine & Biology, 55(21):6411, 2010. 1
  46. 46.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Proceedings of the International Conference on Neural Information Processing Systems (NeurIPS), 2017. 2
  47. 47.Haixin Wang, Tianhao Zhang, Muzhi Yu, Jinan Sun, Wei Ye, Chen Wang, and Shikun Zhang. Stacking networks dynamically for image restoration based on the plug-and-play framework. In Proceedings of the European Conference on Computer Vision (ECCV), 2020. 2
  48. 48.Yuqing Wang, Zhaoliang Xu, Xinlong Wang, Chunhua Shen, Baoshan Cheng, Hao Shen, and Huaxia Xia. End-to-end video instance segmentation with Transformers. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2021. 2
  49. 49.Zhendong Wang, Xiaodong Cun, Jianmin Bao, Wengang Zhou, Jianzhuang Liu, and Houqiang Li. Uformer: A general U-shaped Transformer for image restoration. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2022. 3
  50. 50.Zhuoyuan Wu, Jian Zhang, and Chong Mou. Dense deep unfolding network with 3D-CNN prior for snapshot compressive sensing. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), 2021. 1
  51. 51.Zhuoyuan Wu, Zhenyu Zhang, Jiechong Song, and Jian Zhang. Spatial-temporal synergic prior driven unfolding network for snapshot compressive imaging. In Proceedings of IEEE International Conference on Multimedia and Expo (ICME), 2021. 1
  52. 52.Di You, Jingfen Xie, and Jian Zhang. ISTA-Net++: Flexible deep unfolding network for compressive sensing. In Proceedings of IEEE International Conference on Multimedia and Expo (ICME), 2021. 2
  53. 53.Di You, Jian Zhang, Jingfen Xie, Bin Chen, and Siwei Ma. COAST: Controllable arbitrary-sampling network for compressive sensing. IEEE Transactions on Image Processing, 30:6066–6080, 2021. 1, 2, 5, 6, 7
  54. 54.Jian Zhang and Bernard Ghanem. ISTA-Net: Interpretable optimization-inspired deep network for image compressive sensing. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018. 1, 2, 5, 6, 7
  55. 55.Jian Zhang, Chen Zhao, and Wen Gao. Optimization-inspired compact deep compressive sensing. IEEE Journal of Selected Topics in Signal Processing, 14(4):765–774, 2020. 2
  56. 56.Jian Zhang, Chen Zhao, Debin Zhao, and Wen Gao. Image compressive sensing recovery using adaptively learned sparsifying basis via L0 minimization. Signal Processing, 103:114–126, 2014. 2
  57. 57.Jian Zhang, Debin Zhao, and Wen Gao. Group-based sparse representation for image restoration. IEEE Transactions on Image Processing, 23(8):3336–3351, 2014. 2
  58. 58.Kai Zhang, Luc Van Gool, and Radu Timofte. Deep unfolding network for image super-resolution. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2020. 2
  59. 59.Zhilin Zhang, Tzyy-Ping Jung, Scott Makeig, and Bhaskar D Rao. Compressed sensing for energy-efficient wireless telemonitoring of noninvasive fetal ECG via block sparse bayesian learning. IEEE Transactions on Biomedical Engineering, 60(2):300–309, 2012. 1
  60. 60.Zhonghao Zhang, Yipeng Liu, Jiani Liu, Fei Wen, and Ce Zhu. AMP-Net: Denoising-based deep unfolding for compressive image sensing. IEEE Transactions on Image Processing, 30:1487–1500, 2020. 1, 2, 5, 6, 7
  61. 61.Chen Zhao, Siwei Ma, and Wen Gao. Image compressive-sensing recovery using structured laplacian sparsity in DCT domain and multi-hypothesis prediction. In Proceedings of IEEE International Conference on Multimedia and Expo (ICME), 2014. 2
  62. 62.Chen Zhao, Siwei Ma, Jian Zhang, Ruiqin Xiong, and Wen Gao. Video compressive sensing reconstruction via reweighted residual sparsity. IEEE Transactions on Circuits and Systems for Video Technology, 27(6):1182–1195, 2016. 2
  63. 63.Chen Zhao, Jian Zhang, Siwei Ma, and Wen Gao. Non-convex Lp nuclear norm based ADMM framework for compressed sensing. In Proceedings of Data Compression Conference (DCC), 2016. 2
  64. 64.Chen Zhao, Jian Zhang, Ronggang Wang, and Wen Gao. CREAM: CNN-REgularized ADMM framework for compressive-sensed image reconstruction. IEEE Access, 6:76838–76853, 2018. 2
  65. 65.Lichen Zhao, Daigang Cai, Lu Sheng, and Dong Xu. 3DVG-Transformer: Relation modeling for visual grounding on point clouds. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), 2021. 4
  66. 66.Xizhou Zhu, Weijie Su, Lewei Lu, Bin Li, Xiaogang Wang, and Jifeng Dai. Deformable DETR: Deformable transformers for end-to-end object detection. In Proceedings of the International Conference on Learning Representations (ICLR), 2021. 2

Citation

MLA
Song, J., et al. “Optimization-Inspired Cross-Attention Transformer for Compressive Sensing”. arXiv, 2023, http://arxiv.org/abs/2304.13986v1.
APA
Song, J., Mou, C., Wang, S., Ma, S., & Zhang, J. (2023). Optimization-Inspired Cross-Attention Transformer for Compressive Sensing. arXiv. http://arxiv.org/abs/2304.13986v1
Chicago
Song, J., C. Mou, S. Wang, S. Ma, and J. Zhang. 2023. “Optimization-Inspired Cross-Attention Transformer for Compressive Sensing”. arXiv. http://arxiv.org/abs/2304.13986v1.
Harvard
Song, J. et al. (2023) “Optimization-Inspired Cross-Attention Transformer for Compressive Sensing”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2304.13986v1.
Vancouver
1. Song J, Mou C, Wang S, Ma S, Zhang J (2023) Optimization-Inspired Cross-Attention Transformer for Compressive Sensing. arXiv

BibTeX

@article{song2023optimization,
  title = {Optimization-Inspired Cross-Attention Transformer for Compressive Sensing},
  author = {Song, Jiechong and Mou, Chong and Wang, Shiqi and Ma, Siwei and Zhang, Jian},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2304.13986v1},
  eprint = {2304.13986}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE