Optimization-Inspired Cross-Attention Transformer for Compressive Sensing
Jiechong SongChong MouShiqi WangSiwei MaJian Zhang
Proposes a lightweight deep unfolding framework that integrates cross-attention mechanisms directly into the iterative optimization steps of compressive sensing, preserving inter-stage feature information to achieve state-of-the-art image reconstruction with substantially fewer parameters.
Compressive sensing is a critical technology used across medical imaging, remote monitoring, and camera systems to capture and store signals efficiently using far fewer measurements than traditional methods. While modern deep unfolding networks have improved image reconstruction by merging classical optimization algorithms with deep neural networks, existing solutions face notable bottlenecks. They typically require massive numbers of parameters and heavy computational budgets, while suffering from feature information loss because they pass data between processing stages primarily at the image level rather than retaining rich internal representations.
To overcome these limitations, the article evaluates and demonstrates a lightweight reconstruction model called the Optimization-inspired Cross-attention Transformer Unfolding Framework. The main objective is to establish an interpretable, highly accurate deep unfolding framework that operates directly in feature space and significantly reduces memory and computational requirements.
The authors conducted extensive computational experiments to validate this approach. The system incorporates an iterative module consisting of a Dual Cross Attention sub-module—which includes an Inertia-Supplied Cross Attention block and a Projection-Guided Cross Attention block—paired with a Feed-Forward Network. The model was trained on 400 images from the standard BSD500 benchmark dataset and tested on two widely recognized evaluation datasets, Set11 and Urban100, across various compression ratios ranging from 10% to 50%.
The experimental findings show that the proposed framework consistently outperforms existing state-of-the-art models across standard image quality metrics. On the Set11 benchmark, the enhanced configuration of the framework achieved an average reconstruction quality improvement of 0.29 dB to 3.91 dB over nine competing modern methods. Visual evaluations revealed sharper edges and structural details compared to alternative approaches. Crucially, the model achieved these results while dramatically reducing operational overhead: at a 10% sampling ratio, the base framework required only 0.40 million parameters and 189.3 billion floating-point operations, compared to 16.90 million parameters and over 13,391 billion operations for leading alternatives. Additional robustness tests demonstrated that the architecture maintains stable reconstruction quality when subjected to varying levels of Gaussian noise.
These results demonstrate that high-performance image reconstruction does not require computationally heavy models. By successfully passing multi-channel feature information across iterations and integrating inertia forces into the optimization steps, the proposed framework lowers deployment costs, decreases hardware memory demands, and accelerates processing times without sacrificing image fidelity. This makes advanced compressive sensing significantly more practical for resource-constrained edge devices and real-time medical or remote imaging systems.
Stakeholders and engineering teams developing imaging pipelines should consider adopting cross-attention feature-space unfolding frameworks to optimize efficiency and reconstruction fidelity. Before deploying to production environments, organizations should conduct pilot testing on domain-specific hardware to evaluate real-time throughput. A key limitation noted is the reliance on synthetic image datasets and artificial Gaussian noise rather than open, real-world compressive sensing datasets. Future efforts should focus on validating performance on specialized real-world hardware, adapting the framework to broader image restoration inverse problems, and extending its capabilities to video processing applications.
- Paper: Algorithm Unrolling: Interpretable, Efficient Deep Learning for Signal and Image Processing, Vishal Monga et al. (2019). Provides the foundational principles of algorithm unrolling/unfolding, which the source relies on to construct its interpretable optimization-driven neural network architecture.
- Paper: Restormer: Efficient Transformer for High-Resolution Image Restoration, Syed Waqas Zamir et al. (2022). Introduces high-resolution feature restoration using efficient transformer blocks that directly inform the source's cross-attention mechanisms.
- Paper: Cross Aggregation Transformer for Image Restoration, Zheng Chen et al. (2022). Demonstrates cross-aggregation and rectangular window attention strategies for image restoration, establishing key concepts adapted in the source paper's dual cross-attention sub-modules.
- Paper: SwinIR: Image Restoration Using Swin Transformer, Jingyun Liang et al. (2021). Establishes a baseline for transformer-driven image restoration using shifted window attention and residual feature extraction.
- Paper: Learning Deep CNN Denoiser Prior for Image Restoration, Kai Zhang et al. (2017). Introduces the integration of deep denoiser priors within model-based optimization algorithms for general inverse imaging problems.
- Paper: Learning a Variational Network for Reconstruction of Accelerated MRI Data, Kerstin Hammernik et al. (2017). Provides a seminal example of embedding iterative physical sensing and compressed sensing optimization steps directly into trainable neural networks.
- Paper: Frequency-Aware Transformer for Learned Image Compression, Han Li et al. (2024). Extends efficient transformer architectures to learned image compression by introducing frequency-aware anisotropic attention mechanisms.
- Paper: TransNeXt: Robust Foveal Visual Perception for Vision Transformers, Dai Shi (2024). Builds on lightweight attention and multi-scale feature aggregation principles to design robust vision transformer backbones that mitigate depth degradation.
- Paper: C3: High-Performance and Low-Complexity Neural Compression from a Single Image or Video, Hyunjik Kim et al. (2024). Explores ultra-low complexity neural compression models tailored to individual media instances, extending the efficiency goals explored in the source.
- Paper: Fast Tensor Completion via Approximate Richardson Iteration, Mehrdad Ghadiri et al. (2025). Advances iterative optimization for high-dimensional recovery problems by developing fast approximate Richardson iterations for tensor completion.
