MambaRaw: Selective State Space Modeling for Efficient 4K Raw Image Reconstruction

Peize LiFanhu ZengTongda XuXingguo XuXinjie ZhangXingtong GeHaotian ZhangYan Wang

article2026arXiv1 citations

Presents MambaRaw, a selective state space framework that replaces quadratic attention with tiled scanning to reconstruct 4K raw images from embedded JPEG previews with improved fidelity and reduced coding latency.

Listen

Capturing uncompressed, high-resolution raw images is essential for advanced computational photography, but storing and transmitting raw data demands significant storage capacity and network bandwidth. While standard in-camera JPEG preview images provide an accessible, low-cost visual reference, existing metadata-based frameworks that reconstruct raw signals from these previews struggle with computational bottlenecks. Conventional context models rely on convolutional neural networks that lack broad contextual awareness or attention mechanisms whose computational demands grow prohibitively at high resolutions such as 4K. The article addresses the need for a practical, low-latency framework capable of high-fidelity raw image reconstruction.

The main objective of the article is to demonstrate MambaRaw, a JPEG-conditioned metadata reconstruction framework that integrates linear-time state space models into entropy parameter estimation. The framework is designed to achieve superior reconstruction fidelity and lower data transmission rates while significantly reducing computational overhead and processing latency.

To achieve this, the authors designed a spatial-energy coupled context modeling mechanism comprised of two lightweight components: a tile-based selective scanning module that applies advanced modeling only to high-information image regions, and an energy-aware refinement module that calibrates features according to the energy distribution of raw signals. The approach was evaluated through extensive benchmark experiments across three distinct camera sensor datasets (Samsung, Olympus, and Sony) from the NUS benchmark, as well as the AdobeFiveK photographic dataset, comparing rate-distortion performance, latency, computational operations, and peak memory usage against leading metadata-based methods.

The experimental findings show substantial improvements in reconstruction fidelity, efficiency, and hardware requirements. First, the proposed framework improves raw image reconstruction quality by 1.2 to 1.4 decibels in peak signal-to-noise ratio across camera subsets while requiring equal or lower metadata bitrates compared to top baselines. Second, under true 4K resolution testing, the method reduces total computational floating-point operations by approximately 56% and lowers end-to-end coding latency by about 9%. Third, at full 4K resolution, peak runtime memory consumption decreases from 22.8 gigabytes in standard baselines to 10.2 gigabytes, safely fitting within standard consumer-grade hardware limits. Finally, on low-bitrate benchmarks, the framework requires approximately 16% fewer metadata bits while maintaining higher fidelity than previous state-of-the-art models.

These findings indicate that high-resolution raw reconstruction can be deployed efficiently without enterprise-level server hardware or severe latency penalties. By demonstrating that selective state space modeling captures global image context at linear computational complexity, the article shows that devices can preserve professional-grade raw image fidelity over constrained networks. This shift reduces bandwidth costs and operational memory risks while maintaining visual precision in textured image regions.

For practical adoption, organizations developing computational photography, mobile imaging pipelines, or cloud photo storage systems should consider implementing selective state space entropy modeling to replace heavy attention-based or purely convolutional context models. Stakeholders should pursue two follow-up initiatives: expanding the framework to multi-frame raw video processing to capitalize on temporal redundancies, and developing hardware-aware optimizations to deploy selective state space processing directly onto mobile hardware accelerators for real-time mobile photography.

The conclusions are supported by consistent multi-dataset and multi-camera evaluations across various rate-distortion settings. However, decision-makers should note that primary benchmark evaluations followed standard research protocols on downscaled imagery before full-resolution validation, and the current implementation is restricted to static single-frame captures rather than video sequences. Confidence in the reported fidelity gains and memory reductions remains high for single-frame photographic applications.

No sufficiently relevant recommendations were found.

Cover for MambaRaw: Selective State Space Modeling for Efficient 4K Raw Image Reconstruction

Abstract

In-camera JPEG previews are ubiquitous in raw image formats and provide an sRGB reference at negligible storage cost. Although existing metadata-based reconstruction frameworks can exploit this side information when recovering raw images, their context models often become computationally expensive especially at high resolution, eg, 4K raw image, given that attention mechanisms scale quadratically with feature maps, hindering its practical application. To address these limitations, we propose MambaRaw, a JPEG-conditioned metadata-based raw image reconstruction framework that uses State Space Models (SSMs) to estimate entropy parameters efficiently. Our key contribution comprises a Spatial-Energy Coupled Context Modeling mechanism with two lightweight modules: (1) TileMambaBlock, which performs Mamba-style selective scanning only on information-dense tiles to improve the efficiency; and (2) Energy-Aware Refinement (EAR), an identity-initialized residual module that enhance feature representation to match the long-tail energy distribution of raw signals. Extensive experiments on three camera datasets (Sony, Olympus, Samsung) show consistent improvements over strong metadata-based baselines and set a new state of the art for JPEG-guided raw reconstruction with great efficiency. Notably, at low metadata bitrates, MambaRaw increases PSNR by 1.2--1.4 dB and reduces end-to-end coding latency by about 9%. Code is released at this https URL.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Learned Image Compression
  • 2.2 Metadata-based RAW Reconstruction
  • 2.3 State Space Models for Vision
  • 3 Method
  • 3.1 Problem Setup
  • 3.2 Motivation and Overview
  • 3.3 JPEG-Conditioned Reconstruction Backbone
  • 3.4 Efficiency: TileMambaBlock with Energy-Guided Selection
  • 3.5 Effectiveness: Energy-Aware Refinement (EAR)
  • 3.6 Training Strategy
  • 4 Experiments
  • 4.1 Experimental Setup
  • 4.2 Main Results
  • 4.3 Ablation Study
  • 4.4 Qualitative Visualization
  • 5 Conclusion
  • References
  • 0.A More Details
  • 0.A.1 Network Architecture
  • 0.A.2 Training Strategy
  • 0.A.3 Efficiency Measurement Details
  • 0.A.4 Evaluation Metrics
  • 0.A.5 Dataset Protocol Details (NUS)
  • 0.B More Experimental Results
  • 0.B.1 Hyperparameter Sensitivity
  • 0.B.2 Detailed Ablations Across Camera Subsets
  • 0.C More Visualization Results

Knowls

  1. Knowl 1 — MambaRaw Architecture for JPEG-Conditioned Raw Image Reconstruction

    model/method

    MambaRaw is a neural reconstruction framework that recovers a full-resolution scene-referred raw image xraw∈[0,1]3×H×W\mathbf{x}_{\text{raw}} \in [0, 1]^{3 \times H \times W} from an aligned in-camera preview xjpg∈[0,1]3×H×W\mathbf{x}_{\text{jpg}} \in [0, 1]^{3 \times H \times W} and a compact transmitted metadata bitstream s\mathbf{s}.

    The architecture builds upon a two-level variational autoencoder (VAE) hyperprior backbone with learned context modeling. It consists of:

    1. A JPEG-conditioned analysis transform gag_a mapping (xraw,xjpg)(\mathbf{x}_{\text{raw}}, \mathbf{x}_{\text{jpg}}) to latent representation y\mathbf{y}.
    2. Hyper-analysis hah_a and hyper-synthesis hsh_s transforms that extract and decode hyperlatents z\mathbf{z} into side information U=hs(z^)\mathbf{U} = h_s(\hat{\mathbf{z}}).
    3. A JPEG-conditioned synthesis transform gsg_s reconstructing x^\hat{\mathbf{x}} from quantized latents y^\hat{\mathbf{y}}.

    To incorporate JPEG guidance without quadratic attention overhead, preview features are bilinearly interpolated to the spatial resolution of layer ll, denoted xjpg(l)=ϕl(xjpg(l−1))\mathbf{x}_{\text{jpg}}^{(l)} = \phi_l(\mathbf{x}_{\text{jpg}}^{(l-1)}), and concatenated along the channel dimension:

    F~(l)=[M(l−1)(F~(l−1)),xjpg(l)]\tilde{\mathbf{F}}^{(l)} = \left[\mathcal{M}^{(l-1)}(\tilde{\mathbf{F}}^{(l-1)}), \mathbf{x}_{\text{jpg}}^{(l)}\right]

    where M(l−1)\mathcal{M}^{(l-1)} is the layer mapping function and [⋅,⋅][\cdot, \cdot] denotes concatenation.

    Inside the Level-1 entropy parameter network, intermediate features F~\tilde{\mathbf{F}} pass through an input projection ψepin\psi_{\text{ep}}^{\text{in}}, a TileMambaBlock module for content-adaptive long-range spatial context modeling, and an Energy-Aware Refinement (EAR) module for channel-wise energy calibration, before producing independent Gaussian parameters (μ,log⁡σ)(\boldsymbol{\mu}, \log \boldsymbol{\sigma}) via an output head.

  2. Knowl 2 — Mixed-Scale Context Modeling Inference Algorithm

    algorithm

    The mixed-scale inference pipeline executes spatial context modeling by dynamically selecting high-energy tiles for 2D State Space Model (SSM) processing and applying Energy-Aware Refinement (EAR).

    Input: JPEG-conditioned feature F~∈RC0×H×W\tilde{\mathbf{F}} \in \mathbb{R}^{C_0 \times H \times W}, tile dimension TT, keep ratio ρ∈(0,1]\rho \in (0, 1]
    Output: Refined context feature tensor F′∈RC×H×W\mathbf{F}' \in \mathbb{R}^{C \times H \times W}
    Fin←ψepin(F~)\mathbf{F}_{\text{in}} \leftarrow \psi_{\text{ep}}^{\text{in}}(\tilde{\mathbf{F}})
    Nh←⌈H/T⌉,Nw←⌈W/T⌉,Nt←Nh×NwN_h \leftarrow \lceil H / T \rceil, \quad N_w \leftarrow \lceil W / T \rceil, \quad N_t \leftarrow N_h \times N_w
    if (H≤T and W≤T) or (ρ≥1.0)(H \le T \text{ and } W \le T) \text{ or } (\rho \ge 1.0) then
        Fc←MambaBlock(Fin)\mathbf{F}_c \leftarrow \text{MambaBlock}(\mathbf{F}_{\text{in}})
    else
        Pad Fin\mathbf{F}_{\text{in}} to multiples of TT along spatial dimensions to obtain Fin,pad\mathbf{F}_{\text{in,pad}}
        Reshape Fin,pad\mathbf{F}_{\text{in,pad}} into non-overlapping tiles {ti}i=1Nt\{\mathbf{t}_i\}_{i=1}^{N_t} where ti∈RC×T×T\mathbf{t}_i \in \mathbb{R}^{C \times T \times T}
        for i←1i \leftarrow 1 to NtN_t do
            Si←1CT2∑c=1C∑h=1T∑w=1Tti[c,h,w]2S_i \leftarrow \frac{1}{C T^2} \sum_{c=1}^C \sum_{h=1}^T \sum_{w=1}^T \mathbf{t}_i[c, h, w]^2
        end for
        k←max⁡(1,⌊ρNt⌋)k \leftarrow \max(1, \lfloor \rho N_t \rfloor)
        Itop←TopKIndices({Si}i=1Nt,k)\mathcal{I}_{\text{top}} \leftarrow \text{TopKIndices}(\{S_i\}_{i=1}^{N_t}, k)
        for i←1i \leftarrow 1 to NtN_t do
            if i∈Itopi \in \mathcal{I}_{\text{top}} then
                ti′←MambaBlock(ti)\mathbf{t}'_i \leftarrow \text{MambaBlock}(\mathbf{t}_i)
            else
                ti′←ti\mathbf{t}'_i \leftarrow \mathbf{t}_i
            end if
        end for
        Assemble {ti′}i=1Nt\{\mathbf{t}'_i\}_{i=1}^{N_t} and crop padding to obtain Fc∈RC×H×W\mathbf{F}_c \in \mathbb{R}^{C \times H \times W}
    end if
    e←1C∑j=1C(Fc,j)2∈R1×H×W\mathbf{e} \leftarrow \frac{1}{C} \sum_{j=1}^C (\mathbf{F}_{c, j})^2 \in \mathbb{R}^{1 \times H \times W}
    g←σ(Conv1×1(e))∈RC×H×W\mathbf{g} \leftarrow \sigma(\text{Conv}_{1\times 1}(\mathbf{e})) \in \mathbb{R}^{C \times H \times W}
    ΔF←Conv1×1(ReLU(Conv1×1(Fc)))\Delta \mathbf{F} \leftarrow \text{Conv}_{1\times 1}(\text{ReLU}(\text{Conv}_{1\times 1}(\mathbf{F}_c)))
    F′←Fc+g⊙ΔF\mathbf{F}' \leftarrow \mathbf{F}_c + \mathbf{g} \odot \Delta \mathbf{F}
    return F′\mathbf{F}'

    Default hyperparameters: tile size T=64T = 64, keep ratio ρ=0.5\rho = 0.5, internal channel width C=192C = 192.

  3. Knowl 3 — TileMambaBlock with Energy-Guided Selective Scanning

    model/method

    TileMambaBlock reduces computation in high-resolution context modeling by restricting State Space Model (SSM) selective scanning to high-entropy foreground tiles while skipping uniform background tiles.

    Given the context input feature Fin∈RC×H×W\mathbf{F}_{\text{in}} \in \mathbb{R}^{C \times H \times W}, the feature map is partitioned into non-overlapping tiles of size T×TT \times T, resulting in a set of Nt=⌈H/T⌉×⌈W/T⌉N_t = \lceil H/T \rceil \times \lceil W/T \rceil tiles T={t1,…,tNt}\mathcal{T} = \{\mathbf{t}_1, \dots, \mathbf{t}_{N_t}\}, where ti∈RC×T×T\mathbf{t}_i \in \mathbb{R}^{C \times T \times T}.

    The information density of each tile is evaluated using its spatial L2L_2 energy score:

    Si=1CT2∑c=1C∑h=1T∑w=1Tti[c,h,w]2S_i = \frac{1}{C T^2} \sum_{c=1}^C \sum_{h=1}^T \sum_{w=1}^T \mathbf{t}_i[c, h, w]^2

    Tiles are sorted by SiS_i, and the top-kk subset S\mathcal{S} is selected with k=⌊ρNt⌋k = \lfloor \rho N_t \rfloor, where ρ∈(0,1]\rho \in (0, 1] is the keep ratio (default ρ=0.5\rho = 0.5). The block processes tiles as follows:

    ti′={MambaBlock(ti),i∈S,ti,otherwise\mathbf{t}'_i = \begin{cases} \text{MambaBlock}(\mathbf{t}_i), & i \in \mathcal{S}, \\ \mathbf{t}_i, & \text{otherwise} \end{cases}

    The internal MambaBlock\text{MambaBlock} applies the 2D Visual State Space (VSS) operator with a state expansion factor of 2. It unrolls the 2D tile along four scanning directions (left-to-right, right-to-left, top-to-bottom, bottom-to-top), processes each sequence with a discrete linear state space model:

    ht+1=Aht+Bxt,yt=Chth_{t+1} = \mathbf{A}h_t + \mathbf{B}x_t, \quad y_t = \mathbf{C}h_t

    and merges the four directional scans via element-wise summation.

  4. Knowl 4 — Energy-Aware Refinement (EAR) Module

    model/method

    Energy-Aware Refinement (EAR) is a spatial-energy refurbishment module that modulates feature representations to match the long-tail energy distribution of raw sensor signals while preserving local spatial granularity.

    Given an intermediate context feature map Fc∈RC×H×W\mathbf{F}_c \in \mathbb{R}^{C \times H \times W}, EAR computes the per-pixel energy map across all CC channels:

    e=1C∑j=1C(Fc,j)2∈R1×H×W\mathbf{e} = \frac{1}{C} \sum_{j=1}^C (\mathbf{F}_{c, j})^2 \in \mathbb{R}^{1 \times H \times W}

    where Fc,j\mathbf{F}_{c, j} is the jj-th channel slice of Fc\mathbf{F}_c.

    The spatial energy map e\mathbf{e} is mapped to a channel-wise gating tensor g\mathbf{g}, and a non-linear residual perturbation ΔF\Delta \mathbf{F} is computed from Fc\mathbf{F}_c:

    g=σ(Conv1×1(e))∈RC×H×W\mathbf{g} = \sigma\left(\text{Conv}_{1\times 1}(\mathbf{e})\right) \in \mathbb{R}^{C \times H \times W}

    ΔF=Conv1×1(ReLU(Conv1×1(Fc)))∈RC×H×W\Delta \mathbf{F} = \text{Conv}_{1\times 1}\left(\text{ReLU}\left(\text{Conv}_{1\times 1}(\mathbf{F}_c)\right)\right) \in \mathbb{R}^{C \times H \times W}

    F′=Fc+g⊙ΔF\mathbf{F}' = \mathbf{F}_c + \mathbf{g} \odot \Delta \mathbf{F}

    where σ(⋅)\sigma(\cdot) denotes the sigmoid activation function and ⊙\odot represents element-wise multiplication.

    To ensure numerical stability during early training when the entropy model is uncalibrated, the weights and biases of the final 1×11 \times 1 convolution generating ΔF\Delta \mathbf{F} are initialized to zero, making EAR an exact identity mapping F′=Fc\mathbf{F}' = \mathbf{F}_c at initialization.

  5. Knowl 5 — Rate-Distortion Optimization and Entropy Parameter Head

    equation

    The end-to-end training of MambaRaw minimizes a rate-distortion loss across multiple operating points parameterized by λ\lambda:

    L=R(y^)+R(z^)+λ⋅D(xraw,x^)\mathcal{L} = R(\hat{\mathbf{y}}) + R(\hat{\mathbf{z}}) + \lambda \cdot D(\mathbf{x}_{\text{raw}}, \hat{\mathbf{x}})

    where R(y^)R(\hat{\mathbf{y}}) is the estimated metadata bitrate for quantized latents y^\hat{\mathbf{y}}, R(z^)R(\hat{\mathbf{z}}) is the bitrate for quantized hyperlatents z^\hat{\mathbf{z}}, D(xraw,x^)D(\mathbf{x}_{\text{raw}}, \hat{\mathbf{x}}) measures the distortion in raw-linear color space (normalized to [0,1][0, 1]), and λ∈{0.02,0.24,0.8,1.5,2.0,5.0,10.0,20.0}\lambda \in \{0.02, 0.24, 0.8, 1.5, 2.0, 5.0, 10.0, 20.0\} controls the rate-distortion trade-off.

    The distribution of the quantized latents y^\hat{\mathbf{y}} is modeled as conditionally independent Gaussian distributions whose mean μ\boldsymbol{\mu} and scale σ\boldsymbol{\sigma} are predicted by an entropy head ψepout\psi_{\text{ep}}^{\text{out}} from the refined context feature F′\mathbf{F}' and the hyperprior side information U=hs(z^)\mathbf{U} = h_s(\hat{\mathbf{z}}):

    (μ,log⁡σ)=ψepout(F′,U)(\boldsymbol{\mu}, \log \boldsymbol{\sigma}) = \psi_{\text{ep}}^{\text{out}}(\mathbf{F}', \mathbf{U})

  6. Knowl 6 — Full 4K Ultra-High-Definition Raw Reconstruction Benchmark

    data/table

    Evaluation on full native 4K resolution inputs (3840×21603840 \times 2160, λ=0.8\lambda = 0.8) comparing MambaRaw with the Beyond-R2LCM baseline shows substantial reductions in computational complexity and memory usage alongside improved reconstruction quality.

    Method FLOPs (G) ↓\downarrow Mem (GB) ↓\downarrow Time (ms) ↓\downarrow PSNR (dB) ↑\uparrow SSIM ↑\uparrow
    Beyond-R2LCM 5420.7 22.8 3125 55.21 0.9982
    MambaRaw (Ours) 2380.5 10.2 2859 56.58 0.9991

    MambaRaw reduces floating-point operations (FLOPs) by 56.1%, decreases peak memory consumption by 55.3% (bringing 4K raw reconstruction into a 10.2 GB consumer GPU memory footprint compared to 22.8 GB), lowers end-to-end wall-clock latency by 8.5%, and achieves a +1.37 dB gain in PSNR.

  7. Knowl 7 — Rate-Distortion Performance on NUS Dataset Across Camera Sensors

    data/table

    Quantitative rate-distortion evaluation on 4×4\times downsampled raw images from the NUS dataset across three sensor subsets (Samsung NX2000, Olympus E-PL6, and Sony SLT-A57). Bitrate (bpp) measures metadata only, excluding the shared base JPEG preview.

    Method Input bpp ↓\downarrow Samsung NX2000 Olympus E-PL6 Sony SLT-A57
    Space PSNR ↑\uparrow SSIM ↑\uparrow PSNR ↑\uparrow SSIM ↑\uparrow PSNR ↑\uparrow SSIM ↑\uparrow
    SAM 16-bit 0.7500 47.03 0.9962 49.35 0.9978 50.44 0.9982
    CAM 8-bit 0.8438 46.24 0.9964 46.84 0.9966 47.66 0.9974
    CAM 16-bit 0.8438 48.08 0.9968 50.71 0.9975 50.49 0.9973
    CAM* 8-bit 0.8438 49.40 0.9972 50.37 0.9976 51.63 0.9983
    CAM* 16-bit 0.8438 49.57 0.9975 51.54 0.9980 53.11 0.9985
    R2LCM 8-bit 1.2250 50.78 0.9986 53.56 0.9993 53.87 0.9991
    Beyond-R2LCM 8-bit 0.3763 56.74 0.9996 59.04 0.9997 58.21 0.9996
    MambaRaw (Ours) 8-bit 0.3612 57.91 0.9997 60.45 0.9998 59.58 0.9997
    • denotes CAM evaluated with online test-time fine-tuning. MambaRaw achieves a 1.17 dB improvement on Samsung NX2000, 1.41 dB on Olympus E-PL6, and 1.37 dB on Sony SLT-A57 compared to Beyond-R2LCM while requiring 4.0% lower metadata bitrate (0.3612 vs. 0.3763 bpp).
  8. Knowl 8 — Reconstruction Fidelity on AdobeFiveK Software ISP Benchmark

    data/table

    Evaluation on the AdobeFiveK benchmark under the Software ISP setup (4,500 training / 500 testing pairs) at original resolution. Metadata bitrates are reported in bits per pixel (bpp).

    Method bpp ↓\downarrow PSNR (dB) ↑\uparrow SSIM ↑\uparrow
    InvISP N/A 52.69 0.9994
    SAM 9.566e-4 49.61 0.9987
    SAM 9.521e-3 54.76 0.9995
    CAM 8.438e-1 56.72 0.9996
    R2LCM (w/o metadata) N/A 53.03 0.9993
    R2LCM 4.901e-4 58.14 0.9997
    Beyond-R2LCM 3.760e-4 58.44 0.9997
    MambaRaw (Ours) 3.150e-4 58.55 0.9997
    R2LCM 1.045e-2 59.02 0.9994
    Beyond-R2LCM 2.916e-3 59.09 0.9997
    MambaRaw (Ours) 2.450e-3 59.18 0.9998

    In the low-bitrate regime (<5×10−4< 5\times 10^{-4} bpp), MambaRaw improves PSNR by 0.11 dB over Beyond-R2LCM while using 16.2% fewer metadata bits (3.150×10−43.150\times 10^{-4} vs. 3.760×10−43.760\times 10^{-4} bpp).

  9. Knowl 9 — Ablation of Module Components and Context Modeling Blocks

    empirical result

    Progressive ablation on the Sony SLT-A57 subset (λ=0.8 \lambda = 0.8) isolates the structural contributions of EAR and TileMambaBlock context modeling against CNN and Transformer alternatives:

    1. Component Contributions:

      • Baseline (Beyond-R2LCM): 58.21 dB PSNR, 0.9996 SSIM, 563 ms latency.
      • Baseline + EAR: 58.55 dB PSNR (+0.34 dB), 0.9996 SSIM, 570 ms latency (+7 ms).
      • Baseline + EAR + Dense SSM: 59.61 dB PSNR (+1.06 dB over +EAR), 0.9997 SSIM, 584 ms latency.
      • MambaRaw (Baseline + EAR + TileMambaBlock with ρ=0.5\rho=0.5): 59.58 dB PSNR, 0.9997 SSIM, 515 ms latency. Tile-wise selective scanning reduces runtime by 11.8% (584 ms →\to 515 ms) while retaining 97% of the dense SSM's PSNR gain.
    2. Context Block Architecture Comparison:

      • ResBlock (CNN): 58.45 dB PSNR, 0.9996 SSIM, 480 ms latency.
      • Swin Transformer (Window Attention): 59.52 dB PSNR, 0.9997 SSIM, 620 ms latency.
      • SSM Block (MambaRaw): 59.58 dB PSNR, 0.9997 SSIM, 515 ms latency. The SSM block outperforms the Swin Transformer in PSNR (+0.06 dB) while running 16.9% faster (515 ms vs. 620 ms) due to linear complexity.
  10. Knowl 10 — Tile Selection Metrics and Hyperparameter Sensitivity Analysis

    empirical result

    Evaluations on the Sony SLT-A57 subset at λ=0.8\lambda = 0.8 demonstrate the optimal configuration for tile selection:

    1. Selection Metric Comparison:

      • Random Selection: 59.25 dB PSNR, 0.9996 SSIM, 515 ms latency.
      • Gradient Magnitude: 59.52 dB PSNR, 0.9997 SSIM, 528 ms latency.
      • Local Entropy: 59.55 dB PSNR, 0.9997 SSIM, 542 ms latency.
      • L2L_2 Energy (Ours): 59.58 dB PSNR, 0.9997 SSIM, 515 ms latency. L2L_2 energy yields the highest reconstruction quality while minimizing computation time (tile score calculation takes only 10 ms / 1.9% of runtime).
    2. Tile Size (TT) Sensitivity:

      • T=16T=16: 59.60 dB PSNR, 544 ms.
      • T=32T=32: 59.59 dB PSNR, 528 ms.
      • T=64T=64: 59.58 dB PSNR, 515 ms.
      • T=128T=128: 59.50 dB PSNR, 507 ms. T=64T=64 achieves near-optimal reconstruction quality while maintaining low overhead.
    3. Keep Ratio (ρ\rho) Sensitivity:

      • ρ=0.25\rho=0.25: 59.42 dB PSNR, 482 ms.
      • ρ=0.50\rho=0.50: 59.58 dB PSNR, 515 ms.
      • ρ=0.75\rho=0.75: 59.60 dB PSNR, 556 ms.
      • ρ=1.00\rho=1.00 (Dense): 59.61 dB PSNR, 584 ms. PSNR saturates beyond ρ=0.50\rho=0.50, making ρ=0.50\rho=0.50 the optimal trade-off between quality (+0.16 dB over ρ=0.25\rho=0.25) and speed (69 ms faster than dense scanning).
  11. Knowl 11 — Limitations and Scope of MambaRaw

    limitation

    MambaRaw has two primary limitations as identified by the authors:

    1. Single-Frame Constraint: The current architecture is designed solely for single-frame raw image reconstruction and does not exploit inter-frame temporal redundancies present in burst raw captures or raw video sequences.
    2. Hardware Deployment: The selective tile-scanning operations lack dedicated low-level hardware kernel mappings optimized for mobile image signal processor (ISP) neural accelerators, which are required for real-time edge processing in commercial camera pipelines.

Coverage note — None was omitted; all key architectural components, algorithms, mathematical formulations, benchmarks (NUS, AdobeFiveK, 4K), ablation studies, and limitations were fully extracted.

References

  1. 1.Ballé, J., Laparra, V., Simoncelli, E.P.: End-to-end optimized image compression. arXiv preprint arXiv:1611.01704 (2016) 3
  2. 2.Ballé, J., Minnen, D., Singh, S., Hwang, S.J., Johnston, N.: Variational image compression with a scale hyperprior. arXiv preprint arXiv:1802.01436 (2018) 2, 3
  3. 3.Bychkovsky, V., Paris, S., Chan, E., Durand, F.: Learning photographic global tonal adjustment with a database of input/output image pairs. In: CVPR 2011. pp. 97–104. IEEE (2011) 10
  4. 4.Chen, H., Han, W., Zheng, H., Shen, J.: Rawmamba: Unified srgb-to-raw de-rendering with state space model. arXiv preprint arXiv:2411.11717 (2024) 4
  5. 5.Chen, Y., Qin, H., Zhang, Z., Magno, M., Benini, L., Li, Y.: Q-mambair: Accurate quantized mamba for efficient image restoration. arXiv preprint arXiv:2503.21970 (2025) 4
  6. 6.Chen, Y., Lyu, Z., He, B., Hu, H., Wang, Q., Tian, Y., Song, L., Zhang, W., Lu, G.: Cmic: Content-adaptive mamba for learned image compression. arXiv preprint arXiv:2508.02192 (2025) 4
  7. 7.Cheng, D., Prasad, D.K., Brown, M.S.: Illuminant estimation for color constancy: why spatial-domain methods work and the role of the color distribution. Journal of the Optical Society of America A 31(5), 1049–1058 (2014) 10
  8. 8.Cheng, Z., Sun, H., Takeuchi, M., Katto, J.: Learned image compression with discretized gaussian mixture likelihoods and attention modules. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. pp. 7939–7948 (2020) 3
  9. 9.Dao, T., Gu, A.: Transformers are ssms: Generalized models and efficient algorithms through structured state space duality. arXiv preprint arXiv:2405.21060 (2024) 4
  10. 10.Gao, G., You, P., Pan, R., Han, S., Zhang, Y., Dai, Y., Lee, H.: Neural image compression via attentional multi-scale back projection and frequency decomposition. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. pp. 14677–14686 (2021) 3
  11. 11.Gu, A., Dao, T.: Mamba: Linear-time sequence modeling with selective state spaces. In: First conference on language modeling (2024) 4
  12. 12.Gu, A., Goel, K., Ré, C.: Efficiently modeling long sequences with structured state spaces. arXiv preprint arXiv:2111.00396 (2021) 4
  13. 13.Guo, H., Guo, Y., Zha, Y., Zhang, Y., Li, W., Dai, T., Xia, S.T., Li, Y.: Mambairv2: Attentive state space restoration. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 28124–28133 (2025) 4
  14. 14.Guo, Z., Zhang, Z., Feng, R., Chen, Z.: Causal contextual prediction for learned image compression. IEEE Transactions on Circuits and Systems for Video Technology 32(4), 2329–2341 (2021) 3
  15. 15.Hatamizadeh, A., Kautz, J.: Mambavision: A hybrid mamba-transformer vision backbone. In: Proceedings of the Computer Vision and Pattern Recognition Conference. pp. 25261–25270 (2025) 4
  16. 16.He, D., Yang, Z., Peng, W., Ma, R., Qin, H., Wang, Y.: Elic: Efficient learned image compression with unevenly grouped space-channel contextual adaptive coding. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. pp. 5718–5727 (2022) 3
  17. 17.He, D., Zheng, Y., Sun, B., Wang, Y., Qin, H.: Checkerboard context model for efficient learned image compression. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. pp. 14771–14780 (2021) 3
  18. 18.Hu, J., Shen, L., Sun, G.: Squeeze-and-excitation networks. In: Proceedings of the IEEE conference on computer vision and pattern recognition. pp. 7132–7141 (2018) 8
  19. 19.Huang, T., Pei, X., You, S., Wang, F., Qian, C., Xu, C.: Localmamba: Visual state space model with windowed selective scan. In: European conference on computer vision. pp. 12–22. Springer (2024) 4
  20. 20.Li, K., Li, X., Wang, Y., He, Y., Wang, Y., Wang, L., Qiao, Y.: Videomamba: State space model for efficient video understanding. In: European conference on computer vision. pp. 237–255. Springer (2024) 15
  21. 21.Li, M., Ma, K., You, J., Zhang, D., Zuo, W.: Efficient and effective context-based convolutional entropy modeling for image compression. IEEE Transactions on Image Processing 29, 5900–5911 (2020) 3
  22. 22.Liu, J., Sun, H., Katto, J.: Learned image compression with mixed transformer-cnn architectures. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. pp. 14388–14397 (2023) 3
  23. 23.Liu, Y., Tian, Y., Zhao, Y., Yu, H., Xie, L., Wang, Y., Ye, Q., Jiao, J., Liu, Y.: Vmamba: Visual state space model. Advances in neural information processing systems 37, 103031–103063 (2024) 4, 5, 20
  24. 24.Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., Guo, B.: Swin transformer: Hierarchical vision transformer using shifted windows. In: Proceedings of the IEEE/CVF international conference on computer vision. pp. 10012–10022 (2021) 14
  25. 25.Ma, C., Wang, Z., Liao, R., Ye, Y.: A cross channel context model for latents in deep image compression. arXiv preprint arXiv:2103.02884 (2021) 3
  26. 26.Ma, J., Li, F., Wang, B.: U-mamba: Enhancing long-range dependency for biomedical image segmentation. arXiv preprint arXiv:2401.04722 (2024) 4
  27. 27.Minnen, D., Ballé, J., Toderici, G.D.: Joint autoregressive and hierarchical priors for learned image compression. Advances in neural information processing systems 31 (2018) 2, 3
  28. 28.Minnen, D., Singh, S.: Channel-wise autoregressive entropy models for learned image compression. In: 2020 IEEE International Conference on Image Processing (ICIP). pp. 3339–3343. IEEE (2020) 3
  29. 29.Nam, S., Punnappurath, A., Brubaker, M.A., Brown, M.S.: Learning srgb-to-raw-rgb de-rendering with content-aware metadata. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. pp. 17704–17713 (2022) 4, 10, 11, 12, 14, 21, 24
  30. 30.Patel, Y., Appalaraju, S., Manmatha, R.: Saliency driven perceptual image compression. In: Proceedings of the IEEE/CVF winter conference on applications of computer vision. pp. 227–236 (2021) 3
  31. 31.Punnappurath, A., Brown, M.S.: Spatially aware metadata for raw reconstruction. In: Proceedings of the IEEE/CVF winter conference on applications of computer vision. pp. 218–226 (2021) 2, 4, 10, 12
  32. 32.Qin, S., Lu, Y., Zhou, Y., Li, J., Ren, Y., Xue, Y., Xia, S.T., Chen, B.: Freqsic: Frequency-aware stereo image compression with bi-directional checkerboard context model. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 19393–19402 (2026) 3
  33. 33.Qin, S., Wang, J., Zhou, Y., Chen, B., Luo, T., An, B., Dai, T., Xia, S.T., Wang, Y.: Cassic: Towards content-adaptive state-space models for learned image compression. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. pp. 15727–15736 (2025) 4
  34. 34.Qin, S., Zhang, X., Liu, Z., Wang, J., Chen, B., Li, J., Ren, Y., Xia, S.T., Zhang, J.: Mambasic: Mamba-based stereo image compression with bi-directional multi-reference entropy model. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 5306–5315 (2026) 4
  35. 35.Shi, Y., Xia, B., Jin, X., Wang, X., Zhao, T., Xia, X., Xiao, X., Yang, W.: Vmambair: Visual state space model for image restoration. IEEE Transactions on Circuits and Systems for Video Technology 35(6), 5560–5574 (2025) 4
  36. 36.Tian, Y., Ling, X., Geng, C., Hu, Q., Lu, G., Zha, G.: Smc++: Masked learning of unsupervised video semantic compression. IEEE Transactions on Pattern Analysis and Machine Intelligence (2025) 3
  37. 37.Tian, Y., Lu, G., Min, X., Che, Z., Zhai, G., Guo, G., Gao, Z.: Self-conditioned probabilistic learning of video rescaling. In: Proceedings of the IEEE/CVF international conference on computer vision. pp. 4490–4499 (2021) 3
  38. 38.Tian, Y., Lu, G., Yan, Y., Zhai, G., Chen, L., Gao, Z.: A coding framework and benchmark towards low-bitrate video understanding. IEEE Transactions on Pattern Analysis and Machine Intelligence 46(8), 5852–5872 (2024) 3
  39. 39.Tian, Y., Lu, G., Zhai, G.: Free-vsc: Free semantics from visual foundation models for unsupervised video semantic compression. In: European Conference on Computer Vision. pp. 163–183. Springer (2024) 3
  40. 40.Tian, Y., Lu, G., Zhai, G., Gao, Z.: Non-semantics suppressed mask learning for unsupervised video semantic compression. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. pp. 13610–13622 (2023) 3
  41. 41.Wallace, G.K.: The jpeg still picture compression standard. Communications of the ACM 34(4), 30–44 (1991) 2
  42. 42.Wang, Y., Yu, Y., Yang, W., Guo, L., Chau, L.P., Kot, A.C., Wen, B.: Raw image reconstruction with learned compact metadata. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 18206–18215 (2023) 2, 4, 6, 10, 12
  43. 43.Wang, Y., Yu, Y., Yang, W., Guo, L., Chau, L.P., Kot, A.C., Wen, B.: Beyond learned metadata-based raw image reconstruction. International Journal of Computer Vision 132(12), 5514–5533 (2024) 4, 6, 10, 12, 13, 14, 19, 20, 21, 24
  44. 44.Warenkorb, L.R.: Information technology-high efficiency coding and media delivery in heterogeneous environments-part 3: 3d audio (2015) 2
  45. 45.Wu, C., Wang, L., Zheng, Z., Cui, Y., Yang, Z., Chen, X., Zhang, Y., Jiang, W., Xia, J.: Scan clusters, not pixels: A cluster-centric paradigm for efficient ultra-high-definition image restoration. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 15528–15537 (2026) 3
  46. 46.Xing, Y., Qian, Z., Chen, Q.: Invertible image signal processing. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. pp. 6287–6296 (2021) 4, 12
  47. 47.Zeng, F., Tang, H., Shao, Y., Chen, S., Shao, L., Wang, Y.: Mambaic: State space models for high-performance learned image compression. In: Proceedings of the Computer Vision and Pattern Recognition Conference. pp. 18041–18050 (2025) 4, 20
  48. 48.Zhang, J., Nguyen, A.T., Han, X., Trinh, V.Q.H., Qin, H., Samaras, D., Hosseini, M.S.: 2dmamba: Efficient state space model for image representation with applications on giga-pixel whole slide image classification. In: Proceedings of the Computer Vision and Pattern Recognition Conference. pp. 3583–3592 (2025) 4
  49. 49.Zhou, Y., Zhou, P., Ng, T.K.: Efficient cascaded multiscale adaptive network for image restoration. In: European Conference on Computer Vision. pp. 92–110. Springer (2024) 3
  50. 50.Zhu, L., Liao, B., Zhang, Q., Wang, X., Liu, W., Wang, X.: Vision mamba: Efficient visual representation learning with bidirectional state space model. arXiv preprint arXiv:2401.09417 (2024) 4
  51. 51.Zhu, Y., Yang, Y., Cohen, T.: Transformer-based transform coding. In: International conference on learning representations (2022) 3
  52. 52.Zou, R., Song, C., Zhang, Z.: The devil is in the details: Window-based attention for image compression. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. pp. 17492–17501 (2022) 3

Citation

MLA
Li, P., et al. “MambaRaw: Selective State Space Modeling for Efficient 4K Raw Image Reconstruction”. arXiv, 2026, http://arxiv.org/abs/2606.24479v1.
APA
Li, P., Zeng, F., Xu, T., Xu, X., Zhang, X., Ge, X., Zhang, H., & Wang, Y. (2026). MambaRaw: Selective State Space Modeling for Efficient 4K Raw Image Reconstruction. arXiv. http://arxiv.org/abs/2606.24479v1
Chicago
Li, P., F. Zeng, T. Xu, et al. 2026. “MambaRaw: Selective State Space Modeling for Efficient 4K Raw Image Reconstruction”. arXiv. http://arxiv.org/abs/2606.24479v1.
Harvard
Li, P. et al. (2026) “MambaRaw: Selective State Space Modeling for Efficient 4K Raw Image Reconstruction”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2606.24479v1.
Vancouver
1. Li P, Zeng F, Xu T, Xu X, Zhang X, Ge X, Zhang H, Wang Y (2026) MambaRaw: Selective State Space Modeling for Efficient 4K Raw Image Reconstruction. arXiv

BibTeX

@article{li2026mambaraw,
  title = {MambaRaw: Selective State Space Modeling for Efficient 4K Raw Image Reconstruction},
  author = {Li, Peize and Zeng, Fanhu and Xu, Tongda and Xu, Xingguo and Zhang, Xinjie and Ge, Xingtong and Zhang, Haotian and Wang, Yan},
  year = {2026},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2606.24479v1},
  eprint = {2606.24479}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/