SNR-Aware Low-light Image Enhancement

Xiaogang XuRuixing WangChi-Wing FuJiaya Jia

article2022CVPR732 citations

Proposes an SNR-guided hybrid architecture that dynamically combines transformers for noisy, low-SNR regions and convolutional operations for higher-SNR areas, outperforming existing state-of-the-art methods across seven low-light benchmark datasets.

Listen

Low-light imaging presents significant challenges for automated systems and human operators across night-time security surveillance, autonomous navigation, and consumer photography. In dimly lit settings, images suffer from severe visibility loss and non-uniform noise distributions. Standard neural network approaches apply uniform enhancement operations across the entire scene, which frequently distorts colors, fails to resolve details in pitch-black areas, or amplifies background noise.

The article demonstrates a new framework designed to solve these issues through spatially varying enhancement guided by signal-to-noise ratio (SNR)—a standard metric measuring the ratio between meaningful image information and background noise. The authors evaluate this architecture across multiple public benchmarks and human perceptual tests to establish whether dynamically adjusting local and non-local operations according to regional noise levels outperforms current methods.

The framework pairs an SNR estimation method with a dual-branch neural architecture in its deepest layer. It utilizes standard convolutional blocks for short-range operations in relatively clear, high-signal regions and a specialized transformer module for long-range operations in heavily degraded, low-signal regions. Crucially, the system introduces a selective self-attention mechanism that masks out extremely noisy tokens, preventing corrupted areas from degrading long-range feature extraction. The outputs are then adaptively blended using the estimated SNR map. The framework was evaluated on seven benchmark datasets (such as LOL, SID, SMID, and SDSD) against more than a dozen state-of-the-art baselines, followed by a blind perceptual study involving 100 participants evaluating smartphone imagery.

The experimental findings show clear, consistent advantages across all benchmarks. First, the proposed framework achieved higher quantitative fidelity than all comparative methods, reaching 24.61 dB on the LOL-v1 dataset compared to 24.14 dB for the best previous convolutional method and 16.27–16.36 dB for standard visual transformers. Second, ablation studies showed that the complete system outperformed variants that removed either the long-range transformer branch or the short-range convolutional branch, confirming that combining both mechanisms is essential. Third, the user study confirmed statistically significant human preference for the framework over leading baselines in sharpness, noise suppression, color vividness, and natural realism. Finally, tests indicated the model remains robust regardless of the specific denoising algorithm used to estimate the initial noise map.

These results establish that low-light enhancement cannot rely on uniform processing or standard, unconstrained vision transformers, which inadvertently amplify severe noise across global contexts. By implementing spatially adaptive operations, organizations deploying night-time computer vision systems can achieve superior visual clarity and downstream object recognition without requiring specialized sensor hardware upgrades. The model significantly reduces visual artifacts, lowers failure risks in low-visibility autonomous tasks, and delivers higher operational reliability.

Stakeholders and development teams working on low-light enhancement pipelines should evaluate adopting SNR-guided feature extraction as an upgrade over uniform convolutional or standard transformer models. Before widespread operational deployment, engineering teams should conduct targeted pilot evaluations on specific field sensor data. The primary limitations of the current study are its focus on static RGB images and the reliance on approximation heuristics for noise estimation in nearly black regions. Future development should prioritize expanding the framework into low-light video streams by incorporating temporal tracking alongside spatial adaptation, as well as testing generative models to realistically synthesize visual information in completely black environments.

Cover for SNR-Aware Low-light Image Enhancement

Abstract

This paper presents a new solution for low-light image enhancement by collectively exploiting Signal-to-Noise Ratio-aware transformers and convolutional models to dynamically enhance pixels with spatial-varying operations. They are long-range operations for image regions of extremely low Signal-to-Noise-Ratio (SNR) and short-range operations for other regions. We propose to take an SNR prior to guide the feature fusion and formulate the SNR-aware transformer with a new self-attention model to avoid tokens from noisy image regions of very low SNR. Extensive experiments show that our framework consistently achieves better performance than SOTA approaches on seven representative benchmarks with the same structure. Also, we conducted a large-scale user study with 100 participants to verify the superior perceptual quality of our results. The code is available at https://github.com/dvlab-research/SNR-Aware-Low-Light-Enhance.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Our Method
  • 3.1. Long- and Short-range Branches
  • 3.2. SNR-based Spatially-varying Feature Fusion
  • 3.3. SNR-guided Attention in Transformer
  • 3.4. Loss Function
  • 4. Experiments
  • 4.1. Datasets and Implementation Details
  • 4.2. Comparison with Current Methods
  • 4.3. Ablation Study
  • 4.4. Influence of SNR Prior
  • 5. Conclusion
  • References

Knowls

  1. Knowl 1 — Dual-Branch Signal-to-Noise-Ratio-Aware Framework for Low-Light Image Enhancement

    model/method

    The SNR-aware low-light enhancement framework adaptively applies long-range operations to low Signal-to-Noise Ratio (SNR) regions and short-range operations to high-SNR regions.

    Given an input low-light image I∈RH×W×3I \in \mathbb{R}^{H \times W \times 3}, an encoder comprising three convolutional stages (with strides 1, 2, and 2 and LeakyReLU activations) followed by a residual block extracts a deep feature representation F∈Rh×w×CF \in \mathbb{R}^{h \times w \times C}, where h=H/4h = H/4 and w=W/4w = W/4.

    In the deepest layer, the feature FF is processed by two parallel branches:

    1. Short-range branch: Implemented with convolutional residual blocks to capture local image context and details, producing short-range feature Fs∈Rh×w×C\mathcal{F}_s \in \mathbb{R}^{h \times w \times C}.
    2. Long-range branch: Implemented with a patch-based SNR-aware transformer to capture global context via self-attention without propagating noise from low-SNR areas, producing long-range feature Fl∈Rh×w×C\mathcal{F}_l \in \mathbb{R}^{h \times w \times C}.

    The features Fs\mathcal{F}_s and Fl\mathcal{F}_l are dynamically merged via SNR-guided feature fusion into a combined feature F∈Rh×w×C\mathcal{F} \in \mathbb{R}^{h \times w \times C}. A decoder symmetric to the encoder (using pixel-shuffle layers for upsampling) with skip connections from the encoder processes F\mathcal{F} to predict an additive residual map R∈RH×W×3R \in \mathbb{R}^{H \times W \times 3}. The final enhanced image is computed as:

    I′=I+RI' = I + R

  2. Knowl 2 — SNR-Aware Masked Self-Attention Mechanism

    model/method

    To prevent noisy, low-SNR image regions from corrupting attention computations across the image, the SNR-aware transformer applies masked multi-head self-attention (MSA) based on local SNR estimates.

    The feature map F∈Rh×w×CF \in \mathbb{R}^{h \times w \times C} is partitioned into m=(h/p)×(w/p)m = (h/p) \times (w/p) non-overlapping patches of size p×p×Cp \times p \times C. The normalized SNR prior map S′∈[0,1]h×wS' \in [0, 1]^{h \times w} is partitioned into mm corresponding patches, and the spatial average value within patch ii is calculated as Si∈RS_i \in \mathbb{R} for i∈{1,…,m}i \in \{1, \dots, m\}.

    A binary patch validity mask vector S∈{0,1}m\mathcal{S} \in \{0, 1\}^m is obtained by thresholding with a threshold parameter ss:

    Si={0,Si<s1,Si≥s,i∈{1,…,m}\mathcal{S}_i = \begin{cases} 0, & S_i < s \\ 1, & S_i \ge s \end{cases}, \quad i \in \{1, \dots, m\}

    A mask matrix Smask′∈{0,1}m×mS_{\text{mask}}' \in \{0, 1\}^{m \times m} is formed by stacking mm identical row copies of S\mathcal{S}. For layer ii and attention head bb with query features Qi,b=qiWbqQ_{i,b} = q_i W_b^q, key features Ki,b=kiWbkK_{i,b} = k_i W_b^k, and value features Vi,b=viWbvV_{i,b} = v_i W_b^v (where qi,ki,vi∈Rm×(p2C)q_i, k_i, v_i \in \mathbb{R}^{m \times (p^2 C)} are LayerNorm projections, Wbq,Wbk,Wbv∈R(p2C)×CkW_b^q, W_b^k, W_b^v \in \mathbb{R}^{(p^2 C) \times C_k} are projection matrices, and CkC_k is the channel dimension per head), the SNR-guided attention is computed as:

    Attentioni,b(Qi,b,Ki,b,Vi,b)=Softmax(Qi,bKi,bTdk+(1−Smask′)σ)Vi,b\text{Attention}_{i,b}(Q_{i,b}, K_{i,b}, V_{i,b}) = \text{Softmax}\left(\frac{Q_{i,b} K_{i,b}^T}{\sqrt{d_k}} + (1 - S_{\text{mask}}')\sigma\right) V_{i,b}

    where dkd_k is the normalization scale factor and σ=−109\sigma = -10^9. Setting σ=−109\sigma = -10^9 ensures that tokens located in regions with Sj<sS_j < s receive zero attention weight after Softmax, eliminating noise propagation from low-SNR patches.

  3. Knowl 3 — Single-Image Signal-to-Noise Ratio Map Prior Estimation

    model/method

    The framework estimates an SNR prior map S∈RH×WS \in \mathbb{R}^{H \times W} from a single RGB input image I∈RH×W×3I \in \mathbb{R}^{H \times W \times 3} by treating noise as high-frequency spatial discontinuities between adjacent pixels.

    First, the grayscale image Ig∈RH×WI_g \in \mathbb{R}^{H \times W} is computed from II. A baseline clean approximation I^g\widehat{I}_g is generated using a no-learning denoising operation:

    I^g=denoise(Ig)\widehat{I}_g = \text{denoise}(I_g)

    where local pixel averaging (local means) is used by default. The estimated spatial noise map N∈RH×WN \in \mathbb{R}^{H \times W} is computed as the absolute difference between the noisy grayscale image and its denoised counterpart:

    N=∣Ig−I^g∣N = |I_g - \widehat{I}_g|

    The pixel-wise SNR prior map S∈RH×WS \in \mathbb{R}^{H \times W} is then calculated as:

    S=I^gNS = \frac{\widehat{I}_g}{N}

    This approximate SNR map serves as a spatial guidance prior for both attention masking in the transformer and feature fusion across branches.

  4. Knowl 4 — SNR-Guided Spatially-Varying Feature Fusion

    equation

    To dynamically adjust the balance between short-range local convolutional operations and long-range global transformer operations according to regional degradation levels, the estimated SNR prior map S∈RH×WS \in \mathbb{R}^{H \times W} is bilinearly interpolated to the feature spatial dimensions h×wh \times w and normalized to [0,1][0, 1], yielding S′∈[0,1]h×wS' \in [0, 1]^{h \times w}.

    The short-range feature Fs∈Rh×w×C\mathcal{F}_s \in \mathbb{R}^{h \times w \times C} and long-range feature Fl∈Rh×w×C\mathcal{F}_l \in \mathbb{R}^{h \times w \times C} are fused via spatially varying linear interpolation:

    F=Fs⊙S′+Fl⊙(1−S′)\mathcal{F} = \mathcal{F}_s \odot S' + \mathcal{F}_l \odot (1 - S')

    where ⊙\odot represents element-wise multiplication broadcast across the channel dimension CC. Pixels in high-SNR regions (S′→1S' \to 1) are dominated by local convolutional features Fs\mathcal{F}_s, whereas pixels in low-SNR regions (S′→0S' \to 0) rely on non-local context from the transformer features Fl\mathcal{F}_l.

  5. Knowl 5 — Composite Charbonnier and Perceptual Reconstruction Loss

    equation

    The low-light enhancement network is trained end-to-end using a weighted combination of a Charbonnier reconstruction loss and a VGG-based perceptual loss:

    L=Lr+λLvggL = L_r + \lambda L_{vgg}

    The Charbonnier loss LrL_r between the network output I′I' and ground-truth image I^\widehat{I} is formulated as:

    Lr=∥I′−I^∥22+ϵ2L_r = \sqrt{\lVert I' - \widehat{I} \rVert_2^2 + \epsilon^2}

    where ϵ=10−3\epsilon = 10^{-3}.

    The perceptual loss LvggL_{vgg} measures feature-space L1L_1 distance using deep representations extracted by a pretrained VGG network Φ\Phi:

    Lvgg=∥Φ(I′)−Φ(I^)∥1L_{vgg} = \lVert \Phi(I') - \Phi(\widehat{I}) \rVert_1

    where λ\lambda is a hyperparameter balancing the reconstruction and perceptual loss terms.

  6. Knowl 6 — Low-Light Enhancement Performance on LOL Benchmark Datasets

    data/table

    The SNR-aware framework was evaluated against traditional, Retinex-based, convolutional, and vision transformer enhancement methods on the LOL-v1 (15 test pairs), LOL-v2-real (100 test pairs), and LOL-v2-synthetic datasets using Peak Signal-to-Noise Ratio (PSNR, dB) and Structural Similarity Index (SSIM).

    Dataset LOL-v1 LOL-v2-real LOL-v2-synthetic
    Method PSNR SSIM PSNR SSIM PSNR SSIM
    Dong 16.72 0.580 17.26 0.527 16.90 0.749
    LIME 16.76 0.560 15.24 0.470 16.88 0.776
    MF 18.79 0.640 18.73 0.559 17.50 0.751
    SRIE 11.86 0.500 17.34 0.686 14.50 0.616
    BIMEF 13.86 0.580 17.85 0.653 17.20 0.713
    DRD 16.77 0.560 15.47 0.567 17.13 0.798
    RRM 13.88 0.660 17.34 0.686 17.15 0.727
    SID 14.35 0.436 13.24 0.442 15.04 0.610
    DeepUPE 14.38 0.446 13.27 0.452 15.08 0.623
    KIND 20.87 0.800 14.74 0.641 13.29 0.578
    DeepLPF 15.28 0.473 14.10 0.480 16.02 0.587
    FIDE 18.27 0.665 16.85 0.678 15.20 0.612
    LPNet 21.46 0.802 17.80 0.792 19.51 0.846
    MIR-Net 24.14 0.830 20.02 0.820 21.94 0.876
    RF 15.23 0.452 14.05 0.458 15.97 0.632
    3DLUT 14.35 0.445 17.59 0.721 18.04 0.800
    A3DLUT 14.77 0.458 18.19 0.745 18.92 0.838
    Band 20.13 0.830 20.29 0.831 23.22 0.927
    EG 17.48 0.650 18.23 0.617 16.57 0.734
    Retinex 18.23 0.720 18.37 0.723 16.55 0.652
    Sparse 17.20 0.640 20.06 0.816 22.05 0.905
    IPT 16.27 0.504 19.80 0.813 18.30 0.811
    Uformer 16.36 0.507 18.82 0.771 19.66 0.871
    Ours 24.61 0.842 21.48 0.849 24.14 0.928

    The SNR-aware model achieves the highest PSNR and SSIM across all three LOL benchmarks, outperforming pure transformer restoration baselines (IPT by +8.34 dB on LOL-v1; Uformer by +8.25 dB on LOL-v1) and convolutional enhancement architectures.

  7. Knowl 7 — Low-Light Enhancement Performance on Extreme-Dark and Dynamic Video Benchmarks

    data/table

    The framework was evaluated on extreme low-light raw-to-RGB datasets (SID Sony subset and full SMID) and on static subsets of the SDSD dataset (indoor and outdoor splits) using PSNR (dB) and SSIM.

    Dataset SID SMID SDSD-indoor SDSD-outdoor
    Method PSNR SSIM PSNR SSIM PSNR SSIM PSNR SSIM
    DRD 16.48 0.578 22.83 0.684 20.84 0.617 20.96 0.629
    SID 16.97 0.591 24.78 0.718 23.29 0.703 24.90 0.693
    DeepUPE 17.01 0.604 23.91 0.690 21.70 0.662 21.94 0.698
    KIND 18.02 0.583 22.18 0.634 21.95 0.672 21.97 0.654
    DeepLPF 18.07 0.600 24.36 0.688 22.21 0.664 22.76 0.658
    FIDE 18.34 0.578 24.42 0.692 22.41 0.659 22.20 0.629
    LPNet 20.08 0.598 26.55 0.772 23.87 0.841 22.09 0.629
    MIR-Net 20.84 0.605 25.66 0.762 24.38 0.864 27.13 0.837
    RF 16.44 0.596 23.11 0.681 20.97 0.655 21.21 0.689
    3DLUT 20.11 0.592 23.86 0.678 21.66 0.655 21.89 0.649
    A3DLUT 20.32 0.595 24.56 0.684 22.39 0.656 22.95 0.692
    Band 19.02 0.577 26.60 0.781 24.08 0.868 25.77 0.841
    EG 17.23 0.543 22.62 0.674 20.02 0.604 20.10 0.616
    Retinex 18.44 0.581 25.88 0.744 23.17 0.696 23.84 0.743
    Sparse 18.68 0.606 25.48 0.766 23.25 0.863 25.28 0.804
    IPT 20.53 0.561 27.03 0.783 26.11 0.831 27.55 0.850
    Uformer 18.54 0.577 27.20 0.792 23.17 0.859 23.85 0.748
    Ours 22.87 0.625 28.49 0.805 29.44 0.894 28.66 0.866

    The method outperforms all 17 competing baseline approaches on all four datasets, demonstrating resilience to extreme sensor noise in the RGB domain.

  8. Knowl 8 — Ablation Analysis of Architecture Components

    data/table

    An ablation study evaluated the performance contributions of individual model components across all seven benchmarks:

    • Ours w/o L: Removes the long-range transformer branch (pure convolutional network).
    • Ours w/o S: Removes the short-range convolutional branch (pure transformer network with SNR attention).
    • Ours w/o SA: Removes both the short-range convolutional branch and SNR-guided attention (pure baseline transformer in the deepest layer).
    • Ours w/o A: Retains dual branches and SNR fusion, but removes SNR-guided attention masking from the transformer.
    • Ours: Full model with dual branches, SNR-guided masked self-attention, and SNR-guided feature fusion.
    LOL-v1 LOL-v2-real LOL-v2-synth SID SMID SDSD-in SDSD-out
    Configuration PSNR SSIM PSNR SSIM PSNR SSIM PSNR SSIM PSNR SSIM PSNR SSIM PSNR SSIM
    Ours w/o L 16.27 0.638 16.98 0.687 20.81 0.881 19.10 0.593 26.20 0.776 22.24 0.818 20.03 0.713
    Ours w/o S 23.06 0.828 18.98 0.790 23.47 0.919 22.30 0.604 27.00 0.768 28.13 0.884 25.43 0.823
    Ours w/o SA 20.67 0.752 18.85 0.765 21.88 0.842 21.02 0.544 27.01 0.774 25.78 0.839 24.57 0.832
    Ours w/o A 21.86 0.760 19.40 0.782 22.23 0.866 21.19 0.550 26.87 0.769 27.36 0.874 26.62 0.857
    Ours (Full) 24.61 0.842 21.48 0.849 24.14 0.928 22.87 0.625 28.49 0.805 29.44 0.894 28.37 0.862

    The full configuration achieves the highest metrics across all datasets. Dual-branch modeling yields substantial improvements over single-branch designs (e.g., +8.34 dB on LOL-v1 over convolutional-only "Ours w/o L"), and SNR-guided attention provides consistent gains over unmasked attention (e.g., +2.75 dB on LOL-v1 over "Ours w/o A").

  9. Knowl 9 — Insensitivity to the Choice of Denoising Operator for SNR Estimation

    empirical result

    The low-light enhancement framework was evaluated using three different no-learning denoising algorithms for estimating the SNR prior map I^g=denoise(Ig)\widehat{I}_g = \text{denoise}(I_g):

    1. Local Means (local pixel averaging)
    2. Non-local Means
    3. BM3D

    Testing across all seven benchmark datasets (LOL-v1, LOL-v2-real, LOL-v2-synthetic, SID, SMID, SDSD-indoor, and SDSD-outdoor) revealed nearly identical PSNR and SSIM performance regardless of which denoising operator was used. Because the downstream transformer masking and feature fusion require only relative spatial discrimination of signal-to-noise levels rather than ground-truth SNR values, fast local pixel averaging is sufficient, avoiding the computational overhead of advanced denoising methods like BM3D while outperforming competing baseline models.

  10. Knowl 10 — Perceptual Quality User Study on Real Smartphone Low-Light Captures

    empirical result

    A user study was conducted with 100 participants evaluating the visual enhancement quality on 30 low-light images captured by iPhone X and Huawei P30 smartphones across diverse scenes (roads, parks, libraries, schools, portraits), where more than 50% of the pixels had luminance values below 30%.

    The SNR-aware model (trained on SDSD-outdoor) was compared against the five strongest baseline methods (MIR-Net, IPT, Uformer, Band, LPNet) on six perceptual questions using a 1-to-5 Likert scale:

    1. Detail perceivability (Q1)
    2. Color vividness (Q2)
    3. Visual realism (Q3)
    4. Freedom from overexposure (Q4)
    5. Freedom from noise (Q5)
    6. Overall rating (Q6)

    The proposed framework received a significantly higher distribution of top ratings (score 5) and fewer lowest ratings (score 1) across all six categories. Paired two-tailed tt-tests between the proposed framework and each baseline across all questions produced pp-values with p<0.001p < 0.001, confirming statistically significant superiority in human perceptual preference.

  11. Knowl 11 — Limitations and Future Directions in SNR-Aware Low-Light Enhancement

    limitation

    The SNR-aware low-light enhancement framework has three identified limitations:

    1. Semantic agnostic spatial variation: The spatial adaptation relies entirely on pixel/patch-level Signal-to-Noise Ratio (SNR) estimates without incorporating higher-level semantic scene understanding to guide enhancement operations.
    2. Lack of temporal modeling: The method is formulated for static 2D images and does not incorporate temporal coherence mechanisms required for processing continuous low-light video streams.
    3. Degradation in near-zero photon regions: In extremely dark, near-black areas where physical photons are almost completely absent, deterministic reconstruction and spatial feature aggregation reach fundamental information limits, requiring generative prior models to synthesize plausible content.

Coverage note — None omitted; all core architectural mechanisms, equations, empirical benchmarks across 7 datasets, ablation configurations, user studies, and limitations from the paper are represented.

References

  1. 1.Antoni Buades, Bartomeu Coll, and J-M. Morel. A non-local algorithm for image denoising. In IEEE Conf. Comput. Vis. Pattern Recog., 2005. 4, 8
  2. 2.Jianrui Cai, Shuhang Gu, and Lei Zhang. Learning a deep single image contrast enhancer from multi-exposure images. IEEE Trans. Image Process., 2018. 2
  3. 3.Damon M. Chandler and Sheila S. Hemami. VSNR: A wavelet-based visual signal-to-noise ratio for natural images. IEEE Trans. Image Process., 2007. 1
  4. 4.Chen Chen, Qifeng Chen, Minh N. Do, and Vladlen Koltun. Seeing motion in the dark. In Int. Conf. Comput. Vis., 2019. 2, 5
  5. 5.Chen Chen, Qifeng Chen, Jia Xu, and Vladlen Koltun. Learning to see in the dark. In IEEE Conf. Comput. Vis. Pattern Recog., 2018. 2, 5, 6, 7
  6. 6.Hanting Chen, Yunhe Wang, Tianyu Guo, Chang Xu, Yiping Deng, Zhenhua Liu, Siwei Ma, Chunjing Xu, Chao Xu, and Wen Gao. Pre-trained image processing transformer. In IEEE Conf. Comput. Vis. Pattern Recog., 2021. 2, 3, 5, 6, 7
  7. 7.Yu-Sheng Chen, Yu-Ching Wang, Man-Hsin Kao, and Yung-Yu Chuang. Deep Photo Enhancer: Unpaired learning for image enhancement from photographs with GANs. In IEEE Conf. Comput. Vis. Pattern Recog., 2018. 2
  8. 8.Kostadin Dabov, Alessandro Foi, Vladimir Katkovnik, and Karen Egiazarian. Image denoising with block-matching and 3D filtering. In Image Processing: Algorithms and Systems, Neural Networks, and Machine Learning, 2006. 4, 8
  9. 9.Xuan Dong, Guan Wang, Yi Pang, Weixin Li, Jiangtao Wen, Wei Meng, and Yao Lu. Fast efficient algorithm for enhancement of low lighting video. In Int. Conf. Multimedia and Expo, 2011. 5, 6
  10. 10.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. In Int. Conf. Learn. Represent., 2021. 3
  11. 11.Xueyang Fu, Delu Zeng, Yue Huang, Yinghao Liao, Xinghao Ding, and John Paisley. A fusion-based enhancing method for weakly illuminated images. Signal Processing, 2016. 5, 6
  12. 12.Xueyang Fu, Delu Zeng, Yue Huang, Xiao-Ping Zhang, and Xinghao Ding. A weighted variational model for simultaneous reflectance and illumination estimation. In IEEE Conf. Comput. Vis. Pattern Recog., 2016. 1, 5, 6
  13. 13.Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. Generative Adversarial Nets. In Adv. Neural Inform. Process. Syst., 2014. 8
  14. 14.Chunle Guo, Chongyi Li, Jichang Guo, Chen Change Loy, Junhui Hou, Sam Kwong, and Runmin Cong. Zero-reference deep curve estimation for low-light image enhancement. In IEEE Conf. Comput. Vis. Pattern Recog., 2020. 2
  15. 15.Xiaojie Guo, Yu Li, and Haibin Ling. LIME: Low-light image enhancement via illumination map estimation. IEEE Trans. Image Process., 2016. 1, 5, 6
  16. 16.Kai Han, Yunhe Wang, Hanting Chen, Xinghao Chen, Jianyuan Guo, Zhenhua Liu, Yehui Tang, An Xiao, Chunjing Xu, Yixing Xu, et al. A survey on visual transformer. IEEE Trans. Pattern Anal. Mach. Intell., 2022. 3
  17. 17.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In IEEE Conf. Comput. Vis. Pattern Recog., 2016. 1
  18. 18.Yukun Huang, Zheng-Jun Zha, Xueyang Fu, Richang Hong, and Liang Li. Real-world person re-identification via degradation invariance learning. In IEEE Conf. Comput. Vis. Pattern Recog., 2020. 1
  19. 19.Andrey Ignatov, Nikolay Kobyshev, Radu Timofte, Kenneth Vanhoey, and Luc Van Gool. WESPE: Weakly supervised photo enhancer for digital cameras. In IEEE Conf. Comput. Vis. Pattern Recog., 2018. 2
  20. 20.Yifan Jiang, Xinyu Gong, Ding Liu, Yu Cheng, Chen Fang, Xiaohui Shen, Jianchao Yang, Pan Zhou, and Zhangyang Wang. EnlightenGAN: Deep light enhancement without paired supervision. IEEE Trans. Image Process., 2021. 2, 5, 6, 7
  21. 21.Salman Khan, Muzammal Naseer, Munawar Hayat, Syed Waqas Zamir, Fahad Shahbaz Khan, and Mubarak Shah. Transformers in vision: A survey. ACM Computing Surveys (CSUR), 2021. 3
  22. 22.Hanul Kim, Su-Min Choi, Chang-Su Kim, and Yeong Jun Koh. Representative color transform for image enhancement. In Int. Conf. Comput. Vis., 2021. 2, 6
  23. 23.Diederik P. Kingma and Jimmy Ba. Adam: A method for stochastic optimization. arXiv:1412.6980, 2014. 5
  24. 24.Satoshi Kosugi and Toshihiko Yamasaki. Unpaired image enhancement featuring reinforcement-learning-controlled image editing software. In AAAI, 2020. 5, 6, 7
  25. 25.Wei-Sheng Lai, Jia-Bin Huang, Narendra Ahuja, and Ming-Hsuan Yang. Fast and accurate image super-resolution with deep laplacian pyramid networks. IEEE Trans. Pattern Anal. Mach. Intell., 2018. 5
  26. 26.Jiaqian Li, Juncheng Li, Faming Fang, Fang Li, and Guixu Zhang. Luminance-aware pyramid network for low-light image enhancement. IEEE Trans. Multimedia, 2020. 5, 6, 7
  27. 27.Jing Li, Stan Z. Li, Quan Pan, and Tao Yang. Illumination and motion-based video enhancement for night surveillance. In IEEE International Workshop on Visual Surveillance and Performance Evaluation of Tracking and Surveillance, 2005. 1
  28. 28.Mading Li, Jiaying Liu, Wenhan Yang, Xiaoyan Sun, and Zongming Guo. Structure-revealing low-light image enhancement via robust Retinex model. IEEE Trans. Image Process., 2018. 2, 5, 6
  29. 29.Risheng Liu, Long Ma, Jiaao Zhang, Xin Fan, and Zhongxuan Luo. Retinex-inspired unrolling with cooperative prior architecture search for low-light image enhancement. In IEEE Conf. Comput. Vis. Pattern Recog., 2021. 1, 2, 5, 6, 7
  30. 30.Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. In Int. Conf. Comput. Vis., 2021. 3
  31. 31.Kin Gwn Lore, Adedotun Akintayo, and Soumik Sarkar. LLNet: A deep autoencoder approach to natural low-light image enhancement. Pattern Recognition, 2017. 2
  32. 32.Mehdi Mirza and Simon Osindero. Conditional Generative Adversarial Nets. arXiv:1411.1784, 2014. 8
  33. 33.Sean Moran, Pierre Marza, Steven McDonagh, Sarah Parisot, and Gregory Slabaugh. DeepLPF: Deep local parametric filters for image enhancement. In IEEE Conf. Comput. Vis. Pattern Recog., 2020. 2, 5, 6, 7
  34. 34.Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al. Pytorch: An imperative style, high-performance deep learning library. In Adv. Neural Inform. Process. Syst., 2019. 5
  35. 35.Zia-ur Rahman, Daniel J. Jobson, and Glenn A. Woodell. Retinex processing for automatic image enhancement. IEEE Trans. Image Process., 2004. 2
  36. 36.Wenzhe Shi, Jose Caballero, Ferenc Huszar, Johannes Totz, Andrew P. Aitken, Rob Bishop, Daniel Rueckert, and Zehan Wang. Real-time single image and video super-resolution using an efficient sub-pixel convolutional neural network. In IEEE Conf. Comput. Vis. Pattern Recog., 2016. 5
  37. 37.Karen Simonyan and Andrew Zisserman. Very deep convolutional networks for large-scale image recognition. In Int. Conf. Learn. Represent., 2014. 5
  38. 38.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Adv. Neural Inform. Process. Syst., 2017. 1, 3
  39. 39.Ruixing Wang, Xiaogang Xu, Chi-Wing Fu, Jiangbo Lu, Bei Yu, and Jiaya Jia. Seeing dynamic scene in the dark: High-quality video dataset with mechatronic alignment. In Int. Conf. Comput. Vis., 2021. 2, 5, 7
  40. 40.Ruixing Wang, Qing Zhang, Chi-Wing Fu, Xiaoyong Shen, Wei-Shi Zheng, and Jiaya Jia. Underexposed photo enhancement using deep illumination estimation. In IEEE Conf. Comput. Vis. Pattern Recog., 2019. 2, 5, 6, 7
  41. 41.Shuhang Wang, Jin Zheng, Hai-Miao Hu, and Bo Li. Naturalness preserved enhancement algorithm for non-uniform illumination images. IEEE Trans. Image Process., 2013. 1
  42. 42.Tao Wang, Yong Li, Jingyang Peng, Yipeng Ma, Xian Wang, Fenglong Song, and Youliang Yan. Real-time image enhancer via learnable spatial-aware 3D lookup tables. In Int. Conf. Comput. Vis., 2021. 2, 5, 6, 7
  43. 43.Zhou Wang, Alan C. Bovik, Hamid R. Sheikh, and Eero P. Simoncelli. Image quality assessment: from error visibility to structural similarity. IEEE Trans. Image Process., 2004. 6
  44. 44.Zhendong Wang, Xiaodong Cun, Jianmin Bao, and Jianzhuang Liu. Uformer: A general U-Shaped transformer for image restoration. In IEEE Conf. Comput. Vis. Pattern Recog., 2022. 3, 5, 6, 7
  45. 45.Chen Wei, Wenjing Wang, Wenhan Yang, and Jiaying Liu. Deep Retinex decomposition for low-light enhancement. In Brit. Mach. Vis. Conf., 2018. 2, 5, 6, 7
  46. 46.Haiping Wu, Bin Xiao, Noel Codella, Mengchen Liu, Xiyang Dai, Lu Yuan, and Lei Zhang. CvT: Introducing convolutions to vision transformers. In Int. Conf. Comput. Vis., 2021. 3
  47. 47.Bing Xu, Naiyan Wang, Tianqi Chen, and Mu Li. Empirical evaluation of rectified activations in convolutional network. In ICML, 2015. 5
  48. 48.Ke Xu, Xin Yang, Baocai Yin, and Rynson WH. Lau. Learning to restore low-light images via decomposition-and-enhancement. In IEEE Conf. Comput. Vis. Pattern Recog., 2020. 1, 2, 5, 6, 7
  49. 49.Jianzhou Yan, Stephen Lin, Bing Kang Sing, and Xiaoou Tang. A learning-to-rank approach for image color enhancement. In IEEE Conf. Comput. Vis. Pattern Recog., 2014. 2
  50. 50.Zhicheng Yan, Hao Zhang, Baoyuan Wang, Sylvain Paris, and Yizhou Yu. Automatic photo adjustment using deep neural networks. ACM Trans. Graph., 2016. 2
  51. 51.Wenhan Yang, Shiqi Wang, Yuming Fang, Yue Wang, and Jiaying Liu. From fidelity to perceptual quality: A semi-supervised approach for low-light image enhancement. In IEEE Conf. Comput. Vis. Pattern Recog., 2020. 2
  52. 52.Wenhan Yang, Shiqi Wang, Yuming Fang, Yue Wang, and Jiaying Liu. Band representation-based semi-supervised low-light image enhancement: Bridging the gap between signal fidelity and perceptual quality. IEEE Trans. Image Process., 2021. 2, 5, 6, 7
  53. 53.Wenhan Yang, Wenjing Wang, Haofeng Huang, Shiqi Wang, and Jiaying Liu. Sparse gradient regularized deep Retinex network for robust low-light image enhancement. IEEE Trans. Image Process., 2021. 2, 5, 6, 7
  54. 54.Susu Yao, Weisi Lin, EePing Ong, and Zhongkang Lu. Contrast Signal-to-Noise Ratio for image quality assessment. In IEEE Int. Conf. Image Process., 2005. 1
  55. 55.Zhenqiang Ying, Ge Li, and Wen Gao. A bio-inspired multi-exposure fusion framework for low-light image enhancement. arXiv:1711.00591, 2017. 5, 6
  56. 56.Zhenqiang Ying, Ge Li, Yurui Ren, Ronggang Wang, and Wenmin Wang. A new low-light image enhancement algorithm using camera response model. In Int. Conf. Comput. Vis., 2017. 1
  57. 57.Kun Yuan, Shaopeng Guo, Ziwei Liu, Aojun Zhou, Fengwei Yu, and Wei Wu. Incorporating convolution designs into visual transformers. In Int. Conf. Comput. Vis., 2021. 3
  58. 58.Li Yuan, Yunpeng Chen, Tao Wang, Weihao Yu, Yujun Shi, Zihang Jiang, Francis EH. Tay, Jiashi Feng, and Shuicheng Yan. Tokens-to-Token ViT: Training vision transformers from scratch on ImageNet. In Int. Conf. Comput. Vis., 2021. 3
  59. 59.Syed Waqas Zamir, Aditya Arora, Salman Khan, Munawar Hayat, Fahad Shahbaz Khan, Ming-Hsuan Yang, and Ling Shao. Learning enriched features for real image restoration and enhancement. In Eur. Conf. Comput. Vis., 2020. 2, 5, 6, 7
  60. 60.Hui Zeng, Jianrui Cai, Lida Li, Zisheng Cao, and Lei Zhang. Learning image-adaptive 3D lookup tables for high performance photo enhancement in real-time. IEEE Trans. Pattern Anal. Mach. Intell., 2020. 2, 5, 6, 7
  61. 61.Yonghua Zhang, Jiawan Zhang, and Xiaojie Guo. Kindling the darkness: A practical low-light image enhancer. In ACM Int. Conf. Multimedia, 2019. 5, 6, 7
  62. 62.Lin Zhao, Shao-Ping Lu, Tao Chen, Zhenglu Yang, and Ariel Shamir. Deep symmetric network for underexposed image enhancement with recurrent attentional learning. In Int. Conf. Comput. Vis., 2021. 2, 6
  63. 63.Chuanjun Zheng, Daming Shi, and Wentian Shi. Adaptive unfolding total variation network for low-light image enhancement. In Int. Conf. Comput. Vis., 2021. 2

Citation

MLA
Xu, X., et al. “SNR-Aware Low-light Image Enhancement”. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022, pp. 17693–703, https://doi.org/10.1109/CVPR52688.2022.01719.
APA
Xu, X., Wang, R., Fu, C.-W., & Jia, J. (2022). SNR-Aware Low-light Image Enhancement. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 17693–17703. https://doi.org/10.1109/CVPR52688.2022.01719
Chicago
Xu, X., R. Wang, C.-W. Fu, and J. Jia. 2022. “SNR-Aware Low-light Image Enhancement”. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 17693–703. https://doi.org/10.1109/CVPR52688.2022.01719.
Harvard
Xu, X. et al. (2022) “SNR-Aware Low-light Image Enhancement”, 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp. 17693–17703. Available at: https://doi.org/10.1109/CVPR52688.2022.01719.
Vancouver
1. Xu X, Wang R, Fu C-W, Jia J (2022) SNR-Aware Low-light Image Enhancement. In: 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp 17693–17703

BibTeX

@inproceedings{Xu_2022, title={SNR-Aware Low-light Image Enhancement}, url={http://dx.doi.org/10.1109/CVPR52688.2022.01719}, DOI={10.1109/cvpr52688.2022.01719}, booktitle={2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)}, publisher={IEEE}, author={Xu, Xiaogang and Wang, Ruixing and Fu, Chi-Wing and Jia, Jiaya}, year={2022}, month=June, pages={17693–17703} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE