Fourier Priors-Guided Diffusion for Zero-Shot Joint Low-Light Enhancement and Deblurring

Xiaoqian LvShengping ZhangChenyang WangYichen ZhengBineng ZhongChongyi LiLiqiang Nie

article2024CVPR107 citations

Proposes FourierDiff, a zero-shot framework that integrates Fourier amplitude and phase priors into pre-trained diffusion models to jointly brighten dark images and remove motion blur without requiring paired training data or explicit degradation assumptions.

Listen

Low-light photography frequently suffers from simultaneous degradation: limited illumination and severe motion blur caused by long camera exposure times. Traditional image processing methods usually tackle brightness enhancement and deblurring separately or rely heavily on synthetic training data and assumed degradation models. When applied to complex real-world night photography, existing techniques often create severe visual artifacts, introduce color distortion, or fail to remove blur.

The article introduces and evaluates FourierDiff, a novel zero-shot framework that performs joint low-light enhancement and motion deblurring without requiring paired training data or predefined degradation assumptions. The framework leverages the frequency domain by separating brightness information into Fourier amplitudes and structural content into Fourier phases, using a pre-trained diffusion model to restore degraded images.

The authors evaluated the framework across standard synthetic datasets (12,000 image pairs from LOL-Blur) and real-world benchmarks (482 test images from RealBlur) using standard blind perceptual quality metrics (NIQE, PI, BRISQUE, and MUSIQ) alongside full-reference metrics. An unconditional diffusion model pre-trained on ImageNet served as the foundation, paired with an alternating optimization strategy designed to iteratively refine blur kernel estimation during the diffusion reverse sampling process. A subjective user study involving 40 participants across 20 test scenes was also conducted to gauge visual preference.

The evaluation yielded several key findings. First, the proposed framework outperformed all state-of-the-art baselines across every no-reference perceptual quality metric on both synthetic and real-world datasets, achieving top scores such as a RealBlur BRISQUE score of 26.39 compared to 34.80–45.89 for competing sequential and joint pipelines. Second, the method achieved results comparable to fully supervised methods on synthetic benchmarks while generalizing substantially better to real-world night scenes without generating artificial halos or excessive noise. Third, human evaluation showed a decisive preference for the approach, with participants favoring the proposed method over competing methods in 70.5% to 97.0% of visual comparisons. Finally, an ablation analysis demonstrated that updating blur kernel estimates during diffusion sampling produced clear performance improvements, with optimal efficiency and quality achieved at an alternating step interval of 200.

These results demonstrate that separating image restoration tasks into distinct frequency components allows pre-trained generative models to restore complex, real-world image degradations without the risk of overfitting associated with synthetic training pairs. Eliminating the requirement for specialized paired datasets significantly reduces development costs and deployment risks for computer vision applications operating in adverse, low-light environments, such as surveillance, night-time action recognition, and autonomous navigation.

Organizations seeking to enhance image quality in low-light environments should consider adopting frequency-guided zero-shot diffusion pipelines over rigid, supervised networks, especially where diverse and unmodeled real-world blur is present. To implement this approach effectively, practitioners should incorporate adjustable brightness parameters to tune outputs according to specific operational needs.

While the framework shows high efficacy on standard night-time imagery, confidence should be tempered in scenarios involving extreme darkness, where the complete loss of initial content guidance degrades structural reconstruction. Additionally, because the architecture relies on iterative diffusion sampling over hundreds of steps, high computational latency remains a limitation that currently precludes real-time deployment without further optimization.

Cover for Fourier Priors-Guided Diffusion for Zero-Shot Joint Low-Light Enhancement and Deblurring

Abstract

Existing joint low-light enhancement and deblurring methods learn pixel-wise mappings from paired synthetic data, which results in limited generalization in real-world scenes. While some studies explore the rich generative prior of pre-trained diffusion models, they typically rely on the assumed degradation process and cannot handle unknown real-world degradations well. To address these problems, we propose a novel zero-shot framework, FourierDiff, which embeds Fourier priors into a pre-trained diffusion model to harmoniously handle the joint degradation of luminance and structures. FourierDiff is appealing in its relaxed requirements on paired training data and degradation assumptions. The key zero-shot insight is motivated by image characteristics in the Fourier domain: most luminance information concentrates on amplitudes while structure and content information are closely related to phases. Based on this observation, we decompose the sampled results of the reverse diffusion process in the Fourier domain and take advantage of the amplitude of the generative prior to align the enhanced brightness with the distribution of natural images. To yield a sharp and content-consistent enhanced result, we further design a spatial-frequency alternating optimization strategy to progressively refine the phase of the input. Extensive experiments demonstrate the superior effectiveness of the proposed method, especially in real-world scenes. The code is available at https://github.com/aipixel/FourierDiff.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 2.1. Low-Light Image Enhancement
  • 2.2. Image Deblurring
  • 2.3. Diffusion-based Image Restoration
  • 3. Preliminary
  • 4. Methodology
  • 4.1. Fourier Priors-Guided Diffusion Sampling
  • 4.2. Spatial-Frequency Alternating Optimization
  • 5. Experiments
  • 5.1. Datasets and Evaluation Metrics
  • 5.2. Implementation Details
  • 5.3. Comparison with State-of-the-art Methods
  • 5.4. User Study
  • 5.5. Ablation Study
  • 6. Conclusion
  • References

Knowls

  1. Knowl 1 — Fourier Domain Amplitude-Phase Separation Principle in FourierDiff

    model/method

    FourierDiff is a zero-shot joint low-light enhancement and deblurring framework that leverages the Fourier domain decoupling of image luminance and structure to guide an unconditional pre-trained diffusion model without paired training data or explicit degradation assumptions.

    In the 2D discrete Fourier domain, swapping the amplitude and phase of two images reveals distinct physical roles:

    1. Luminance and Amplitude: Luminance and global illumination priors concentrate primarily on the Fourier amplitude spectrum AA. The reverse sampling states of a diffusion model pre-trained on natural images provide rich natural illumination priors in their amplitude spectrum.

    2. Structure, Content, and Phase: Semantic content and structural details are primarily governed by the Fourier phase spectrum PP. The degraded input image's phase spectrum preserves scene content and encodes blur patterns as repetitive high-frequency edge artifacts.

    FourierDiff uses this separation to extract natural brightness from the generative prior via Fourier amplitude while preserving spatial content and iteratively removing blur via Fourier phase manipulation.

  2. Knowl 2 — Fourier Priors-Guided Diffusion Sampling

    model/method

    In FourierDiff, reverse diffusion sampling is guided by recombining the Fourier amplitude of intermediate diffusion predictions with the Fourier phase of the degraded input image at each time step tt.

    Given the current noisy state xt∈RH×W×C\mathbf{x}_t \in \mathbb{R}^{H \times W \times C} at diffusion step tt, the clean image estimate x0∣t\mathbf{x}_{0|t} is calculated via Tweedie's formula:

    x0∣t=1αˉt(xt−ϵθ(xt,t)1−αˉt)\mathbf{x}_{0|t} = \frac{1}{\sqrt{\bar{\alpha}_t}}\left(\mathbf{x}_t - \boldsymbol{\epsilon}_\theta(\mathbf{x}_t, t)\sqrt{1 - \bar{\alpha}_t}\right)

    where ϵθ(⋅)\boldsymbol{\epsilon}_\theta(\cdot) is the pre-trained unconditional score network, αt=1−βt\alpha_t = 1 - \beta_t, and αˉt=∏i=0tαi\bar{\alpha}_t = \prod_{i=0}^t \alpha_i.

    Both x0∣t\mathbf{x}_{0|t} and the current refined input guide yt\mathbf{y}_t are transformed via the 2D Fast Fourier Transform (FFT):

    (Ayt,Pyt)=FFT(yt)(A_{\mathbf{y}_t}, P_{\mathbf{y}_t}) = \text{FFT}(\mathbf{y}_t)

    (Ax0∣t,Px0∣t)=FFT(x0∣t)(A_{\mathbf{x}_{0|t}}, P_{\mathbf{x}_{0|t}}) = \text{FFT}(\mathbf{x}_{0|t})

    where AA and PP denote amplitude and phase spectra, respectively. To combine generative illumination priors with spatial content fidelity, the rectified clean estimate x^0∣t\hat{\mathbf{x}}_{0|t} is reconstructed by combining amplitudes and retaining the input phase via the Inverse Fast Fourier Transform (IFFT):

    x^0∣t=IFFT(γAx0∣t+Ayt,Pyt)\hat{\mathbf{x}}_{0|t} = \text{IFFT}(\gamma A_{\mathbf{x}_{0|t}} + A_{\mathbf{y}_t}, P_{\mathbf{y}_t})

    where γ\gamma is an adaptive brightness scaling parameter. The next reverse diffusion state xt−1\mathbf{x}_{t-1} is sampled from the conditional distribution:

    pθ(xt−1∣xt,x^0∣t)=N(xt−1;μt(xt,x^0∣t),σt2I)p_\theta(\mathbf{x}_{t-1} \mid \mathbf{x}_t, \hat{\mathbf{x}}_{0|t}) = \mathcal{N}\left(\mathbf{x}_{t-1}; \boldsymbol{\mu}_t(\mathbf{x}_t, \hat{\mathbf{x}}_{0|t}), \sigma_t^2 \mathbf{I}\right)

    with mean and variance defined as:

    μt(xt,x^0∣t)=αˉt−1βt1−αˉtx^0∣t+αt(1−αˉt−1)1−αˉtxt\boldsymbol{\mu}_t(\mathbf{x}_t, \hat{\mathbf{x}}_{0|t}) = \frac{\sqrt{\bar{\alpha}_{t-1}}\beta_t}{1 - \bar{\alpha}_t}\hat{\mathbf{x}}_{0|t} + \frac{\sqrt{\alpha_t}(1 - \bar{\alpha}_{t-1})}{1 - \bar{\alpha}_t}\mathbf{x}_t

    σt2=1−αˉt−11−αˉtβt\sigma_t^2 = \frac{1 - \bar{\alpha}_{t-1}}{1 - \bar{\alpha}_t}\beta_t

  3. Knowl 3 — Adaptive Non-Reference Brightness Control Loss

    equation

    To accommodate subjective user preferences and diverse exposure conditions, FourierDiff modulates the generative prior's amplitude using a learnable adaptive brightness factor γ\gamma, which is optimized at each reverse diffusion step using a non-reference brightness control objective LbriL_{\text{bri}}:

    Lbri=1R∑n=1R∣Itn−E∣L_{\text{bri}} = \frac{1}{R} \sum_{n=1}^R \left| I_t^n - E \right|

    where:

    • RR is the total number of non-overlapping local patches of size 16×1616 \times 16 pixels across the rectified sampled image x^0∣t\hat{\mathbf{x}}_{0|t}.
    • Itn∈[0,1]I_t^n \in [0, 1] represents the average intensity of local patch nn in x^0∣t\hat{\mathbf{x}}_{0|t}.
    • E∈[0,1]E \in [0, 1] is a target gray-level exposure constant in the RGB color space (set to 0.50.5 by default).

    Minimizing LbriL_{\text{bri}} optimizes γ\gamma in x^0∣t=IFFT(γAx0∣t+Ayt,Pyt)\hat{\mathbf{x}}_{0|t} = \text{IFFT}(\gamma A_{\mathbf{x}_{0|t}} + A_{\mathbf{y}_t}, P_{\mathbf{y}_t}), aligning the local brightness distribution with the specified target EE.

  4. Knowl 4 — Phase-Guided Spatial-Frequency Alternating Optimization

    model/method

    To remove motion blur from the phase guidance during diffusion sampling, FourierDiff employs a spatial-frequency alternating optimization strategy that iteratively refines the input image phase and blur kernel.

    1. Phase Autocorrelation: Given the phase spectrum PyP_{\mathbf{y}} of an image, the autocorrelation A(∣Py∣)\mathcal{A}(|P_{\mathbf{y}}|) of the reconstructed magnitude ∣Py∣|P_{\mathbf{y}}| is computed in the frequency domain as:

    A(∣Py∣)=IFFT(FFT(∣Py∣)⊙FFT(∣Py∣)‾)\mathcal{A}(|P_{\mathbf{y}}|) = \text{IFFT}\left(\text{FFT}(|P_{\mathbf{y}}|) \odot \overline{\text{FFT}(|P_{\mathbf{y}}|)}\right)

    where ⊙\odot represents element-wise multiplication and (⋅)‾\overline{(\cdot)} denotes complex conjugation. The blur kernel k\mathbf{k} is directly extracted from this autocorrelation function.

    1. Spatial Optimization: Using the estimated kernel k\mathbf{k} and degraded image y\mathbf{y}, a latent sharp image y^\hat{\mathbf{y}} is obtained by solving the regularized deblurring problem:

    min⁡y^,k∥k⊗y^−y∥22+λ1∥k∥22+λ2h(∇y^)\min_{\hat{\mathbf{y}}, \mathbf{k}} \|\mathbf{k} \otimes \hat{\mathbf{y}} - \mathbf{y}\|_2^2 + \lambda_1 \|\mathbf{k}\|_2^2 + \lambda_2 h(\nabla \hat{\mathbf{y}})

    where ⊗\otimes denotes 2D convolution, h(⋅)h(\cdot) is a truncated-quadratic gradient regularization term, and λ1,λ2\lambda_1, \lambda_2 are weighting parameters (set to λ1=2,λ2=0.005\lambda_1 = 2, \lambda_2 = 0.005).

    1. Kernel Update via Diffusion Feedback: As the sampled estimate x^0∣t\hat{\mathbf{x}}_{0|t} becomes sharper over reverse diffusion, its Fourier phase Px^0∣tP_{\hat{\mathbf{x}}_{0|t}} provides improved kernel estimates kt=A(∣Px^0∣t∣)\mathbf{k}_t = \mathcal{A}(|P_{\hat{\mathbf{x}}_{0|t}}|). The running blur kernel is updated via temporal moving average:

    k=(1−1t)k+1tkt\mathbf{k} = \left(1 - \frac{1}{t}\right)\mathbf{k} + \frac{1}{t}\mathbf{k}_t

    This optimization is executed periodically at interval steps NN during reverse sampling.

  5. Knowl 5 — FourierDiff Zero-Shot Sampling Algorithm

    algorithm

    FourierDiff executes reverse diffusion sampling while alternating with spatial-frequency phase optimization to generate a sharp, normally illuminated image from a low-light blurry input.

    Input: Degraded image y\mathbf{y}, total diffusion steps TT, alternating optimization interval step NN, regularization weights λ1,λ2\lambda_1, \lambda_2
    Output: Enhanced and deblurred image x0\mathbf{x}_0
    (Ay,Py)←FFT(y)(A_{\mathbf{y}}, P_{\mathbf{y}}) \leftarrow \text{FFT}(\mathbf{y})
    k←A(∣Py∣)\mathbf{k} \leftarrow \mathcal{A}(|P_{\mathbf{y}}|)
    y^←y\hat{\mathbf{y}} \leftarrow \mathbf{y}
    xT∼N(0,I)\mathbf{x}_T \sim \mathcal{N}(\mathbf{0}, \mathbf{I})
    for t=T,T−1,…,1t = T, T-1, \dots, 1 do
        x0∣t←1αˉt(xt−ϵθ(xt,t)1−αˉt)\mathbf{x}_{0|t} \leftarrow \frac{1}{\sqrt{\bar{\alpha}_t}}\left(\mathbf{x}_t - \boldsymbol{\epsilon}_\theta(\mathbf{x}_t, t)\sqrt{1 - \bar{\alpha}_t}\right)
        yt←y^\mathbf{y}_t \leftarrow \hat{\mathbf{y}}
        (Ax0∣t,Px0∣t)←FFT(x0∣t)(A_{\mathbf{x}_{0|t}}, P_{\mathbf{x}_{0|t}}) \leftarrow \text{FFT}(\mathbf{x}_{0|t})
        (Ayt,Pyt)←FFT(yt)(A_{\mathbf{y}_t}, P_{\mathbf{y}_t}) \leftarrow \text{FFT}(\mathbf{y}_t)
        x^0∣t←IFFT(γAx0∣t+Ayt,Pyt)\hat{\mathbf{x}}_{0|t} \leftarrow \text{IFFT}(\gamma A_{\mathbf{x}_{0|t}} + A_{\mathbf{y}_t}, P_{\mathbf{y}_t})
        Sample xt−1∼pθ(xt−1∣xt,x^0∣t)\mathbf{x}_{t-1} \sim p_\theta(\mathbf{x}_{t-1} \mid \mathbf{x}_t, \hat{\mathbf{x}}_{0|t})
        
        if t mod N==0t \bmod N == 0 then
            kt←A(∣Px^0∣t∣)\mathbf{k}_t \leftarrow \mathcal{A}(|P_{\hat{\mathbf{x}}_{0|t}}|)
            k←(1−1t)k+1tkt\mathbf{k} \leftarrow \left(1 - \frac{1}{t}\right)\mathbf{k} + \frac{1}{t}\mathbf{k}_t
            Solve min⁡y^,k∥k⊗y^−y∥22+λ1∥k∥22+λ2h(∇y^)\min_{\hat{\mathbf{y}}, \mathbf{k}} \|\mathbf{k} \otimes \hat{\mathbf{y}} - \mathbf{y}\|_2^2 + \lambda_1\|\mathbf{k}\|_2^2 + \lambda_2 h(\nabla \hat{\mathbf{y}})
        end if
    end for
    return x0\mathbf{x}_0

    The algorithm operates over T=1000T = 1000 diffusion steps with an optimization interval of N=200N = 200. The score network ϵθ\boldsymbol{\epsilon}_\theta is an unconditional diffusion model pre-trained on ImageNet at 256×256256 \times 256 resolution.

  6. Knowl 6 — Experimental Configuration and Benchmarks for Joint Low-Light Deblurring

    experimental setup

    FourierDiff is evaluated on two standard benchmarks for joint low-light enhancement and deblurring:

    • LOL-Blur Dataset: Contains 12,000 synthetic low-light blurry and normal-light sharp image pairs with varying illumination levels and motion blur patterns.
    • RealBlur Dataset: A real-world motion deblurring benchmark. Following standard protocols, a test split of 482 real-world night blurry images across diverse night scenes is evaluated.

    Implementation Setup:

    • Backbone: Unconditional 256×256256 \times 256 diffusion model pre-trained on ImageNet.
    • Hardware & Framework: PyTorch on a single NVIDIA GeForce RTX 3090 GPU.
    • Hyperparameters: Total diffusion steps T=1000T = 1000, alternating optimization interval N=200N = 200, default target exposure E=0.5E = 0.5, regularization parameters λ1=2\lambda_1 = 2, λ2=0.005\lambda_2 = 0.005.
    • Extreme Darkness Preprocessing: For inputs with extreme underexposure, a practical exposure correction (PEC) module with a small exposure parameter is applied as a warm-start to prevent total loss of phase content guidance.

    Evaluation Metrics:

    • Full-Reference: PSNR (dB) and SSIM (evaluated on LOL-Blur).
    • No-Reference: NIQE (lower is better), Perceptual Index PI (lower is better), BRISQUE (lower is better), and MUSIQ (higher is better).
  7. Knowl 7 — Quantitative Evaluation on the LOL-Blur Dataset

    data/table

    On the synthetic LOL-Blur dataset, FourierDiff is evaluated against cascaded enhancement-then-deblurring pipelines, deblurring-then-enhancement pipelines, and the supervised joint baseline LEDNet. Cascaded baselines include combinations of Zero-DCE++, RetinexDIP, GDP, Chen et al., W-DIP, and GRL.

    Method NIQE ↓\downarrow PI ↓\downarrow BRISQUE ↓\downarrow MUSIQ ↑\uparrow PSNR ↑\uparrow SSIM ↑\uparrow
    Enhancement →\rightarrow Deblurring
    Zero-DCE++ →\rightarrow GRL 4.27 5.05 42.21 53.36 18.45 0.59
    RetinexDIP →\rightarrow GRL 4.59 5.38 47.45 49.61 13.65 0.55
    GDP →\rightarrow GRL 4.31 4.81 41.03 56.42 17.72 0.66
    Deblurring →\rightarrow Enhancement
    Chen →\rightarrow Zero-DCE++ 4.76 4.97 47.10 51.64 17.43 0.51
    Chen →\rightarrow GDP 4.87 4.70 49.83 55.19 16.52 0.56
    W-DIP →\rightarrow Zero-DCE++ 4.82 4.32 37.58 47.04 16.52 0.42
    W-DIP →\rightarrow GDP 5.03 4.10 35.60 50.10 15.69 0.46
    GRL →\rightarrow Zero-DCE++ 4.28 5.13 44.36 55.96 18.90 0.64
    GRL →\rightarrow GDP 4.32 4.90 43.25 58.92 18.16 0.70
    Joint Restoration
    LEDNet (Supervised) 3.99 5.07 42.59 59.64 25.74 0.85
    FourierDiff (Ours, Zero-Shot) 3.80 3.88 33.13 62.46 20.53 0.71

    FourierDiff achieves the best scores across all perceptual no-reference metrics (NIQE 3.80, PI 3.88, BRISQUE 33.13, MUSIQ 62.46), outperforming both cascaded zero-shot/unsupervised approaches and the fully supervised LEDNet in perceptual realness.

  8. Knowl 8 — Quantitative Evaluation on Real-World Night Blurry Images (RealBlur)

    data/table

    On the 482 real-world night blurry test images of the RealBlur dataset, FourierDiff is evaluated without paired ground truth using no-reference perceptual image quality metrics.

    Method NIQE ↓\downarrow PI ↓\downarrow BRISQUE ↓\downarrow MUSIQ ↑\uparrow
    Enhancement →\rightarrow Deblurring
    Zero-DCE++ →\rightarrow GRL 3.33 4.55 30.46 42.88
    RetinexDIP →\rightarrow GRL 3.35 4.29 30.90 44.79
    GDP →\rightarrow GRL 3.26 4.54 28.96 39.35
    Deblurring →\rightarrow Enhancement
    Chen →\rightarrow Zero-DCE++ 4.88 4.88 45.89 49.96
    Chen →\rightarrow GDP 4.67 4.95 45.60 47.19
    W-DIP →\rightarrow Zero-DCE++ 4.23 4.31 35.60 41.43
    W-DIP →\rightarrow GDP 4.06 4.20 33.00 38.68
    GRL →\rightarrow Zero-DCE++ 3.70 4.71 34.80 45.50
    GRL →\rightarrow GDP 3.58 4.61 33.10 43.22
    Joint Restoration
    LEDNet 3.72 5.03 42.31 49.45
    FourierDiff (Ours) 3.25 3.36 26.39 52.24

    FourierDiff achieves top performance across all four metrics (NIQE 3.25, PI 3.36, BRISQUE 26.39, MUSIQ 52.24). Supervised methods such as LEDNet experience performance degradation when applied to unseen real-world scenes, whereas FourierDiff generalizes without degradation modeling assumptions.

  9. Knowl 9 — Ablation of Spatial-Frequency Alternating Optimization Interval Step N

    data/table

    The impact of the spatial-frequency alternating optimization (SFA) interval step NN is evaluated on the RealBlur dataset across five settings. The baseline w/o SFA denotes refining the input image phase once prior to diffusion sampling and holding it constant throughout all T=1000T=1000 steps.

    Setting w/o SFA N=500N = 500 N=200N = 200 N=100N = 100 N=1N = 1
    NIQE ↓\downarrow 3.62 3.37 3.25 3.26 3.19
    PI ↓\downarrow 3.87 3.61 3.36 3.39 3.34
    BRISQUE ↓\downarrow 31.65 29.72 26.39 26.12 25.14
    MUSIQ ↑\uparrow 48.81 49.87 52.24 52.41 51.11

    Decreasing the interval from N=500N=500 to N=200N=200 yields substantial perceptual gains (e.g., BRISQUE drops from 29.7229.72 to 26.3926.39, MUSIQ rises from 49.8749.87 to 52.2452.24). Decreasing the interval further from N=200N=200 to N=100N=100 or N=1N=1 produces marginal gains at significant computational expense. Thus, N=200N=200 achieves the optimal trade-off between restoration quality and sampling efficiency.

  10. Knowl 10 — Limitations of FourierDiff

    limitation

    FourierDiff exhibits two principal limitations:

    1. Severe Content Degradation in Extreme Darkness: In scenes with near-total absence of light, the phase spectrum of the input image contains insufficient structural information, leading to degraded content guidance and sub-optimal restoration.
    2. Inference Latency: Because the framework relies on iterative sampling over T=1000T = 1000 diffusion steps coupled with periodic regularized optimization, its inference latency prevents real-time processing.

Coverage note — None was omitted; all key methodological contributions, equations, algorithmic procedures, experimental setups, comparative benchmark tables, ablation studies, and stated limitations are fully covered.

References

  1. 1.Yochai Blau and Tomer Michaeli. The perception-distortion tradeoff. In CVPR, pages 6228–6237, 2018. 6
  2. 2.Gustav Bredell, Ertunc Erdil, Bruno Weber, and Ender Konukoglu. Wiener guided DIP for unsupervised blind image deconvolution. In WACV, pages 3046–3055, 2023. 3, 6, 7
  3. 3.Yuanhao Cai, Hao Bian, Jing Lin, Haoqian Wang, Radu Timofte, and Yulun Zhang. Retinexformer: One-stage retinex-based transformer for low-light image enhancement. In ICCV, pages 12504–12513, 2023. 1, 3
  4. 4.Chen Chen, Qifeng Chen, Jia Xu, and Vladlen Koltun. Learning to see in the dark. In CVPR, pages 3291–3300, 2018. 3
  5. 5.Liangyu Chen, Xin Lu, Jie Zhang, Xiaojie Chu, and Chengpeng Chen. HINet: Half instance normalization network for image restoration. In CVPRW, pages 182–192, 2021. 3
  6. 6.Liang Chen, Jiawei Zhang, Songnan Lin, Faming Fang, and Jimmy S. Ren. Blind deblurring for saturated images. In CVPR, pages 6308–6316, 2021. 3, 6, 7
  7. 7.Liang Chen, Jiawei Zhang, Jinshan Pan, Songnan Lin, Faming Fang, and Jimmy S. Ren. Learning a non-blind deblurring network for night blurry images. In CVPR, pages 10542–10550, 2021. 3, 8
  8. 8.Liang Chen, Jiawei Zhang, Zhenhua Li, Yunxuan Wei, Faming Fang, Jimmy Ren, and Jinshan Pan. Deep richardson–lucy deconvolution for low-light image deblurring. IJCV, pages 1–18, 2023. 3
  9. 9.Rui Chen, Jiajun Chen, Zixi Liang, Huaien Gao, and Shan Lin. Darklight networks for action recognition in the dark. In CVPRW, pages 846–852, 2021. 1
  10. 10.Zhi Chen, Zijun Fan, Yongjie Li, Huaien Gao, and Shan Lin. Z-domain entropy adaptable flex for semi-supervised action recognition in the dark. In CVPRW, pages 4258–4265, 2022. 1
  11. 11.Sung-Jin Cho, Seo-Won Ji, Jun-Pyo Hong, Seung-Won Jung, and Sung-Jea Ko. Rethinking coarse-to-fine approach in single image deblurring. In ICCV, pages 4621–4630, 2021. 1, 3
  12. 12.Jooyoung Choi, Sungwon Kim, Yonghyun Jeong, Youngjune Gwon, and Sungroh Yoon. ILVR: conditioning method for denoising diffusion probabilistic models. In ICCV, pages 14347–14356, 2021. 3
  13. 13.Hyungjin Chung, Jeongsol Kim, Michael Thompson McCann, Marc Louis Klasky, and Jong Chul Ye. Diffusion posterior sampling for general noisy inverse problems. In ICLR, 2023. 3
  14. 14.Ziteng Cui, Guo-Jun Qi, Lin Gu, Shaodi You, Zenghui Zhang, and Tatsuya Harada. Multitask AET with orthogonal tangent regularity for dark object detection. In ICCV, pages 2533–2542, 2021. 1
  15. 15.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. ImageNet: A large-scale hierarchical image database. In CVPR, pages 248–255, 2009. 7
  16. 16.Prafulla Dhariwal and Alexander Quinn Nichol. Diffusion models beat gans on image synthesis. In NeurIPS, pages 8780–8794, 2021. 7
  17. 17.Ben Fei, Zhaoyang Lyu, Liang Pan, Junzhe Zhang, Weidong Yang, Tianyue Luo, Bo Zhang, and Bo Dai. Generative diffusion prior for unified image restoration and enhancement. In CVPR, pages 9935–9946, 2023. 1, 2, 3, 6, 7
  18. 18.Dong Gong, Jie Yang, Lingqiao Liu, Yanning Zhang, Ian D. Reid, Chunhua Shen, Anton van den Hengel, and Qinfeng Shi. From motion blur to motion flow: A deep learning solution for removing heterogeneous motion blur. In CVPR, pages 3806–3815, 2017. 3
  19. 19.Chunle Guo, Chongyi Li, Jichang Guo, Chen Change Loy, Junhui Hou, Sam Kwong, and Runmin Cong. Zero-reference deep curve estimation for low-light image enhancement. In CVPR, pages 1777–1786, 2020. 1, 3
  20. 20.Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. In NeurIPS, pages 6840–6851, 2020. 3, 4
  21. 21.Zhe Hu, Sunghyun Cho, Jue Wang, and Ming-Hsuan Yang. Deblurring low-light images with light streaks. In CVPR, pages 3382–3389, 2014. 3
  22. 22.Hai Jiang, Ao Luo, Haoqiang Fan, Songchen Han, and Shuaicheng Liu. Low-light image enhancement with wavelet-based diffusion models. ACM TOG, 42(6):1–14, 2023. 3
  23. 23.Bahjat Kawar, Michael Elad, Stefano Ermon, and Jiaming Song. Denoising diffusion restoration models. In NeurIPS, 2022. 2, 3
  24. 24.Junjie Ke, Qifei Wang, Yilin Wang, Peyman Milanfar, and Feng Yang. MUSIQ: multi-scale image quality transformer. In ICCV, pages 5128–5137, 2021. 6
  25. 25.Orest Kupyn, Volodymyr Budzan, Mykola Mykhailych, Dmytro Mishkin, and Jiri Matas. DeblurGAN: Blind motion deblurring using conditional adversarial networks. In CVPR, pages 8183–8192, 2018. 3
  26. 26.Orest Kupyn, Tetiana Martyniuk, Junru Wu, and Zhangyang Wang. DeblurGAN-v2: Deblurring (orders-of-magnitude) faster and better. In ICCV, pages 8877–8886, 2019. 1, 3
  27. 27.Edwin H Land. The retinex theory of color vision. Sci. Amer., 237(6):108–129, 1977. 3
  28. 28.Chongyi Li, Chunle Guo, Linghao Han, Jun Jiang, Ming-Ming Cheng, Jinwei Gu, and Chen Change Loy. Low-light image and video enhancement using deep learning: A survey. IEEE TPAMI, 44(12):9396–9416, 2022. 3
  29. 29.Chongyi Li, Chunle Guo, and Chen Change Loy. Learning to enhance low-light image via zero-reference deep curve estimation. IEEE TPAMI, 44(8):4225–4238, 2022. 3, 6, 7
  30. 30.Chongyi Li, Chun-Le Guo, Man Zhou, Zhexin Liang, Shangchen Zhou, Ruicheng Feng, and Chen Change Loy. Embedding fourier for ultra-high-definition low-light image enhancement. In ICLR, 2023. 1, 3
  31. 31.Haoying Li, Ziran Zhang, Tingting Jiang, Peng Luo, Huajun Feng, and Zhihai Xu. Real-world deep local motion deblurring. In AAAI, pages 1314–1322, 2023. 3
  32. 32.Xin Li, Yulin Ren, Xin Jin, Cuiling Lan, Xingrui Wang, Wenjun Zeng, Xinchao Wang, and Zhibo Chen. Diffusion models for image restoration and enhancement–a comprehensive survey. arXiv preprint arXiv:2308.09388, 2023. 3
  33. 33.Yawei Li, Yuchen Fan, Xiaoyu Xiang, Denis Demandolx, Rakesh Ranjan, Radu Timofte, and Luc Van Gool. Efficient and explicit modelling of image hierarchies for image restoration. In CVPR, pages 18278–18289, 2023. 1, 3, 6, 7
  34. 34.Jiaying Liu, Dejia Xu, Wenhan Yang, Minhao Fan, and Haofeng Huang. Benchmarking low-light image enhancement and beyond. IJCV, 129(4):1153–1184, 2021. 1, 3
  35. 35.Risheng Liu, Long Ma, Jiaao Zhang, Xin Fan, and Zhongxuan Luo. Retinex-inspired unrolling with cooperative prior architecture search for low-light image enhancement. In CVPR, pages 10561–10570, 2021. 3
  36. 36.Andreas Lugmayr, Martin Danelljan, Andres Romero, Fisher Yu, Radu Timofte, and Luc Van Gool. Repaint: Inpainting using denoising diffusion probabilistic models. In CVPR, pages 11451–11461, 2022. 3
  37. 37.Ziwei Luo, Fredrik K Gustafsson, Zheng Zhao, Jens Sjölund, and Thomas B Schön. Refusion: Enabling large-size realistic image restoration with latent-space diffusion models. In CVPR, pages 1680–1691, 2023. 3
  38. 38.Xiaoqian Lv, Shengping Zhang, Qinglin Liu, Haozhe Xie, Bineng Zhong, and Huiyu Zhou. BacklitNet: A dataset and network for backlit image enhancement. CVIU, 218:103403, 2022. 3
  39. 39.Long Ma, Tengyu Ma, Risheng Liu, Xin Fan, and Zhongxuan Luo. Toward fast, flexible, and robust low-light image enhancement. In CVPR, pages 5627–5636, 2022. 1, 3, 6, 7
  40. 40.Long Ma, Tianjiao Ma, Xinwei Xue, Xin Fan, Zhongxuan Luo, and Risheng Liu. Practical exposure correction: Great truths are always simple. arXiv preprint arXiv:2212.14245, 2022. 7
  41. 41.Xintian Mao, Yiming Liu, Fengze Liu, Qingli Li, Wei Shen, and Yan Wang. Intriguing findings of frequency selection for image deblurring. In AAAI, pages 1905–1913, 2023. 5
  42. 42.Tom Mertens, Jan Kautz, and Frank Van Reeth. Exposure fusion. In Proc. Pacific Conf. Comput. Graph. Appl., pages 382–390, 2007. 5
  43. 43.Anish Mittal, Anush Krishna Moorthy, and Alan Conrad Bovik. No-reference image quality assessment in the spatial domain. IEEE TIP, 21(12):4695–4708, 2012. 6
  44. 44.Anish Mittal, Rajiv Soundararajan, and Alan C. Bovik. Making a “completely blind” image quality analyzer. IEEE Sign. Process. Letters, 20(3):209–212, 2013. 6
  45. 45.Seungjun Nah, Tae Hyun Kim, and Kyoung Mu Lee. Deep multi-scale convolutional neural network for dynamic scene deblurring. In CVPR, pages 257–265, 2017. 3
  46. 46.Liyuan Pan, Richard I. Hartley, Miaomiao Liu, and Yuchao Dai. Phase-only image based kernel estimation for single image blind deblurring. In CVPR, pages 6034–6043, 2019. 5
  47. 47.Dongwei Ren, Kai Zhang, Qilong Wang, Qinghua Hu, and Wangmeng Zuo. Neural blind deconvolution using deep priors. In CVPR, pages 3338–3347, 2020. 3
  48. 48.Jaesung Rim, Haeyun Lee, Jucheol Won, and Sunghyun Cho. Real-world blur dataset for learning and benchmarking deblurring algorithms. In ECCV, pages 184–201, 2020. 6
  49. 49.Chitwan Saharia, Jonathan Ho, William Chan, Tim Salimans, David J. Fleet, and Mohammad Norouzi. Image super-resolution via iterative refinement. IEEE TPAMI, 45(4): 4713–4726, 2023. 3
  50. 50.Christian J. Schuler, Michael Hirsch, Stefan Harmeling, and Bernhard Schölkopf. Learning to deblur. IEEE TPAMI, 38 (7):1439–1451, 2016. 3
  51. 51.Jascha Sohl-Dickstein, Eric A. Weiss, Niru Maheswaranathan, and Surya Ganguli. Deep unsupervised learning using nonequilibrium thermodynamics. In ICML, pages 2256–2265, 2015. 3
  52. 52.Jiaming Song, Chenlin Meng, and Stefano Ermon. Denoising diffusion implicit models. In ICLR, 2021. 3
  53. 53.Jian Sun, Wenfei Cao, Zongben Xu, and Jean Ponce. Learning a convolutional neural network for non-uniform motion blur removal. In CVPR, pages 769–777, 2015. 3
  54. 54.Xin Tao, Hongyun Gao, Xiaoyong Shen, Jue Wang, and Jiaya Jia. Scale-recurrent network for deep image deblurring. In CVPR, pages 8174–8182, 2018. 3
  55. 55.Dmitry Ulyanov, Andrea Vedaldi, and Victor S. Lempitsky. Deep image prior. In CVPR, pages 9446–9454, 2018. 3
  56. 56.Ruixing Wang, Qing Zhang, Chi-Wing Fu, Xiaoyong Shen, Wei-Shi Zheng, and Jiaya Jia. Underexposed photo enhancement using deep illumination estimation. In CVPR, pages 6849–6857, 2019. 3
  57. 57.Tao Wang, Kaihao Zhang, Tianrun Shen, Wenhan Luo, Björn Stenger, and Tong Lu. Ultra-high-definition low-light image enhancement: A benchmark and transformer-based method. In AAAI, pages 2654–2662, 2023. 3
  58. 58.Yufei Wang, Renjie Wan, Wenhan Yang, Haoliang Li, Lap-Pui Chau, and Alex C. Kot. Low-light image enhancement with normalizing flow. In AAAI, pages 2604–2612, 2022. 3
  59. 59.Yinhuai Wang, Jiwen Yu, and Jian Zhang. Zero-shot image restoration using denoising diffusion null-space model. In ICLR, 2023. 2, 3, 4, 5
  60. 60.Yufei Wang, Yi Yu, Wenhan Yang, Lanqing Guo, Lap-Pui Chau, Alex C Kot, and Bihan Wen. Exposurediffusion: Learning to expose for low-light image enhancement. In ICCV, pages 12438–12448, 2023. 3
  61. 61.Zhou Wang, Alan C. Bovik, Hamid R. Sheikh, and Eero P. Simoncelli. Image quality assessment: from error visibility to structural similarity. IEEE TIP, 13(4):600–612, 2004. 7
  62. 62.Zhendong Wang, Xiaodong Cun, Jianmin Bao, Wengang Zhou, Jianzhuang Liu, and Houqiang Li. Uformer: A general u-shaped transformer for image restoration. In CVPR, pages 17662–17672, 2022. 3
  63. 63.Chen Wei, Wenjing Wang, Wenhan Yang, and Jiaying Liu. Deep retinex decomposition for low-light enhancement. In BMVC, page 155, 2018. 3
  64. 64.Wenhui Wu, Jian Weng, Pingping Zhang, Xu Wang, Wenhan Yang, and Jianmin Jiang. Uretinex-net: Retinex-based deep unfolding network for low-light image enhancement. In CVPR, pages 5891–5900, 2022. 3
  65. 65.Li Xu, Shicheng Zheng, and Jiaya Jia. Unnatural L0 sparse representation for natural image deblurring. In CVPR, pages 1107–1114, 2013. 6
  66. 66.Xiaogang Xu, Ruixing Wang, Chi-Wing Fu, and Jiaya Jia. Snr-aware low-light image enhancement. In CVPR, pages 17693–17703, 2022. 3
  67. 67.Xiaogang Xu, Ruixing Wang, and Jiangbo Lu. Low-light image enhancement via structure modeling and guidance. In CVPR, pages 9893–9903, 2023. 3
  68. 68.Shuzhou Yang, Moxuan Ding, Yanmin Wu, Zihan Li, and Jian Zhang. Implicit neural representation for cooperative low-light image enhancement. In ICCV, pages 12918–12927, 2023. 3
  69. 69.Xunpeng Yi, Han Xu, Hao Zhang, Linfeng Tang, and Jiayi Ma. Diff-retinex: Rethinking low-light image enhancement with a generative diffusion model. In ICCV, pages 12302–12311, 2023. 3
  70. 70.Syed Waqas Zamir, Aditya Arora, Salman H. Khan, Munawar Hayat, Fahad Shahbaz Khan, Ming-Hsuan Yang, and Ling Shao. Multi-stage progressive image restoration. In CVPR, pages 14821–14831, 2021. 3
  71. 71.Syed Waqas Zamir, Aditya Arora, Salman Khan, Munawar Hayat, Fahad Shahbaz Khan, and Ming-Hsuan Yang. Restormer: Efficient transformer for high-resolution image restoration. In CVPR, pages 5718–5729, 2022. 1, 3
  72. 72.Yonghua Zhang, Xiaojie Guo, Jiayi Ma, Wei Liu, and Jiawan Zhang. Beyond brightening low-light images. IJCV, 129(4): 1013–1037, 2021. 3
  73. 73.Zhao Zhang, Huan Zheng, Richang Hong, Mingliang Xu, Shuicheng Yan, and Meng Wang. Deep color consistent network for low-light image enhancement. In CVPR, pages 1889–1898, 2022. 3
  74. 74.Zhihong Zhang, Yuxiao Cheng, Jinli Suo, Liheng Bian, and Qionghai Dai. INFWIDE: image and feature space wiener deconvolution network for non-blind image deblurring in low-light conditions. IEEE TIP, 32:1390–1402, 2023. 3
  75. 75.Zunjin Zhao, Bangshu Xiong, Lei Wang, Qiaofeng Ou, Lei Yu, and Fa Kuang. RetinexDIP: A unified deep framework for low-light image enhancement. IEEE TCSVT, 32 (3):1076–1088, 2022. 3, 7
  76. 76.Dewei Zhou, Zongxin Yang, and Yi Yang. Pyramid diffusion models for low-light image enhancement. In IJCAI, pages 1795–1803, 2023. 3
  77. 77.Shangchen Zhou, Chongyi Li, and Chen Change Loy. LED-Net: Joint low-light enhancement and deblurring in the dark. In ECCV, pages 573–589, 2022. 1, 2, 3, 6, 7

Citation

MLA
Kim, J., and T.-K. Kim. “Arbitrary-Scale Image Generation and Upsampling Using Latent Diffusion Model and Implicit Neural Decoder”. arXiv, 2024, http://arxiv.org/abs/2403.10255v1.
APA
Kim, J., & Kim, T.-K. (2024). Arbitrary-Scale Image Generation and Upsampling using Latent Diffusion Model and Implicit Neural Decoder. arXiv. http://arxiv.org/abs/2403.10255v1
Chicago
Kim, J., and T.-K. Kim. 2024. “Arbitrary-Scale Image Generation and Upsampling Using Latent Diffusion Model and Implicit Neural Decoder”. arXiv. http://arxiv.org/abs/2403.10255v1.
Harvard
Kim, J. and Kim, T.-K. (2024) “Arbitrary-Scale Image Generation and Upsampling using Latent Diffusion Model and Implicit Neural Decoder”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2403.10255v1.
Vancouver
1. Kim J, Kim T-K (2024) Arbitrary-Scale Image Generation and Upsampling using Latent Diffusion Model and Implicit Neural Decoder. arXiv

BibTeX

@article{kim2024arbitrary,
  title = {Arbitrary-Scale Image Generation and Upsampling using Latent Diffusion Model and Implicit Neural Decoder},
  author = {Kim, Jinseok and Kim, Tae-Kyun},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2403.10255v1},
  eprint = {2403.10255}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE