Wavelet-based Fourier Information Interaction with Frequency Diffusion Adjustment for Underwater Image Restoration

Chen ZhaoWeiling CaiChenyu DongChengwei Hu

article2024CVPR237 citations

Proposes WF-Diff, a two-stage underwater image restoration framework that integrates wavelet-Fourier frequency interactions with a frequency residual diffusion adjustment module to effectively correct color distortion and recover fine textures.

Listen

Underwater image restoration is critical for marine robotics, subsea monitoring, and underwater object tracking. Light refraction, absorption, and scattering in aquatic environments cause severe distortions, such as heavy color casts, loss of contrast, and blurred structural details. While deep learning methods have improved restoration quality, existing models primarily operate in standard pixel space. As a result, they fail to leverage frequency-based representations and struggle to recover fine textures without introducing unwanted artifacts.

The article demonstrates a two-stage restoration framework, named WF-Diff, that combines frequency decomposition with targeted diffusion refinement. The objective is to systematically separate color correction from fine detail enhancement across distinct frequency bands, maximizing image clarity and structural precision.

The framework first decomposes underwater images into low-frequency and high-frequency components. A preliminary restoration network applies spatial and frequency fusion blocks to correct low-frequency color components while using transformer blocks to rebuild high-frequency structures. A cross-frequency conditioner shares contextual cues between these paths. In the second stage, a plug-and-play residual diffusion module refines the remaining high- and low-frequency errors against ground-truth targets. The authors validated the method against eight leading restoration models across three benchmark datasets totaling thousands of real-world test and reference images.

The evaluation produced several notable results. First, the proposed framework achieved state-of-the-art restoration quality across the primary benchmark datasets, reaching a peak signal-to-noise ratio of 23.86 on the UIEBD dataset compared to 21.88 for the top existing diffusion baseline. Second, perceptual distortion and distributional error metrics improved significantly, lowering the Fréchet Inception Distance score on UIEBD to 27.85 from the prior baseline of 31.07. Third, ablation testing confirmed that learning residual frequency distributions prevented the visual artifacts and hallucinations common in standard pixel-level diffusion models. Finally, the framework generalized effectively across unreferenced test datasets, yielding superior non-reference quality scores.

These findings indicate that splitting underwater image enhancement into frequency-specific preliminary restoration and residual diffusion substantially enhances visual fidelity. For operations reliant on subsea computer vision, clearer imagery reduces operational risk and improves detection accuracy in turbid conditions. However, the use of two diffusion models increases computational demand: the full framework requires approximately 0.28 seconds of inference time per image, compared to 0.01 to 0.14 seconds for competing alternatives.

Organizations developing subsea vision systems should consider adopting frequency-guided residual architectures for post-processing or high-precision inspection pipelines. Before deploying this architecture in real-time edge environments—such as autonomous underwater vehicles—engineering teams should prioritize reducing the diffusion sampling steps or applying model compression techniques to accelerate inference speed.

  • Paper: An Underwater Image Enhancement Benchmark Dataset and Beyond, Chongyi Li et al. (2019). This paper establishes the UIEBD dataset and baseline enhancement benchmarks that the source directly uses to train and validate its frequency-based restoration framework.
  • Paper: Denoising Diffusion Probabilistic Models, Jonathan Ho et al. (2020). This foundational paper establishes the formulation and training principles of denoising diffusion probabilistic models, which underpin the residual diffusion refinement module utilized in the source.
  • Paper: Fast Underwater Image Enhancement for Improved Visual Perception, Md Jahidul Islam et al. (2019). This work introduces fast, deep learning-based underwater image enhancement and the EUVP dataset, providing foundational context on handling aquatic optical degradation and efficiency trade-offs.
  • Paper: Restormer: Efficient Transformer for High-Resolution Image Restoration, Syed Waqas Zamir et al. (2022). This paper presents efficient transformer blocks for multi-scale image restoration, informing the high-frequency transformer processing components used in the source's preliminary stage.
  • Paper: Denoising Diffusion Restoration Models, Bahjat Kawar et al. (2022). This study demonstrates how diffusion models can be guided through spectral and frequency decompositions to solve image restoration tasks without severe artifacts.
  • Paper: Diffusion Posterior Sampling for General Noisy Inverse Problems, Hyungjin Chung et al. (2022). This work establishes diffusion posterior sampling for noisy inverse imaging problems, providing key theoretical groundwork for guided diffusion-based restoration.
  • Paper: Multi-Stage Progressive Image Restoration, Syed Waqas Zamir et al. (2021). This paper introduces multi-stage progressive restoration architectures with cross-stage feature fusion, inspiring the sequential restoration and refinement structure of WF-Diff.
  • Paper: Dual-Domain Attention for Image Deblurring, Yuning Cui et al. (2023). This study demonstrates decoupling restoration across spatial and frequency domains using attention mechanisms, directly preceding the dual-frequency conditioning design of the source.
Cover for Wavelet-based Fourier Information Interaction with Frequency Diffusion Adjustment for Underwater Image Restoration

Abstract

Underwater images are subject to intricate and diverse degradation, inevitably affecting the effectiveness of underwater visual tasks. However, most approaches primarily operate in the raw pixel space of images, which limits the exploration of the frequency characteristics of underwater images, leading to an inadequate utilization of deep models' representational capabilities in producing high-quality images. In this paper, we introduce a novel Underwater Image Enhancement (UIE) framework, named WF-Diff, designed to fully leverage the characteristics of frequency domain information and diffusion models. WF-Diff consists of two detachable networks: Wavelet-based Fourier information interaction network (WFI2-net) and Frequency Residual Diffusion Adjustment Module (FR-DAM). With our full exploration of the frequency domain information, WFI2-net aims to achieve preliminary enhancement of frequency information in the wavelet space. Our proposed FRDAM can further refine the high- and low-frequency information of the initial enhanced images, which can be viewed as a plug-and-play universal module to adjust the detail of the underwater images. With the above techniques, our algorithm can show SOTA performance on real-world underwater image datasets, and achieves competitive performance in visual quality. The code is available at https://github.com/zhihefang/WF-Diff.

Table of Contents

  • 1. Introduction
  • 2. Related Works
  • 2.1. Underwater Image Enhancement
  • 2.2. Diffusion Model
  • 3. Methodology
  • 3.1. Overall Framework
  • 3.2. Discrete Wavelet and Fourier Transform
  • 3.3. Frequency Preliminary Enhancement
  • 3.4. Cross-Frequency Conditioner
  • 3.5. Frequency Diffusion Adjustment
  • 4. Experiments
  • 4.1. Setup
  • 4.2. Results and Comparisons
  • 4.3. Ablation Study
  • 5. Conclusion
  • References

Knowls

  1. Knowl 1 — WF-Diff Framework Architecture for Underwater Image Restoration

    model/method

    WF-Diff is a two-stage underwater image restoration framework designed to operate across wavelet and Fourier frequency domains while leveraging residual diffusion models. The architecture comprises two detachable networks:

    1. Wavelet-based Fourier Information Interaction Network (WFI2-net): The input image I∈RH×W×cI \in \mathbb{R}^{H \times W \times c} is first decomposed via 2D Discrete Wavelet Transform (DWT) into a low-frequency sub-band ILL∈RH2×W2×cI_{LL} \in \mathbb{R}^{\frac{H}{2} \times \frac{W}{2} \times c} and three high-frequency directional sub-bands {ILH,IHL,IHH}∈RH2×W2×c\{I_{LH}, I_{HL}, I_{HH}\} \in \mathbb{R}^{\frac{H}{2} \times \frac{W}{2} \times c}. WFI2-net uses a parallel encoder-decoder architecture where the high-frequency branch employs Wide Transformer Blocks (WTB) to capture global dependencies and sparse edges, while the low-frequency branch employs Spatial-Frequency Fusion Blocks (SFFB) to adjust Fourier amplitude components. The two branches exchange features at multiple scales via Cross-Frequency Conditioners (CFC), outputting initial restored sub-bands ILL′,ILH′,IHL′,IHH′I'_{LL}, I'_{LH}, I'_{HL}, I'_{HH}.

    2. Frequency Residual Diffusion Adjustment Module (FRDAM): Operating as a plug-and-play refinement module, FRDAM uses two specialized diffusion models—the Low-Frequency Diffusion Branch (LDFB) and the High-Frequency Diffusion Branch (HDFB)—to predict residual corrections I^LL\hat{I}_{LL} and I^(i)\hat{I}_{(i)} (i∈{LH,HL,HH}i \in \{LH, HL, HH\}) conditioned on the initial predictions Ii′I'_i.

    The final enhanced image IfinalI_{final} is generated by applying the Inverse Discrete Wavelet Transform (IDWT) to the sum of the initial predictions and the diffusion-generated residuals:

    Ifinal=IDWT(I(i)′+I^(i),ILL′+I^LL),i∈{LH,HL,HH}I_{final} = \text{IDWT}(I'_{(i)} + \hat{I}_{(i)}, I'_{LL} + \hat{I}_{LL}), \quad i \in \{LH, HL, HH\}

  2. Knowl 2 — Wavelet and Fourier Frequency Decomposition for Underwater Degradation Modeling

    model/method

    WF-Diff isolates distinct degradation modes of underwater images by decomposing inputs into wavelet sub-bands and analyzing their Fourier representations.

    Given an image I∈RH×W×cI \in \mathbb{R}^{H \times W \times c}, 2D Discrete Wavelet Transform (DWT) with Haar wavelets decomposes the input into four half-resolution sub-bands:

    ILL,{ILH,IHL,IHH}=DWT(I)I_{LL}, \{I_{LH}, I_{HL}, I_{HH}\} = \text{DWT}(I)

    where Haar 1D low-pass and high-pass filter vectors are L=12[1,1]TL = \frac{1}{\sqrt{2}}[1, 1]^T and H=12[1,−1]TH = \frac{1}{\sqrt{2}}[1, -1]^T. The low-frequency sub-band ILL∈RH2×W2×cI_{LL} \in \mathbb{R}^{\frac{H}{2} \times \frac{W}{2} \times c} encodes overall content and color information, while ILH,IHL,IHH∈RH2×W2×cI_{LH}, I_{HL}, I_{HH} \in \mathbb{R}^{\frac{H}{2} \times \frac{W}{2} \times c} encode sparse vertical, horizontal, and diagonal texture details.

    For any single-channel 2D spatial representation x∈RH×W×1x \in \mathbb{R}^{H \times W \times 1}, the 2D Discrete Fourier Transform F\mathcal{F} yields the complex frequency representation X(u,v)X(u, v):

    X(u,v)=F(x)(u,v)=1HW∑h=0H−1∑w=0W−1x(h,w)e−j2π(hHu+wWv)X(u, v) = \mathcal{F}(x)(u, v) = \frac{1}{\sqrt{HW}} \sum_{h=0}^{H-1} \sum_{w=0}^{W-1} x(h, w) e^{-j 2\pi \left(\frac{h}{H}u + \frac{w}{W}v\right)}

    which separates into an amplitude spectrum A(X(u,v))A(X(u, v)) and a phase spectrum P(X(u,v))\mathcal{P}(X(u, v)):

    A(X(u,v))=R2(X(u,v))+I2(X(u,v)),P(X(u,v))=arctan⁡(I(X(u,v))R(X(u,v)))A(X(u, v)) = \sqrt{R^2(X(u, v)) + I^2(X(u, v))}, \quad \mathcal{P}(X(u, v)) = \arctan\left(\frac{I(X(u, v))}{R(X(u, v))}\right)

    where R(⋅)R(\cdot) and I(⋅)I(\cdot) represent real and imaginary components. Fourier amplitude-phase swapping between degraded underwater images and clear ground truth demonstrates that underwater color cast is concentrated in the amplitude spectrum of the low-frequency wavelet component ILLI_{LL}, while structural and textural blurriness resides in the high-frequency sub-bands ILH,IHL,IHHI_{LH}, I_{HL}, I_{HH}.

  3. Knowl 3 — Wide Transformer Block for High-Frequency Wavelet Enhancement

    model/method

    The Wide Transformer Block (WTB) models long-range dependencies and local edge features within the high-frequency restoration branch of WFI2-net. Given high-frequency sub-bands {ILH,IHL,IHH}∈RH2×W2×c\{I_{LH}, I_{HL}, I_{HH}\} \in \mathbb{R}^{\frac{H}{2} \times \frac{W}{2} \times c}, convolutional projection yields feature embeddings Tin∈R3×H2×W2×CT_{in} \in \mathbb{R}^{3 \times \frac{H}{2} \times \frac{W}{2} \times C}.

    Each WTB stage processes the prior embedding Ti−1T_{i-1} through layer normalization (Norm\text{Norm}), followed by multi-scale depth-wise convolution WpW_p and 1×11 \times 1 point-wise convolution WdW_d:

    Q,K,V,L=Split(WdWp(Norm(Ti−1)))Q, K, V, L = \text{Split}(W_d W_p(\text{Norm}(T_{i-1})))

    where Split\text{Split} partitions channels into query QQ, key KK, value VV, and local representation LL. The intermediate embedding T^i\hat{T}_i is formed by self-attention (SA\text{SA}) on (Q,K,V)(Q, K, V) and channel attention (CA\text{CA}) on LL with a residual connection:

    T^i=SA(Q,K,V)+CA(L)+Ti−1\hat{T}_i = \text{SA}(Q, K, V) + \text{CA}(L) + T_{i-1}

    The final output TiT_i is computed via a feed-forward network (FFN\text{FFN}) with normalization and residual connection:

    Ti=FFN(Norm(T^i))+T^iT_i = \text{FFN}(\text{Norm}(\hat{T}_i)) + \hat{T}_i

  4. Knowl 4 — Spatial-Frequency Fusion Block for Low-Frequency Wavelet Enhancement

    model/method

    The Spatial-Frequency Fusion Block (SFFB) restores the low-frequency wavelet sub-band ILLI_{LL} by combining dual-domain spatial and Fourier representations.

    SFFB contains two processing units:

    1. Spatial Domain Unit (SDU): Uses multi-scale convolution kernels to expand the spatial receptive field, producing a spatial embedding FsF_s.
    2. Frequency Domain Unit (FDU): Applies Fast Fourier Transform (FFT) to FsF_s to extract its amplitude A(Fs)A(F_s) and phase P(Fs)\mathcal{P}(F_s). Both components pass through separate two-layer 1×11 \times 1 convolutions to produce adjusted amplitude A′(Fs)A'(F_s) and phase P′(Fs)\mathcal{P}'(F_s). The Inverse Fast Fourier Transform (IFFT) converts these modified spectra back to the spatial domain to yield frequency embedding FfF_f.

    The final output embedding FsfF_{sf} is obtained by additive dual-domain fusion:

    Fsf=Fs+FfF_{sf} = F_s + F_f

  5. Knowl 5 — Cross-Frequency Conditioner for Dual-Branch Information Interaction

    model/method

    The Cross-Frequency Conditioner (CFC) enables cross-frequency feature interaction between the high-frequency and low-frequency processing branches of WFI2-net.

    Given high-frequency feature embedding Tin∈R3×H2×W2×CT_{in} \in \mathbb{R}^{3 \times \frac{H}{2} \times \frac{W}{2} \times C} and low-frequency embedding Fin∈RH2×W2×CF_{in} \in \mathbb{R}^{\frac{H}{2} \times \frac{W}{2} \times C}, TinT_{in} is split into directional sub-bands TLH,THL,THH∈RH2×W2×CT_{LH}, T_{HL}, T_{HH} \in \mathbb{R}^{\frac{H}{2} \times \frac{W}{2} \times C} and summed into an aggregated high-frequency representation. Linear 1×11 \times 1 convolutions construct query QQ, key KK, and branch-specific values VTV_T and VFV_F:

    Q=Conv1×1(TLH+THL+THH)Q = \text{Conv}_{1 \times 1}(T_{LH} + T_{HL} + T_{HH}) K=Conv1×1(Fin)K = \text{Conv}_{1 \times 1}(F_{in}) VT=Conv1×1(TLH+THL+THH)V_T = \text{Conv}_{1 \times 1}(T_{LH} + T_{HL} + T_{HH}) VF=Conv1×1(Fin)V_F = \text{Conv}_{1 \times 1}(F_{in})

    The modulated output features ToutT_{out} (for the high-frequency branch) and FoutF_{out} (for the low-frequency branch) are calculated via cross-attention:

    Tout=R(Softmax(QKTdk)VT)T_{out} = \mathcal{R}\left(\text{Softmax}\left(\frac{QK^T}{\sqrt{d_k}}\right) V_T\right) Fout=Softmax(QKTdk)VFF_{out} = \text{Softmax}\left(\frac{QK^T}{\sqrt{d_k}}\right) V_F

    where dkd_k is the channel dimension of QQ, and R\mathcal{R} denotes a channel replication operation that maps the aggregated feature back to the 3-sub-band representation.

  6. Knowl 6 — Frequency Residual Diffusion Adjustment Module (FRDAM)

    model/method

    The Frequency Residual Diffusion Adjustment Module (FRDAM) refines initial frequency predictions Ii′I'_i in the wavelet domain by learning the residual distribution x0=Gi−Ii′x_0 = G_i - I'_i between ground-truth wavelet components GiG_i and initial estimates Ii′I'_i for i∈{LL,LH,HL,HH}i \in \{LL, LH, HL, HH\}. FRDAM contains a Low-Frequency Diffusion Branch (LDFB) and a High-Frequency Diffusion Branch (HDFB).

    Forward Diffusion: For each branch, a Markov chain progressively adds Gaussian noise over time step t∈[1,T]t \in [1, T] with variance schedule βt\beta_t:

    q(xt∣x0)=N(xt;αˉtx0,(1−αˉt)I)q(x_t | x_0) = \mathcal{N}\left(x_t; \sqrt{\bar{\alpha}_t} x_0, (1 - \bar{\alpha}_t) I\right)

    where αt=1−βt\alpha_t = 1 - \beta_t and αˉt=∏s=1tαs\bar{\alpha}_t = \prod_{s=1}^t \alpha_s.

    Reverse Diffusion and Loss: A U-Net ϵθ(xt,xc,t)\epsilon_\theta(x_t, x_c, t) is conditioned on the initial enhanced sub-band xc=Ii′x_c = I'_i to estimate the added noise ϵ\epsilon. The diffusion training objective is:

    Ldm(θ)=∥ϵ−ϵθ(αˉtx0+1−αˉtϵ,xc,t)∥\mathcal{L}_{dm}(\theta) = \left\| \epsilon - \epsilon_\theta\left(\sqrt{\bar{\alpha}_t} x_0 + \sqrt{1 - \bar{\alpha}_t} \epsilon, x_c, t\right) \right\|

    The conditional reverse transition is parameterized as:

    pθ(xt−1∣xt,xc)=N(xt−1;μθ(xt,xc,t),σt2I)p_\theta(x_{t-1} | x_t, x_c) = \mathcal{N}\left(x_{t-1}; \mu_\theta(x_t, x_c, t), \sigma_t^2 I\right) μθ(xt,xc,t)=1αt(xt−βt1−αˉtϵθ(xt,xc,t)),σt2=1−αˉt−11−αˉtβt\mu_\theta(x_t, x_c, t) = \frac{1}{\sqrt{\alpha_t}}\left(x_t - \frac{\beta_t}{\sqrt{1 - \bar{\alpha}_t}} \epsilon_\theta(x_t, x_c, t)\right), \quad \sigma_t^2 = \frac{1 - \bar{\alpha}_{t-1}}{1 - \bar{\alpha}_t} \beta_t

    During inference, starting from random noise ϵs(h)∈R3×H2×W2×3\epsilon_s^{(h)} \in \mathbb{R}^{3 \times \frac{H}{2} \times \frac{W}{2} \times 3} and ϵs(l)∈RH2×W2×3\epsilon_s^{(l)} \in \mathbb{R}^{\frac{H}{2} \times \frac{W}{2} \times 3}, HDFB and LDFB generate the residual corrections I^(i)=FHDFB(ϵs(h),I(i)′)\hat{I}_{(i)} = \mathcal{F}_{HDFB}(\epsilon_s^{(h)}, I'_{(i)}) and I^LL=FLDFB(ϵs(l),ILL′)\hat{I}_{LL} = \mathcal{F}_{LDFB}(\epsilon_s^{(l)}, I'_{LL}).

  7. Knowl 7 — Training Loss Formulation for WFI2-net

    equation

    The preliminary enhancement network WFI2-net is trained with a multi-objective loss combining high-frequency sub-band error, low-frequency Fourier amplitude error, and global adversarial reconstruction.

    Given ground-truth wavelet sub-bands GLL,GLH,GHL,GHHG_{LL}, G_{LH}, G_{HL}, G_{HH} and predicted sub-bands ILL′,ILH′,IHL′,IHH′I'_{LL}, I'_{LH}, I'_{HL}, I'_{HH}:

    1. High-frequency L2L_2 loss: Lh=∑i∈{LH,HL,HH}∥Ii′−Gi∥2\mathcal{L}_h = \sum_{i \in \{LH, HL, HH\}} \|I'_i - G_i\|_2

    2. Low-frequency Fourier amplitude L1L_1 loss: La=∥A(ILL′)−A(GLL)∥1\mathcal{L}_a = \|A(I'_{LL}) - A(G_{LL})\|_1 where A(⋅)A(\cdot) computes the 2D Fourier amplitude spectrum.

    3. Full-image reconstruction loss Lrec\mathcal{L}_{rec}: An adversarial loss based on Wasserstein GAN applied to the preliminary full image reconstruction.

  8. Knowl 8 — Quantitative Evaluation of WF-Diff on Underwater Datasets

    data/table

    WF-Diff was evaluated on the UIEBD test set (190 real-world images), the LSUI test set (504 images), and the non-reference U45 benchmark (45 images), and compared against eight state-of-the-art methods: UIEWD, UWCNN, UIEC2-Net\text{UIEC}^2\text{-Net}, Water-Net, SCNet, U-color, U-shape, and DM-water.

    Dataset / Metric UIEWD UWCNN UIEC2^2-Net Water-Net SCNet U-color U-shape DM-water WF-Diff (Ours)
    UIEBD
    FID ↓\downarrow 85.12 94.44 35.06 37.48 33.66 38.25 46.11 31.07 27.85
    LPIPS ↓\downarrow 0.3956 0.3525 0.2033 0.2116 0.2497 0.2337 0.2264 0.1436 0.1248
    PSNR (dB) ↑\uparrow 14.65 15.40 20.14 19.35 20.41 20.71 21.25 21.88 23.86
    SSIM ↑\uparrow 0.7265 0.7749 0.8215 0.8321 0.8235 0.8411 0.8453 0.8194 0.8730
    LSUI
    FID ↓\downarrow 98.49 100.5 34.51 38.90 158.99 45.06 28.56 27.91 26.75
    LPIPS ↓\downarrow 0.3962 0.3450 0.1432 0.1678 0.2830 0.1230 0.1028 0.1138 0.1096
    PSNR (dB) ↑\uparrow 15.43 18.24 20.86 19.73 22.63 22.91 24.16 27.65 27.26
    SSIM ↑\uparrow 0.7802 0.8465 0.8867 0.8226 0.9176 0.8902 0.9322 0.8867 0.9437
    U45
    UIQM ↑\uparrow 2.458 2.379 2.780 2.957 2.856 3.104 3.151 3.086 3.181
    UCIQE ↑\uparrow 0.583 0.567 0.591 0.601 0.594 0.586 0.592 0.634 0.619

    WF-Diff achieves the best scores on UIEBD across all four full-reference metrics (PSNR 23.86 dB, SSIM 0.8730, FID 27.85, LPIPS 0.1248). On LSUI, WF-Diff attains the highest SSIM (0.9437) and lowest FID (26.75). On the non-reference U45 benchmark, WF-Diff achieves the highest UIQM score of 3.181.

  9. Knowl 9 — Ablation Analysis of WFI2-net Components and FRDAM Diffusion Strategies

    data/table

    Ablation experiments on the UIEBD dataset demonstrate the contribution of individual modules, losses, and diffusion configurations.

    1. WFI2-net Modules and Losses: Complete WFI2-net achieves 21.87 dB PSNR / 0.8622 SSIM. Disabling Self-Attention (SA) in WTB drops PSNR to 20.94 dB (SSIM 0.8541); disabling Channel Attention (CA) drops it to 20.82 dB (SSIM 0.8473); disabling the Spatial Domain Unit (SDU) drops it to 21.11 dB (SSIM 0.8586); and disabling the Frequency Domain Unit (FDU) drops it to 20.23 dB (SSIM 0.8346). For loss terms, removing Lh\mathcal{L}_h drops PSNR to 20.46 dB, removing Lrec\mathcal{L}_{rec} drops PSNR to 20.65 dB, removing Fourier amplitude loss La\mathcal{L}_a drops PSNR to 19.81 dB, and removing the Cross-Frequency Conditioner (CFC) drops PSNR to 20.97 dB.

    2. FRDAM Configuration Comparison:

    Model Standard DM Residual DM (RDM) Wavelet D-L Wavelet D-H CFC PSNR (dB) ↑\uparrow
    A ✓ ×\times ×\times ×\times ×\times 20.86
    B ×\times ✓ ×\times ×\times ×\times 22.37
    C ×\times ✓ ✓ ×\times ×\times 22.32
    D ×\times ✓ ×\times ✓ ×\times 22.58
    E ×\times ✓ ✓ ✓ ×\times 23.44
    F (Full) ×\times ✓ ✓ ✓ ✓ 23.86

    Standard pixel diffusion (Model A) suffers from noise sampling diversity, resulting in 20.86 dB PSNR. Learning residual distributions in pixel space (Model B) improves PSNR to 22.37 dB. Performing residual diffusion independently on high- and low-frequency wavelet sub-bands (Model E) reaches 23.44 dB, and incorporating CFC cross-frequency interaction (Model F) achieves the best restoration performance of 23.86 dB.

  10. Knowl 10 — Inference Latency and Computational Complexity Limitation of WF-Diff

    limitation

    Because WF-Diff utilizes two separate diffusion models (LDFB and HDFB) to adjust low- and high-frequency residuals iteratively, its inference latency is higher than non-diffusion and single-model architectures.

    Metric SCNet U-shape DM-water WFI2-net (Stage 1 only) WF-Diff (Full)
    Inference Time (s) 0.0149 0.0353 0.1441 0.0739 0.2816
    FLOPs (G) 5.88 66.2 - 94.1 -

    WFI2-net alone operates in 0.0739 seconds with 94.1 GFLOPs (compared to 193.7 GFLOPs for WaterNet, 443.8 GFLOPs for Ucolor, and 66.2 GFLOPs for U-shape). When the full diffusion adjustment module FRDAM is executed with 10 implicit sampling steps, inference time increases to 0.2816 seconds per image, which is approximately double that of DM-water (0.1441s) and eight times slower than U-shape (0.0353s).

Coverage note — All core scientific contributions of the paper—including the WFI2-net architecture, the FRDAM residual diffusion module, the cross-frequency conditioner, loss formulations, benchmark evaluations, extensive ablation studies, and computational limitations—are fully documented. Qualitative visual comparison figures were omitted as their conclusions are fully captured by the quantitative data.

References

  1. 1.Derya Akkaynak and Tali Treibitz. Sea-thru: A method for removing water from underwater images. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2019, Long Beach, CA, USA, June 16-20, 2019, pages 1682–1691. Computer Vision Foundation / IEEE, 2019. 3
  2. 2.Derya Akkaynak, Tali Treibitz, Tom Shlesinger, Yossi Loya, Raz Tamir, and David Iluz. What is the space of attenuation coefficients in underwater computer vision? In 2017 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2017, Honolulu, HI, USA, July 21-26, 2017, pages 568–577, 2017. 1
  3. 3.Saeed Anwar, Chongyi Li, and Fatih Porikli. Deep underwater image enhancement. CoRR, abs/1807.03528, 2018. 6
  4. 4.John Yi-Wu Chiang and Ying-Ching Chen. Underwater image enhancement by wavelength compensation and dehazing. IEEE Trans. Image Process., 21(4):1756–1769, 2012. 3
  5. 5.Jooyoung Choi, Sungwon Kim, Yonghyun Jeong, Youngjune Gwon, and Sungroh Yoon. ILVR: conditioning method for denoising diffusion probabilistic models. In 2021 IEEE/CVF International Conference on Computer Vision, ICCV 2021, Montreal, QC, Canada, October 10-17, 2021, pages 14347–14356. IEEE, 2021. 2, 3
  6. 6.Karin de Langis and Junaed Sattar. Realtime multi-diver tracking and re-identification for underwater human-robot collaboration. In 2020 IEEE International Conference on Robotics and Automation, ICRA 2020, Paris, France, May 31 - August 31, 2020, pages 11140–11146, 2020. 1
  7. 7.Cameron Fabbri, Md Jahidul Islam, and Junaed Sattar. Enhancing underwater imagery using generative adversarial networks. In 2018 IEEE International Conference on Robotics and Automation, ICRA 2018, Brisbane, Australia, May 21-25, 2018, pages 7159–7165. IEEE, 2018. 1, 3
  8. 8.Zhenqi Fu, Xiaopeng Lin, Wu Wang, Yue Huang, and Xinghao Ding. Underwater image enhancement via learning water type desensitized representations. In IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP 2022, Virtual and Singapore, 23-27 May 2022, pages 2764–2768. IEEE, 2022. 6
  9. 9.Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. Gans trained by a two time-scale update rule converge to a local nash equilibrium. In Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017, December 4-9, 2017, Long Beach, CA, USA, pages 6626–6637, 2017. 6
  10. 10.Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual, 2020. 2, 3, 6
  11. 11.Huaibo Huang, Ran He, Zhenan Sun, and Tieniu Tan. Wavelet-srnet: A wavelet-based CNN for multi-scale face super resolution. In IEEE International Conference on Computer Vision, ICCV 2017, Venice, Italy, October 22-29, 2017, pages 1698–1706, 2017. 4
  12. 12.Jie Huang, Yajing Liu, Feng Zhao, Keyu Yan, Jinghao Zhang, Yukun Huang, Man Zhou, and Zhiwei Xiong. Deep fourier-based exposure correction network with spatialfrequency interaction. pages 163–180, 2022. 2
  13. 13.Shirui Huang, Keyan Wang, Huan Liu, Jun Chen, and Yunsong Li. Contrastive semi-supervised learning for underwater image restoration via reliable bank. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2023, Vancouver, BC, Canada, June 17-24, 2023, pages 18145–18155. IEEE, 2023. 3
  14. 14.Md Jahidul Islam, Youya Xia, and Junaed Sattar. Fast underwater image enhancement for improved visual perception. IEEE Robotics Autom. Lett., 5(2):3227–3234, 2020. 3
  15. 15.Paulo Drews Jr., Erickson Rangel do Nascimento, F. Moraes, Silvia S. C. Botelho, and Mario F. M. Campos. Transmission estimation in underwater single images. In 2013 IEEE International Conference on Computer Vision Workshops, ICCV Workshops 2013, Sydney, Australia, December 1-8, 2013, pages 825–830, 2013. 1, 3
  16. 16.Eunhee Kang, Won Chang, Jae Jun Yoo, and Jong Chul Ye. Deep convolutional framelet denosing for low-dose CT via wavelet residual network. IEEE Trans. Medical Imaging, 37 (6):1358–1369, 2018. 4
  17. 17.Chongyi Li, Jichang Guo, Runmin Cong, Yanwei Pang, and Bo Wang. Underwater image enhancement by dehazing with minimum information loss and histogram distribution prior. IEEE Trans. Image Process., 25(12):5664–5677, 2016. 1
  18. 18.Chongyi Li, Chunle Guo, Wenqi Ren, Runmin Cong, Junhui Hou, Sam Kwong, and Dacheng Tao. An underwater image enhancement benchmark dataset and beyond. IEEE Trans. Image Process., 29:4376–4389, 2020. 1, 3, 6
  19. 19.Chongyi Li, Saeed Anwar, Junhui Hou, Runmin Cong, Chunle Guo, and Wenqi Ren. Underwater image enhancement via medium transmission-guided multi-color space embedding. IEEE Trans. Image Process., 30:4985–5000, 2021. 3, 6
  20. 20.Hanyu Li, Jingjing Li, and Wei Wang. A fusion adversarial underwater image enhancement network with a public test dataset. arXiv preprint arXiv:1906.06819, 2019. 6
  21. 21.Jie Li, Katherine A. Skinner, Ryan M. Eustice, and Matthew Johnson-Roberson. Watergan: Unsupervised generative network to enable real-time color correction of monocular underwater images. IEEE Robotics Autom. Lett., 3(1):387–394, 2018. 3
  22. 22.Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. In 2021 IEEE/CVF International Conference on Computer Vision, ICCV 2021, Montreal, QC, Canada, October 10-17, 2021, pages 9992–10002. IEEE, 2021. 2
  23. 23.Shilin Lu, Yanzhu Liu, and Adams Wai-Kin Kong. Tf-icon: Diffusion-based training-free cross-domain image composition. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 2294–2305, 2023. 2
  24. 24.Shilin Lu, Zilan Wang, Leyang Li, Yanzhu Liu, and Adams Wai-Kin Kong. Mace: Mass concept erasure in diffusion models. arXiv preprint arXiv:2403.06135, 2024. 2
  25. 25.Ziyin Ma and Changjae Oh. A wavelet-based dual-stream network for underwater image enhancement. In IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP 2022, Virtual and Singapore, 23-27 May 2022, pages 2769–2773. IEEE, 2022. 6
  26. 26.James McMahon and Erion Plaku. Autonomous data collection with timed communication constraints for unmanned underwater vehicles. IEEE Robotics Autom. Lett., 6(2): 1832–1839, 2021. 1
  27. 27.Karen Panetta, Chen Gao, and Sos Agaian. Human-visualsystem-inspired underwater image quality measures. IEEE Journal of Oceanic Engineering, 41(3):541–551, 2015. 7
  28. 28.Lintao Peng, Chunli Zhu, and Liheng Bian. U-shape transformer for underwater image enhancement. IEEE Trans. Image Process., 32:3066–3079, 2023. 1, 3, 6
  29. 29.Yan-Tsung Peng and Pamela C. Cosman. Underwater image restoration based on image blurriness and light absorption. IEEE Trans. Image Process., 26(4):1579–1594, 2017. 1, 3
  30. 30.Yan-Tsung Peng, Keming Cao, and Pamela C. Cosman. Generalization of the dark channel prior for single image restoration. IEEE Trans. Image Process., 27(6):2856–2868, 2018. 3
  31. 31.Priyadharsini Ravisankar, T. Sree Sharmila, and V. Rajendran. A wavelet transform based contrast enhancement method for underwater acoustic images. Multidimens. Syst. Signal Process., 29(4):1845–1859, 2018. 1, 4
  32. 32.Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bj¨orn Ommer. High-resolution image synthesis with latent diffusion models. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2022, New Orleans, LA, USA, June 18-24, 2022, pages 10674– 10685. IEEE, 2022. 2
  33. 33.Chitwan Saharia, William Chan, Huiwen Chang, Chris A. Lee, Jonathan Ho, Tim Salimans, David J. Fleet, and Mohammad Norouzi. Palette: Image-to-image diffusion models. In SIGGRAPH ’22: Special Interest Group on Computer Graphics and Interactive Techniques Conference, Vancouver, BC, Canada, August 7 - 11, 2022, pages 15:1–15:10. ACM, 2022. 3
  34. 34.Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily L. Denton, Seyed Kamyar Seyed Ghasemipour, Raphael Gontijo Lopes, Burcu Karagol Ayan, Tim Salimans, Jonathan Ho, David J. Fleet, and Mohammad Norouzi. Photorealistic text-to-image diffusion models with deep language understanding. In NeurIPS, 2022. 2
  35. 35.Jiaming Song, Chenlin Meng, and Stefano Ermon. Denoising diffusion implicit models. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021, 2021. 2, 3, 6
  36. 36.Yi Tang, Hiroshi Kawasaki, and Takafumi Iwaguchi. Underwater image enhancement by transformer-based diffusion model with non-uniform sampling for skip strategy. In Proceedings of the 31st ACM International Conference on Multimedia, MM 2023, Ottawa, ON, Canada, 29 October 2023- 3 November 2023, pages 5419–5427. ACM, 2023. 1, 2, 3, 6
  37. 37.Pritish M. Uplavikar, Zhenyu Wu, and Zhangyang Wang. All-in-one underwater image enhancement using domainadversarial learning. In IEEE Conference on Computer Vision and Pattern Recognition Workshops, CVPR Workshops 2019, Long Beach, CA, USA, June 16-20, 2019, pages 1–8, 2019. 3
  38. 38.Yi Wang, Hui Liu, and Lap-Pui Chau. Single underwater image restoration using adaptive attenuation-curve prior. IEEE Trans. Circuits Syst. I Regul. Pap., 65-I(3):992–1002, 2018. 3
  39. 39.Yudong Wang, Jichang Guo, Huan Gao, and HuiHui Yue. Uiecˆ2-net: Cnn-based underwater image enhancement using two color space. Signal Process. Image Commun., 96: 116250, 2021. 6
  40. 40.Yinhuai Wang, Jiwen Yu, and Jian Zhang. Zero-shot image restoration using denoising diffusion null-space model. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023, 2023. 2, 3
  41. 41.Zhou Wang, Alan C. Bovik, Hamid R. Sheikh, and Eero P. Simoncelli. Image quality assessment: from error visibility to structural similarity. IEEE Trans. Image Process., 13(4): 600–612, 2004. 6
  42. 42.Jay Whang, Mauricio Delbracio, Hossein Talebi, Chitwan Saharia, Alexandros G. Dimakis, and Peyman Milanfar. Deblurring via stochastic refinement. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2022, New Orleans, LA, USA, June 18-24, 2022, pages 16272– 16282. IEEE, 2022. 3
  43. 43.Jian Yang, Chen Li, and Xuelong Li. Underwater image restoration with light-aware progressive network. In IEEE International Conference on Acoustics, Speech and Signal Processing ICASSP 2023, Rhodes Island, Greece, June 4-10, 2023, pages 1–5. IEEE, 2023. 3
  44. 44.Miao Yang and Arcot Sowmya. An underwater color image quality evaluation metric. IEEE Transactions on Image Processing, 24(12):6062–6071, 2015. 7
  45. 45.Zongyuan Yang, Baolin Liu, Yongping Xiong, Lan Yi, Guibin Wu, Xiaojun Tang, Ziqi Liu, Junjie Zhou, and Xing Zhang. Docdiff: Document enhancement via residual diffusion models. In Proceedings of the 31st ACM International Conference on Multimedia, MM 2023, Ottawa, ON, Canada, 29 October 2023- 3 November 2023, pages 2795– 2806. ACM, 2023. 2, 3
  46. 46.Xunpeng Yi, Han Xu, Hao Zhang, Linfeng Tang, and Jiayi Ma. Diff-retinex: Rethinking low-light image enhancement with A generative diffusion model. CoRR, abs/2308.13164, 2023. 2, 3
  47. 47.Richard Zhang, Phillip Isola, Alexei A. Efros, Eli Shechtman, and Oliver Wang. The unreasonable effectiveness of deep features as a perceptual metric. In 2018 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2018, Salt Lake City, UT, USA, June 18-22, 2018, pages 586– 595, 2018. 6
  48. 48.Chen Zhao, Wei-Ling Cai, Chenyu Dong, and Ziqi Zeng. Toward sufficient spatial-frequency interaction for gradientaware underwater image enhancement. arXiv preprint arXiv:2202.08537, 2023. 2, 4
  49. 49.Chen Zhao, Wei-Ling Cai, and Zheng Yuan. Spectral normalization and dual contrastive regularization for image-toimage translation. The Visual Computer, pages 1–12, 2024. 3
  50. 50.Chen Zhao, Chenyu Dong, and Weiling Cai. Learning a physical-aware diffusion model based on transformer for underwater image enhancement. arXiv preprint arXiv:2403.01497, 2024. 3
  51. 51.Dewei Zhou, Zongxin Yang, and Yi Yang. Pyramid diffusion models for low-light image enhancement. arXiv preprint arXiv:2305.10028, 2023. 2
  52. 52.Dewei Zhou, You Li, Fan Ma, Zongxin Yang, and Yi Yang. Migc: Multi-instance generation controller for text-to-image synthesis. arXiv preprint arXiv:2402.05408, 2024. 2

Citation

MLA
Zhao, C., et al. “Wavelet-based Fourier Information Interaction with Frequency Diffusion Adjustment for Underwater Image Restoration”. arXiv, 2023, http://arxiv.org/abs/2311.16845v1.
APA
Zhao, C., Cai, W., Dong, C., & Hu, C. (2023). Wavelet-based Fourier Information Interaction with Frequency Diffusion Adjustment for Underwater Image Restoration. arXiv. http://arxiv.org/abs/2311.16845v1
Chicago
Zhao, C., W. Cai, C. Dong, and C. Hu. 2023. “Wavelet-based Fourier Information Interaction with Frequency Diffusion Adjustment for Underwater Image Restoration”. arXiv. http://arxiv.org/abs/2311.16845v1.
Harvard
Zhao, C. et al. (2023) “Wavelet-based Fourier Information Interaction with Frequency Diffusion Adjustment for Underwater Image Restoration”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2311.16845v1.
Vancouver
1. Zhao C, Cai W, Dong C, Hu C (2023) Wavelet-based Fourier Information Interaction with Frequency Diffusion Adjustment for Underwater Image Restoration. arXiv

BibTeX

@article{zhao2023wavelet,
  title = {Wavelet-based Fourier Information Interaction with Frequency Diffusion Adjustment for Underwater Image Restoration},
  author = {Zhao, Chen and Cai, Weiling and Dong, Chenyu and Hu, Chengwei},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2311.16845v1},
  eprint = {2311.16845}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE