Wavelet-based Fourier Information Interaction with Frequency Diffusion Adjustment for Underwater Image Restoration
Chen ZhaoWeiling CaiChenyu DongChengwei Hu
Proposes WF-Diff, a two-stage underwater image restoration framework that integrates wavelet-Fourier frequency interactions with a frequency residual diffusion adjustment module to effectively correct color distortion and recover fine textures.
Underwater image restoration is critical for marine robotics, subsea monitoring, and underwater object tracking. Light refraction, absorption, and scattering in aquatic environments cause severe distortions, such as heavy color casts, loss of contrast, and blurred structural details. While deep learning methods have improved restoration quality, existing models primarily operate in standard pixel space. As a result, they fail to leverage frequency-based representations and struggle to recover fine textures without introducing unwanted artifacts.
The article demonstrates a two-stage restoration framework, named WF-Diff, that combines frequency decomposition with targeted diffusion refinement. The objective is to systematically separate color correction from fine detail enhancement across distinct frequency bands, maximizing image clarity and structural precision.
The framework first decomposes underwater images into low-frequency and high-frequency components. A preliminary restoration network applies spatial and frequency fusion blocks to correct low-frequency color components while using transformer blocks to rebuild high-frequency structures. A cross-frequency conditioner shares contextual cues between these paths. In the second stage, a plug-and-play residual diffusion module refines the remaining high- and low-frequency errors against ground-truth targets. The authors validated the method against eight leading restoration models across three benchmark datasets totaling thousands of real-world test and reference images.
The evaluation produced several notable results. First, the proposed framework achieved state-of-the-art restoration quality across the primary benchmark datasets, reaching a peak signal-to-noise ratio of 23.86 on the UIEBD dataset compared to 21.88 for the top existing diffusion baseline. Second, perceptual distortion and distributional error metrics improved significantly, lowering the Fréchet Inception Distance score on UIEBD to 27.85 from the prior baseline of 31.07. Third, ablation testing confirmed that learning residual frequency distributions prevented the visual artifacts and hallucinations common in standard pixel-level diffusion models. Finally, the framework generalized effectively across unreferenced test datasets, yielding superior non-reference quality scores.
These findings indicate that splitting underwater image enhancement into frequency-specific preliminary restoration and residual diffusion substantially enhances visual fidelity. For operations reliant on subsea computer vision, clearer imagery reduces operational risk and improves detection accuracy in turbid conditions. However, the use of two diffusion models increases computational demand: the full framework requires approximately 0.28 seconds of inference time per image, compared to 0.01 to 0.14 seconds for competing alternatives.
Organizations developing subsea vision systems should consider adopting frequency-guided residual architectures for post-processing or high-precision inspection pipelines. Before deploying this architecture in real-time edge environments—such as autonomous underwater vehicles—engineering teams should prioritize reducing the diffusion sampling steps or applying model compression techniques to accelerate inference speed.
- Paper: An Underwater Image Enhancement Benchmark Dataset and Beyond, Chongyi Li et al. (2019). This paper establishes the UIEBD dataset and baseline enhancement benchmarks that the source directly uses to train and validate its frequency-based restoration framework.
- Paper: Denoising Diffusion Probabilistic Models, Jonathan Ho et al. (2020). This foundational paper establishes the formulation and training principles of denoising diffusion probabilistic models, which underpin the residual diffusion refinement module utilized in the source.
- Paper: Fast Underwater Image Enhancement for Improved Visual Perception, Md Jahidul Islam et al. (2019). This work introduces fast, deep learning-based underwater image enhancement and the EUVP dataset, providing foundational context on handling aquatic optical degradation and efficiency trade-offs.
- Paper: Restormer: Efficient Transformer for High-Resolution Image Restoration, Syed Waqas Zamir et al. (2022). This paper presents efficient transformer blocks for multi-scale image restoration, informing the high-frequency transformer processing components used in the source's preliminary stage.
- Paper: Denoising Diffusion Restoration Models, Bahjat Kawar et al. (2022). This study demonstrates how diffusion models can be guided through spectral and frequency decompositions to solve image restoration tasks without severe artifacts.
- Paper: Diffusion Posterior Sampling for General Noisy Inverse Problems, Hyungjin Chung et al. (2022). This work establishes diffusion posterior sampling for noisy inverse imaging problems, providing key theoretical groundwork for guided diffusion-based restoration.
- Paper: Multi-Stage Progressive Image Restoration, Syed Waqas Zamir et al. (2021). This paper introduces multi-stage progressive restoration architectures with cross-stage feature fusion, inspiring the sequential restoration and refinement structure of WF-Diff.
- Paper: Dual-Domain Attention for Image Deblurring, Yuning Cui et al. (2023). This study demonstrates decoupling restoration across spatial and frequency domains using attention mechanisms, directly preceding the dual-frequency conditioning design of the source.
- Paper: Fourier Priors-Guided Diffusion for Zero-Shot Joint Low-Light Enhancement and Deblurring, Xiaoqian Lv et al. (2024). This work builds on frequency-domain priors in diffusion models by leveraging Fourier amplitude and phase separation for zero-shot joint low-light enhancement and deblurring.
- Paper: A Variational Perspective on Solving Inverse Problems with Diffusion Models, Morteza Mardani et al. (2024). This paper introduces a variational perspective (RED-diff) to refine posterior sampling across diffusion stages, extending the theoretical principles of residual diffusion refinement explored in the source.
- Paper: Variational Flow Maps: Make Some Noise for One-Step Conditional Generation, Abbas Mammadov et al. (2026). This work directly addresses the multi-step inference bottleneck highlighted in the source by developing variational flow maps for one-step conditional generation in inverse restoration problems.
