Fast Underwater Image Enhancement for Improved Visual Perception
Md Jahidul IslamYouya XiaJunaed Sattar
Proposes a real-time conditional generative adversarial network and the large-scale EUVP dataset to correct degraded underwater imagery and boost the performance of robotic perception tasks such as object detection and human pose estimation.
Autonomous and remotely operated underwater vehicles face severe operational challenges due to underwater optical distortions, including light absorption, scattering, and color loss that create heavy green or blue hues and low contrast. Traditional physics-based enhancement methods require physical priors such as depth maps and optical water parameters, which are rarely accessible in dynamic robotic missions, while existing deep learning alternatives are computationally too heavy for real-time robotic deployment. The article introduces FUnIE-GAN, an efficient conditional generative adversarial network designed to perform real-time underwater image enhancement and improve downstream visual perception tasks without requiring physical environmental priors.
To train and validate the system, the authors constructed the Enhancement of Underwater Visual Perception (EUVP) dataset, which contains over 20,000 real-world underwater images captured across varied visibility conditions using seven distinct camera setups. The model architecture pairs a streamlined, five-layer encoder-decoder generator featuring skip-connections with a lightweight patch-based discriminator. The training process balances global structural similarity, feature-level content preservation, and local texture rendition for paired data, while utilizing cycle-consistency constraints for unpaired training.
Key findings show that FUnIE-GAN outperforms existing learning-based and physics-based models in both reconstruction quality and processing speed. In quantitative evaluations on 1,000 test images, the paired model achieved the highest peak signal-to-noise ratio and structural similarity scores against all baselines, and human evaluators ranked its outputs as the top visual quality choice in 88.6% of evaluation comparisons. Using enhanced images as inputs directly improved downstream robotic perception tasks, boosting underwater object detection accuracy by 11% to 14%, human body-pose estimation by 22% to 28%, and visual saliency prediction by 26% to 28%. Computationally, the model requires only 17 megabytes of memory and processes 256x256 images at 25.4 frames per second on embedded hardware (NVIDIA Jetson TX2) and 148.5 frames per second on a dedicated graphics card, whereas baseline methods run far too slowly for live operations.
These results demonstrate that visual enhancement can serve as a practical, lightweight pre-processing stage in autonomous robotic pipelines, substantially reducing failure risks in critical subsea inspection, environmental monitoring, and diver-robot collaboration. The authors recommend adopting this model as an integrated visual front-end on visually guided underwater vehicles, while reserving more complex, parameter-heavy architectures for offline processing pipelines. Future work should prioritize stabilizing unpaired training convergence and extending the pipeline to support higher-resolution image inputs and broader multi-agent marine tasks.
Confidence in these findings is strong across standard visibility ranges, but operational caution is advised in extreme conditions. The model exhibits performance degradation and noise amplification when applied to severely degraded, texture-less scenes, and unpaired training configurations remain prone to color inconsistencies due to training instability. Additionally, because the architecture targets perceptual image enhancement rather than exact physical radiance modeling, it does not guarantee ground-truth optical pixel restoration.
- Paper: Image-to-Image Translation with Conditional Adversarial Networks, Phillip Isola et al. (2017). It introduces the foundational conditional GAN framework (pix2pix) for supervised image-to-image translation upon which the underwater enhancement model directly builds.
- Paper: Unpaired Image-to-Image Translation Using Cycle-Consistent Adversarial Networks, Jun-Yan Zhu et al. (2017). It provides the cycle-consistency adversarial training paradigm for unpaired image translation, essential for understanding the source paper's unpaired training pipeline on underwater datasets.
- Paper: Photo-Realistic Single Image Super-Resolution Using a Generative Adversarial Network, Christian Ledig et al. (2017). It establishes perceptual loss functions combining adversarial and content features that inspire the perceptual quality objectives used to supervise underwater enhancement.
- Paper: High-Resolution Image Synthesis and Semantic Manipulation with Conditional GANs, Ting-Chun Wang et al. (2018). It advances conditional GAN architectures with multi-scale discriminator and feature-matching losses that inform high-quality real-time restoration models.
- Paper: An Underwater Image Enhancement Benchmark Dataset and Beyond, Chongyi Li et al. (2019). It provides the benchmark datasets, physical degradation insights, and baseline deep learning architectures for real-world underwater image restoration.
- Paper: AOD-Net: All-in-One Dehazing Network, Boyi Li et al. (2017). It pioneers lightweight end-to-end atmospheric restoration networks designed explicitly for downstream high-level vision tasks such as object detection.
- Paper: DeblurGAN: Blind Motion Deblurring Using Conditional Adversarial Networks, Orest Kupyn et al. (2017). It presents a lightweight conditional GAN framework for fast image restoration evaluated directly on downstream detection tasks.
- Paper: Zero-Reference Deep Curve Estimation for Low-Light Image Enhancement, Chunle Guo et al. (2020). It extends vision-enhancement preprocessing to a zero-reference learning framework using lightweight curve estimation for extreme lighting degradation.
- Paper: Multi-Stage Progressive Image Restoration, Syed Waqas Zamir et al. (2021). It generalizes multi-stage progressive restoration pipelines to achieve state-of-the-art detail preservation across diverse adverse ambient conditions.
- Paper: Pre-Trained Image Processing Transformer, Hanting Chen et al. (2020). It scales up image processing and degradation restoration using large-scale transformer pre-training across multiple vision tasks.
- Paper: Uformer: A General U-Shaped Transformer for Image Restoration, Zhendong Wang et al. (2021). It introduces a hierarchical U-shaped transformer architecture that advances beyond CNN-based models for high-resolution image restoration.
- Paper: MUSIQ: Multi-scale Image Quality Transformer, Junjie Ke et al. (2021). It provides a multi-scale transformer approach to perceptual image quality assessment without requiring destructive image cropping or resizing.
