An Underwater Image Enhancement Benchmark Dataset and Beyond
Chongyi LiChunle GuoWenqi RenRunmin CongJunhui HouSam KwongDacheng Tao
Establishes a standardized real-world benchmark dataset containing 950 underwater images alongside a convolutional baseline network, Water-Net, to rigorously evaluate and advance underwater image restoration.
Underwater imaging plays a critical role in marine biology, archaeology, environmental monitoring, and aquatic robotics. However, underwater photographs frequently suffer from severe color casts, reduced contrast, and heavy haze caused by light absorption and scattering in water. While numerous enhancement algorithms have been proposed in recent years, their real-world effectiveness has remained difficult to gauge because existing evaluations rely heavily on synthetic datasets or small, unrepresentative image samples. The article set out to address this gap by establishing a comprehensive, real-world benchmark to systematically evaluate existing enhancement techniques and to train an effective deep learning model for real-world underwater enhancement.
To achieve this, the article introduces a benchmark consisting of 950 real underwater images captured under diverse lighting and environmental conditions. High-quality reference counterparts were established for 890 of these images through extensive pairwise visual comparisons conducted by 50 human observers across twelve enhancement techniques, while the remaining 60 images without satisfactory outputs were designated as a challenging test set. Using this benchmark, the article evaluated state-of-the-art physical, non-physical, and data-driven methods across visual quality, full-reference metrics, non-reference metrics, and computational runtime. Additionally, the article developed Water-Net, a baseline convolutional neural network trained on this dataset using a gated fusion strategy and a perceptual loss function to combine white balancing, histogram equalization, and gamma correction.
Key findings show that no single existing enhancement technique consistently succeeds across all underwater scenarios. Fusion-based and commercial enhancement tools produced the most visually pleasing references, while physical-model methods frequently failed due to inaccurate light attenuation assumptions. Crucially, the analysis revealed that commonly used non-reference underwater quality metrics often disagree with human visual judgment, rewarding artificial contrast and color shifts that human observers rated poorly. In quantitative and qualitative tests on unseen images, the proposed Water-Net model outperformed both traditional and generative deep learning methods, delivering superior contrast, natural colors, and fast processing speeds of roughly eight frames per second.
These findings indicate that future research and operational deployments must move away from simplified optical models and unreliable non-reference metrics. Instead, engineering teams and researchers should adopt data-driven fusion architectures and physically accurate imaging formulations, while developing better quality metrics aligned with human perception. Although the benchmark significantly advances the field, the authors note limitations: far-distance backscatter remains difficult to fully eliminate, and some human observers may overlook subtle physical scattering artifacts. Moving forward, incorporating depth maps, expanding datasets to include video sequences, and refining reference selection with advanced optical physics will be essential next steps before deploying these systems in highly automated aquatic missions.
- Paper: Benchmarking Single-Image Dehazing and Beyond, Boyi Li et al. (2017). Its systematic dehazing benchmark provides a useful precedent for understanding this paper’s evaluation across algorithms, objective metrics, and human judgments.
- Paper: Underwater Ranker: Learn Which Is Better and How to Be Better, Chunle Guo et al. (2023). Building on a similar set of 890 underwater images and pairwise human comparisons, this work extends benchmark evaluation by learning a perceptual ranker that can guide restoration models.
