Fast Underwater Image Enhancement for Improved Visual Perception

Md Jahidul IslamYouya XiaJunaed Sattar

article2019IEEE Robotics and Automation Letters1,528 citations

Proposes a real-time conditional generative adversarial network and the large-scale EUVP dataset to correct degraded underwater imagery and boost the performance of robotic perception tasks such as object detection and human pose estimation.

Listen

Autonomous and remotely operated underwater vehicles face severe operational challenges due to underwater optical distortions, including light absorption, scattering, and color loss that create heavy green or blue hues and low contrast. Traditional physics-based enhancement methods require physical priors such as depth maps and optical water parameters, which are rarely accessible in dynamic robotic missions, while existing deep learning alternatives are computationally too heavy for real-time robotic deployment. The article introduces FUnIE-GAN, an efficient conditional generative adversarial network designed to perform real-time underwater image enhancement and improve downstream visual perception tasks without requiring physical environmental priors.

To train and validate the system, the authors constructed the Enhancement of Underwater Visual Perception (EUVP) dataset, which contains over 20,000 real-world underwater images captured across varied visibility conditions using seven distinct camera setups. The model architecture pairs a streamlined, five-layer encoder-decoder generator featuring skip-connections with a lightweight patch-based discriminator. The training process balances global structural similarity, feature-level content preservation, and local texture rendition for paired data, while utilizing cycle-consistency constraints for unpaired training.

Key findings show that FUnIE-GAN outperforms existing learning-based and physics-based models in both reconstruction quality and processing speed. In quantitative evaluations on 1,000 test images, the paired model achieved the highest peak signal-to-noise ratio and structural similarity scores against all baselines, and human evaluators ranked its outputs as the top visual quality choice in 88.6% of evaluation comparisons. Using enhanced images as inputs directly improved downstream robotic perception tasks, boosting underwater object detection accuracy by 11% to 14%, human body-pose estimation by 22% to 28%, and visual saliency prediction by 26% to 28%. Computationally, the model requires only 17 megabytes of memory and processes 256x256 images at 25.4 frames per second on embedded hardware (NVIDIA Jetson TX2) and 148.5 frames per second on a dedicated graphics card, whereas baseline methods run far too slowly for live operations.

These results demonstrate that visual enhancement can serve as a practical, lightweight pre-processing stage in autonomous robotic pipelines, substantially reducing failure risks in critical subsea inspection, environmental monitoring, and diver-robot collaboration. The authors recommend adopting this model as an integrated visual front-end on visually guided underwater vehicles, while reserving more complex, parameter-heavy architectures for offline processing pipelines. Future work should prioritize stabilizing unpaired training convergence and extending the pipeline to support higher-resolution image inputs and broader multi-agent marine tasks.

Confidence in these findings is strong across standard visibility ranges, but operational caution is advised in extreme conditions. The model exhibits performance degradation and noise amplification when applied to severely degraded, texture-less scenes, and unpaired training configurations remain prone to color inconsistencies due to training instability. Additionally, because the architecture targets perceptual image enhancement rather than exact physical radiance modeling, it does not guarantee ground-truth optical pixel restoration.

arXiv: 1903.09766
Cover for Fast Underwater Image Enhancement for Improved Visual Perception

Abstract

In this paper, we present a conditional generative adversarial network-based model for real-time underwater image enhancement. To supervise the adversarial training, we formulate an objective function that evaluates the perceptual image quality based on its global content, color, local texture, and style information. We also present EUVP, a large-scale dataset of a paired and unpaired collection of underwater images (of poor' and good' quality) that are captured using seven different cameras over various visibility conditions during oceanic explorations and human-robot collaborative experiments. In addition, we perform several qualitative and quantitative evaluations which suggest that the proposed model can learn to enhance underwater image quality from both paired and unpaired training. More importantly, the enhanced images provide improved performances of standard models for underwater object detection, human pose estimation, and saliency prediction. These results validate that it is suitable for real-time preprocessing in the autonomy pipeline by visually-guided underwater robots. The model and associated training pipelines are available at this https URL.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Automatic Image Enhancement
  • 2.2 Improving Underwater Visual Perception
  • 3 Proposed Model and Dataset
  • 3.1 FUnIE-GAN Architecture
  • 3.2 Objective Function Formulation
  • 3.2.1 Paired Training
  • 3.2.2 Unpaired Training
  • 3.3 EUVP Dataset
  • 4 Experimental Results
  • 4.1 Qualitative Evaluations
  • 4.2 Quantitative Evaluation
  • 4.3 User Study
  • 4.4 Improved Visual Perception
  • 4.5 Limitations and Failure Cases
  • 5 Conclusion
  • References

Knowls

  1. Knowl 1 — FUnIE-GAN Network Architecture

    model/method

    FUnIE-GAN is a fully-convolutional conditional generative adversarial network designed for real-time underwater image enhancement on 256×256×3256 \times 256 \times 3 RGB images.

    Generator Architecture: The generator follows a lightweight U-Net encoder-decoder structure with mirrored skip-connections:

    • Encoder (e1e_1 to e5e_5): Consists of 5 convolutional stages using 4×44 \times 4 2D convolution filters, Leaky-ReLU non-linear activations, and Batch Normalization (BN). Spatial dimensions and channel depths scale as: 128×128×32128 \times 128 \times 32 (e1e_1), 64×64×12864 \times 64 \times 128 (e2e_2), 32×32×25632 \times 32 \times 256 (e3e_3), 16×16×25616 \times 16 \times 256 (e4e_4), down to a bottleneck of 8×8×2568 \times 8 \times 256 feature maps (e5e_5).
    • Decoder (d1d_1 to d5d_5): Consists of 5 stages using 4×44 \times 4 transposed 2D convolutions (deconvolutions), Dropout, and BN. The outputs of encoder layers e1,e2,e3,e4e_1, e_2, e_3, e_4 are concatenated via skip-connections to mirrored decoder layers d5,d4,d3,d2d_5, d_4, d_3, d_2, respectively. Feature map dimensions scale as: 16×16×51216 \times 16 \times 512 (d1d_1), 32×32×51232 \times 32 \times 512 (d2d_2), 64×64×25664 \times 64 \times 256 (d3d_3), 128×128×64128 \times 128 \times 64 (d4d_4), and 256×256×32256 \times 256 \times 32 (d5d_5), yielding a final enhanced RGB output of dimension 256×256×3256 \times 256 \times 3.
    • The entire generator omits fully-connected layers to minimize parameter count and maximize inference speed.

    Discriminator Architecture: The discriminator is a Markovian PatchGAN that evaluates local image patches of size 16×1616 \times 16 rather than the global image, capturing high-frequency local textures and style while reducing computational cost:

    • Input: Concatenation of the input image and the candidate (real or generated) image, forming a 256×256×6256 \times 256 \times 6 tensor.
    • Layers: 4 convolutional layers employing 3×33 \times 3 kernels with a stride of 2, followed by Leaky-ReLU and BN, with channel depths progressing through 32 (128×128128 \times 128), 64 (64×6464 \times 64), 128 (32×3232 \times 32), and 256 (16×1616 \times 16).
    • Output: A 16×16×116 \times 16 \times 1 matrix representing the patch-level averaged validity scores.
  2. Knowl 2 — FUnIE-GAN Paired Objective Function

    equation

    For paired supervision, FUnIE-GAN optimizes a multi-modal objective function that combines a conditional adversarial loss, an L1L_1 global similarity loss, and a VGG-19 deep feature content loss.

    Let XX denote the source domain of distorted underwater images, YY denote the target domain of enhanced images, and ZZ represent random noise. The total minimax training objective is formulated as:

    G∗=arg⁡min⁡Gmax⁡DLcGAN(G,D)+λ1L1(G)+λcLcon(G)G^* = \arg\min_G \max_D \mathcal{L}_{cGAN}(G, D) + \lambda_1 \mathcal{L}_1(G) + \lambda_c \mathcal{L}_{con}(G)

    where the individual loss terms are defined as follows:

    1. Conditional Adversarial Loss: LcGAN(G,D)=EX,Y[log⁡D(Y)]+EX,Y,Z[log⁡(1−D(X,G(X,Z)))]\mathcal{L}_{cGAN}(G, D) = \mathbb{E}_{X,Y}[\log D(Y)] + \mathbb{E}_{X,Y,Z}[\log(1 - D(X, G(X, Z)))] where generator GG attempts to fool discriminator DD, and DD penalizes locally inconsistent textures and styles.

    2. Global Similarity Loss (L1L_1): L1(G)=EX,Y,Z[∥Y−G(X,Z)∥1]\mathcal{L}_1(G) = \mathbb{E}_{X,Y,Z}\left[\|Y - G(X, Z)\|_1\right] which enforces global appearance preservation while mitigating spatial blurring.

    3. Content Loss (LconL_{con}): Lcon(G)=EX,Y,Z[∥Φ(Y)−Φ(G(X,Z))∥2]\mathcal{L}_{con}(G) = \mathbb{E}_{X,Y,Z}\left[\|\Phi(Y) - \Phi(G(X, Z))\|_2\right] where Φ(⋅)\Phi(\cdot) extracts high-level semantic feature representations from the block5_conv2 layer of a pre-trained VGG-19 network to furnish fine texture details.

    The scaling hyperparameters are set empirically to λ1=0.7\lambda_1 = 0.7 and λc=0.3\lambda_c = 0.3.

  3. Knowl 3 — FUnIE-GAN Unpaired Objective Function

    equation

    When pairwise ground-truth target images are unavailable, FUnIE-GAN-UP learns bidirectional domain transformations between distorted underwater images XX and good quality images YY via cycle consistency without requiring explicit paired correspondence.

    The forward generator GF:{X,Z}→YG_F: \{X, Z\} \to Y and reconstruction generator GR:{Y,Z}→XG_R: \{Y, Z\} \to X are jointly trained with two discriminators DYD_Y and DXD_X via the minimax objective:

    GF∗,GR∗=arg⁡min⁡GF,GRmax⁡DY,DXLcGAN(GF,DY)+LcGAN(GR,DX)+λcycLcyc(GF,GR)G_F^*, G_R^* = \arg\min_{G_F, G_R} \max_{D_Y, D_X} \mathcal{L}_{cGAN}(G_F, D_Y) + \mathcal{L}_{cGAN}(G_R, D_X) + \lambda_{cyc} \mathcal{L}_{cyc}(G_F, G_R)

    where the cycle-consistency loss is given by:

    Lcyc(GF,GR)=EX,Y,Z[∥X−GR(GF(X,Z))∥1]+EX,Y,Z[∥Y−GF(GR(Y,Z))∥1]\mathcal{L}_{cyc}(G_F, G_R) = \mathbb{E}_{X,Y,Z}\left[\|X - G_R(G_F(X, Z))\|_1\right] + \mathbb{E}_{X,Y,Z}\left[\|Y - G_F(G_R(Y, Z))\|_1\right]

    and the adversarial losses LcGAN(GF,DY)\mathcal{L}_{cGAN}(G_F, D_Y) and LcGAN(GR,DX)\mathcal{L}_{cGAN}(G_R, D_X) encourage GFG_F and GRG_R to generate outputs matching the marginal distributions of YY and XX, respectively. The cycle-consistency hyperparameter is set empirically to λcyc=0.1\lambda_{cyc} = 0.1. Explicit L1L_1 global similarity or VGG content loss terms are omitted in the unpaired formulation because Lcyc\mathcal{L}_{cyc} computes reconstruction constraints directly in L1L_1 space.

  4. Knowl 4 — EUVP Dataset for Underwater Image Enhancement

    experimental setup

    The Enhancement of Underwater Visual Perception (EUVP) dataset is a benchmark comprising 20,000 underwater images categorized into paired (over 12,000 instances) and unpaired (over 8,000 instances) sets of poor- and good-quality perceptual samples.

    Data Collection & Diversity: Images were collected during oceanic explorations, human-robot collaborative experiments, and curated YouTube video extractions across multiple waterbody types, natural lighting conditions, and scenes. Imagery was acquired using seven optical systems:

    • Multiple GoPro Hero 5 cameras
    • Aqua AUV uEye cameras
    • BlueRobotics low-light HD USB cameras
    • Trident ROV HD cameras

    Dataset Preparation:

    • Unpaired Collection: Six human evaluators inspected image properties (contrast, color cast, sharpness, and foreground object interpretability) to categorize images into 'poor' and 'good' quality subsets reflecting human perceptual preferences.
    • Paired Collection: A CycleGAN model was trained on the unpaired good and poor images to learn the domain degradation transformation. The learned mapping was then applied to distort good underwater images, synthetically generating corresponding paired degraded instances. This set was further augmented with underwater images sampled from ImageNet and Flickr.

    Training Setup: For model training, 11,000 paired instances and 7,500 unpaired instances were used for training, with the remaining instances reserved for validation and testing. Training was performed for 60,000 to 70,000 iterations with a batch size of 8 on four NVIDIA GeForce GTX 1080 GPUs using TensorFlow.

  5. Knowl 5 — PSNR and SSIM Enhancement Performance on EUVP Dataset

    data/table

    The quantitative reconstruction quality of FUnIE-GAN and FUnIE-GAN-UP was evaluated on 1,000 paired test images (256×256256 \times 256 resolution) from the EUVP dataset against baseline physics-based methods (Uw-HL, Mband-EN) and learning-based GAN models (Res-WGAN, Res-GAN, LS-GAN, Pix2Pix, UGAN-P, CycleGAN).

    Reconstruction fidelity is quantified using Peak Signal-to-Noise Ratio (PSNR, in dB) and Structural Similarity Index (SSIM, unitless in [0,1][0, 1]):

    Model PSNR (G(x),yG(x), y) SSIM (G(x),yG(x), y)
    Input (Degraded) 17.27±2.8817.27 \pm 2.88 0.6200±0.0750.6200 \pm 0.075
    Uw-HL 18.85±1.7618.85 \pm 1.76 0.7722±0.0660.7722 \pm 0.066
    Mband-EN 12.11±2.5512.11 \pm 2.55 0.4565±0.0970.4565 \pm 0.097
    Res-WGAN 16.46±1.8016.46 \pm 1.80 0.5762±0.0140.5762 \pm 0.014
    Res-GAN 14.75±2.2214.75 \pm 2.22 0.4685±0.1220.4685 \pm 0.122
    LS-GAN 17.83±2.8817.83 \pm 2.88 0.6725±0.0620.6725 \pm 0.062
    Pix2Pix 20.27±2.6620.27 \pm 2.66 0.7081±0.0690.7081 \pm 0.069
    UGAN-P 19.59±2.5419.59 \pm 2.54 0.6685±0.0750.6685 \pm 0.075
    CycleGAN 17.14±2.6517.14 \pm 2.65 0.6400±0.0800.6400 \pm 0.080
    FUnIE-GAN-UP 21.36±2.1721.36 \pm 2.17 0.8164±0.0460.8164 \pm 0.046
    FUnIE-GAN 21.92±1.07\mathbf{21.92 \pm 1.07} 0.8876±0.068\mathbf{0.8876 \pm 0.068}

    FUnIE-GAN attains the highest PSNR (21.9221.92 dB) and SSIM (0.88760.8876), outperforming all baseline models. FUnIE-GAN-UP (trained without pairs) ranks second, outperforming CycleGAN and deeper paired models including Pix2Pix and UGAN-P.

  6. Knowl 6 — UIQM Evaluation and Loss Component Ablation

    data/table

    Underwater Image Quality Measure (UIQM), a non-reference metric combining colorfulness, sharpness, and contrast, was evaluated on 1,000 paired and 2,000 unpaired test images from the EUVP dataset.

    Model Paired Data UIQM Unpaired Data UIQM
    Input (Degraded) 2.20±0.692.20 \pm 0.69 2.29±0.622.29 \pm 0.62
    Ground Truth 2.91±0.652.91 \pm 0.65 N/A
    Uw-HL 2.62±0.352.62 \pm 0.35 2.75±0.322.75 \pm 0.32
    Mband-EN 2.28±0.872.28 \pm 0.87 2.34±0.452.34 \pm 0.45
    Res-WGAN 2.55±0.642.55 \pm 0.64 2.46±0.672.46 \pm 0.67
    Res-GAN 2.62±0.892.62 \pm 0.89 2.28±0.342.28 \pm 0.34
    LS-GAN 2.37±0.782.37 \pm 0.78 2.59±0.522.59 \pm 0.52
    Pix2Pix 2.65±0.552.65 \pm 0.55 2.76±0.392.76 \pm 0.39
    UGAN-P 2.72±0.752.72 \pm 0.75 2.77±0.342.77 \pm 0.34
    CycleGAN 2.44±0.712.44 \pm 0.71 2.62±0.672.62 \pm 0.67
    FUnIE-GAN-UP 2.56±0.632.56 \pm 0.63 2.81±0.652.81 \pm 0.65
    FUnIE-GAN 2.78±0.43\mathbf{2.78 \pm 0.43} 2.98±0.51\mathbf{2.98 \pm 0.51}

    Ablation Study on Loss Components: An ablation analysis of FUnIE-GAN demonstrates the quantitative contributions of each loss term:

    • Removing the content loss Lcon\mathcal{L}_{con} causes a 1.07%1.07\% drop in UIQM.
    • Removing the global similarity loss L1\mathcal{L}_1 causes a 4.58%4.58\% drop in UIQM.
    • Removing both L1\mathcal{L}_1 and Lcon\mathcal{L}_{con} degrades average UIQM by 17.6%17.6\%, with corresponding steep declines in PSNR and SSIM.
  7. Knowl 7 — Downstream Visual Perception Improvements via FUnIE-GAN Enhancement

    empirical result

    Using FUnIE-GAN as a pre-processing module directly improves the quantitative accuracy of standard computer vision models on degraded underwater scenes across three core tasks:

    1. Underwater Object Detection: Evaluated on a deep visual diver-following detector, enhanced images produce an average accuracy improvement of 11%11\% to 14%14\%.
    2. 2D Multi-Person Body-Pose Estimation: Evaluated using the Part Affinity Fields model, enhanced images yield an average joint estimation improvement of 22%22\% to 28%28\%.
    3. Visual Saliency Prediction: Evaluated using fixation-driven salient object detection models, enhanced images achieve an average performance improvement of 26%26\% to 28%28\%.

    While competing models (UGAN-P, Pix2Pix, Uw-HL, Res-GAN, Res-WGAN) also yield performance gains on these perception tasks, they do so at computational latencies incompatible with real-time robotic deployment.

  8. Knowl 8 — Computational Efficiency and Embedded Runtime Benchmarks

    empirical result

    FUnIE-GAN features a memory footprint of 17 MB and achieves real-time inference speeds across diverse robotic and desktop compute platforms when processing 256×256256 \times 256 image inputs:

    • Embedded Single-Board Computer (NVIDIA Jetson TX2): Operates at 25.425.4 FPS (approx. 39.439.4 ms per frame).
    • Desktop GPU (NVIDIA GeForce GTX 1080): Operates at 148.5148.5 FPS (approx. 6.76.7 ms per frame).
    • Low-Power Mobile Robot CPU (Intel Core-i3 6100U): Operates at 7.97.9 FPS (approx. 126.6126.6 ms per frame).
    • Standard Desktop CPU (Intel Core-i5 3.6 GHz): Operates at >10>10 FPS (approx. 100100 ms per frame).

    In contrast, competing learning-based models (such as UGAN-P, Pix2Pix, CycleGAN, Res-GAN, and Res-WGAN) require between 200200 ms and 400400 ms per frame, and physics-based models (such as Uw-HL and Mband-EN) require 40004000 ms to 50005000 ms per frame on the same CPU platform.

  9. Knowl 9 — User Preference Study on Underwater Image Quality

    data/table

    A double-blind user study was conducted to assess human visual preference across 9 learning-based underwater enhancement models. A total of 78 participants provided 312 comparative evaluations by ranking the top 3 highest-quality outputs among randomized sets of 9 model-enhanced images.

    Model Rank-1 (%) Rank-2 (%) Rank-3 (%)
    FUnIE-GAN 24.50\mathbf{24.50} 68.50\mathbf{68.50} 88.60\mathbf{88.60}
    UGAN-P 21.2521.25 65.7565.75 80.5080.50
    FUnIE-GAN-UP 18.6718.67 48.2548.25 76.1876.18
    Pix2Pix 11.8811.88 45.1545.15 72.4572.45

    FUnIE-GAN received the highest preference across all three tiers, with an 88.60%88.60\% cumulative top-3 ranking score. Original degraded images obtained an average Rank-3 inclusion rate of only 6.67%6.67\%, demonstrating distinct user preference for enhanced imagery.

  10. Knowl 10 — Limitations and Failure Modes of FUnIE-GAN

    limitation

    FUnIE-GAN exhibits several operational limitations and specific failure modes:

    1. Severe Degradation and Low-Texture Scenes: On images with extreme contrast loss or homogeneous, textureless backgrounds, FUnIE-GAN tends to amplify noise artifacts, leading to over-saturation and poor restoration of fine local textures and true color.
    2. Unpaired Adversarial Instability: In the unpaired training regime (FUnIE-GAN-UP), the discriminator can overpower the generator early during optimization, causing vanishing gradients that halt generator convergence and produce inconsistent color tones and hue artifacts.
    3. Perceptual Rather Than Physical Restoration: FUnIE-GAN optimizes for perceptual image quality rather than physical radiance restoration. It does not model scene depth or wavelength-dependent optical attenuation coefficients, meaning it does not guarantee true physical radiance reconstruction.
    4. Resolution Trade-Off: The model's real-time efficiency is achieved by constraining the operational resolution to 256×256256 \times 256, requiring structural architecture modifications (such as additional bottleneck layers) to scale to higher-resolution inputs.

Coverage note — None was omitted; all key contributions including model architecture, paired and unpaired loss formulations, EUVP dataset details, quantitative PSNR/SSIM/UIQM benchmarks, downstream visual perception evaluations, inference speeds, user study data, and failure modes are fully represented.

References

  1. 1.M. Abadi, P. Barham, J. Chen, Z. Chen, A. Davis, J. Dean, et al. TensorFlow: A System for Large-scale Machine Learning. In USENIX Symposium on Operating Systems Design and Implementation (OSDI), pages 265–283, 2016.
  2. 2.D. Akkaynak and T. Treibitz. A Revised Underwater Image Formation Model. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 6723–6732, 2018.
  3. 3.M. Arjovsky, S. Chintala, and L. Bottou. Wasserstein Generative Adversarial Networks. In International Conference on Machine Learning (ICML), pages 214–223, 2017.
  4. 4.D. Berman, D. Levy, S. Avidan, and T. Treibitz. Underwater Single Image Color Restoration using Haze-Lines and a New Quantitative Dataset. arXiv preprint arXiv:1811.01343, 2018.
  5. 5.B. Bingham, B. Foley, H. Singh, R. Camilli, K. Delaporta, R. Eustice, et al. Robotic Tools for Deep Water Archaeology: Surveying an Ancient Shipwreck with an Autonomous Underwater Vehicle. Journal of Field Robotics (JFR), 27(6):702–717, 2010.
  6. 6.BlueRobotics. Low-light HD USB Camera. https://www.bluerobotics.com/, 2016. Accessed: 3-15-2019.
  7. 7.M. Bryson, M. Johnson-Roberson, O. Pizarro, and S. B. Williams. True Color Correction of Autonomous Underwater Vehicle Imagery. Journal of Field Robotics (JFR), 33(6):853–874, 2016.
  8. 8.B. Cai, X. Xu, K. Jia, C. Qing, and D. Tao. DehazeNet: An End-to-end System for Single Image Haze Removal. IEEE Transactions on Image Processing, 25(11):5187–5198, 2016.
  9. 9.Z. Cao, T. Simon, S.-E. Wei, and Y. Sheikh. Realtime Multi-person 2d Pose Estimation using Part Affinity Fields. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 7291–7299, 2017.
  10. 10.Y.-S. Chen, Y.-C. Wang, M.-H. Kao, and Y.-Y. Chuang. Deep Photo Enhancer: Unpaired Learning for Image Enhancement from Photographs with GANs. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 6306–6314. IEEE, 2018.
  11. 11.Z. Cheng, Q. Yang, and B. Sheng. Deep Colorization. In IEEE International Conference on Computer Vision (ICCV), pages 415–423, 2015.
  12. 12.Y. Cho, J. Jeong, and A. Kim. Model-assisted Multiband Fusion for Single Image Enhancement and Applications to Robot Vision. IEEE Robotics and Automation Letters (RA-L), 3(4):2822–2829, 2018.
  13. 13.J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei. ImageNet: A Large-scale Hierarchical Image Database. In Conference on Computer Vision and Pattern Recognition (CVPR), pages 248–255. IEEE, 2009.
  14. 14.G. Dudek, P. Giguere, C. Prahacs, S. Saunderson, J. Sattar, L.-A. Torres-Mendez, Jenkin, et al. Aqua: An Amphibious Autonomous Robot. Computer, 40(1):46–53, 2007.
  15. 15.C. Fabbri, M. J. Islam, and J. Sattar. Enhancing Underwater Imagery using Generative Adversarial Networks. In IEEE International Conference on Robotics and Automation (ICRA), pages 7159–7165. IEEE, 2018.
  16. 16.I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio. Generative Adversarial Nets. In Advances in Neural Information Processing Systems (NIPS), pages 2672–2680, 2014.
  17. 17.GoPro. GoPro Hero 5. https://gopro.com/, 2016. Accessed: 3-15-2019.
  18. 18.Y. Guo, H. Li, and P. Zhuang. Underwater Image Enhancement Using a Multiscale Dense Generative Adversarial Network. IEEE Journal of Oceanic Engineering, 2019.
  19. 19.K. He, J. Sun, and X. Tang. Single Image Haze Removal using Dark Channel Prior. IEEE Transactions on Pattern Analysis and Machine Intelligence, 33(12):2341–2353, 2010.
  20. 20.A. Hore and D. Ziou. Image Quality Metrics: PSNR vs. SSIM. In International Conference on Pattern Recognition, pages 2366–2369. IEEE, 2010.
  21. 21.A. Ignatov, N. Kobyshev, R. Timofte, K. Vanhoey, and L. Van Gool. DSLR-quality Photos on Mobile Devices with Deep Convolutional Networks. In IEEE International Conference on Computer Vision (ICCV), pages 3277–3285, 2017.
  22. 22.S. Ioffe and C. Szegedy. Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift. CoRR, abs/1502.03167, 2015.
  23. 23.M. J. Islam, M. Fulton, and J. Sattar. Toward a Generic Diver-Following Algorithm: Balancing Robustness and Efficiency in Deep Visual Detection. IEEE Robotics and Automation Letters (RA-L), 4(1):113–120, 2018.
  24. 24.M. J. Islam, M. Ho, and J. Sattar. Understanding Human Motion and Gestures for Underwater Human-Robot Collaboration. Journal of Field Robotics (JFR), pages 1–23, 2018.
  25. 25.P. Isola, J.-Y. Zhu, T. Zhou, and A. A. Efros. Image-to-image Translation with Conditional Adversarial Networks. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 1125–1134, 2017.
  26. 26.J. Johnson, A. Alahi, and L. Fei-Fei. Perceptual Losses for Real-time Style Transfer and Super-resolution. In European Conference on Computer Vision (ECCV), pages 694–711. Springer, 2016.
  27. 27.J. Li, X. Liang, Y. Wei, T. Xu, J. Feng, and S. Yan. Perceptual Generative Adversarial Networks for Small Object Detection. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 1222–1230, 2017.
  28. 28.J. Li, K. A. Skinner, R. M. Eustice, and M. Johnson-Roberson. WaterGAN: Unsupervised Generative Network to Enable Real-time Color Correction of Monocular Underwater Images. IEEE Robotics and Automation Letters (RA-L), 3(1):387–394, 2018.
  29. 29.P. Liu, G. Wang, H. Qi, C. Zhang, H. Zheng, and Z. Yu. Underwater Image Enhancement With a Deep Residual Framework. IEEE Access, 7:94614–94629, 2019.
  30. 30.R. Liu, M. Hou, X. Fan, and Z. Luo. Real-world Underwater Enhancement: Challenging, Benchmark and Efficient Solutions. arXiv preprint arXiv:1901.05320, 2019.
  31. 31.H. Lu, Y. Li, and S. Serikawa. Underwater Image Enhancement using Guided Trigonometric Bilateral Filter and Fast Automatic Color Correction. In IEEE International Conference on Image Processing, pages 3412–3416. IEEE, 2013.
  32. 32.A. L. Maas, A. Y. Hannun, and A. Y. Ng. Rectifier Nonlinearities Improve Neural Network Acoustic Models. In International Conference on Machine Learning (ICML), volume 30, page 3, 2013.
  33. 33.X. Mao, Q. Li, H. Xie, R. Y. Lau, Z. Wang, and S. Paul Smalley. Least Squares Generative Adversarial Networks. In IEEE International Conference on Computer Vision (ICCV), pages 2794–2802, 2017.
  34. 34.M. Mirza and S. Osindero. Conditional Generative Adversarial Nets. arXiv preprint arXiv:1411.1784, 2014.
  35. 35.OpenROV. TRIDENT. https://www.openrov.com/, 2017. Accessed: 3-15-2019.
  36. 36.K. Panetta, C. Gao, and S. Agaian. Human-visual-system-inspired Underwater Image Quality Measures. IEEE Journal of Oceanic Engineering, 41(3):541–551, 2016.
  37. 37.Z.-u. Rahman, D. J. Jobson, and G. A. Woodell. Retinex Processing for Automatic Image Enhancement. Journal of Electronic Imaging, 13(1):100–111, 2004.
  38. 38.O. Ronneberger, P. Fischer, and T. Brox. U-Net: Convolutional Networks for Biomedical Image Segmentation. In International Conference on Medical Image Computing and Computer-assisted Intervention, pages 234–241. Springer, 2015.
  39. 39.F. Shkurti, A. Xu, M. Meghjani, J. C. G. Higuera, Y. Girdhar, Giguere, et al. Multi-domain Monitoring of Marine Environments using a Heterogeneous Robot Team. In IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 1747–1753. IEEE, 2012.
  40. 40.W. Wang, J. Shen, X. Dong, and A. Borji. Salient Object Detection Driven by Fixation Prediction. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 1711–1720, 2018.
  41. 41.Z. Wang, A. C. Bovik, H. R. Sheikh, E. P. Simoncelli, et al. Image Quality Assessment: from Error Visibility to Structural Similarity. IEEE Transactions on Image Processing, 13(4):600–612, 2004.
  42. 42.Z. Yi, H. Zhang, P. Tan, and M. Gong. DualGAN: Unsupervised Dual Learning for Image-to-image Translation. In IEEE International Conference on Computer Vision (ICCV), pages 2849–2857, 2017.
  43. 43.X. Yu, Y. Qu, and M. Hong. Underwater-GAN: Underwater Image Restoration via Conditional Generative Adversarial Network. In International Conference on Pattern Recognition, pages 66–75. Springer, 2018.
  44. 44.R. Zhang, P. Isola, and A. A. Efros. Colorful Image Colorization. In European Conference on Computer Vision (ECCV), pages 649–666. Springer, 2016.
  45. 45.S. Zhang, T. Wang, J. Dong, and H. Yu. Underwater Image Enhancement via Extended Multi-scale Retinex. Neurocomputing, 245:1–9, 2017.
  46. 46.J. Zhao, M. Mathieu, and Y. LeCun. Energy-based Generative Adversarial Network. 2017.
  47. 47.J.-Y. Zhu, T. Park, P. Isola, and A. A. Efros. Unpaired Image-to-image Translation using Cycle-consistent Adversarial Networks. In IEEE International Conference on Computer Vision (ICCV), pages 2223–2232, 2017.

Citation

MLA
Islam, M. J., et al. “Fast Underwater Image Enhancement for Improved Visual Perception”. arXiv, 2019, http://arxiv.org/abs/1903.09766v3.
APA
Islam, M. J., Xia, Y., & Sattar, J. (2019). Fast Underwater Image Enhancement for Improved Visual Perception. arXiv. http://arxiv.org/abs/1903.09766v3
Chicago
Islam, M. J., Y. Xia, and J. Sattar. 2019. “Fast Underwater Image Enhancement for Improved Visual Perception”. arXiv. http://arxiv.org/abs/1903.09766v3.
Harvard
Islam, M.J., Xia, Y. and Sattar, J. (2019) “Fast Underwater Image Enhancement for Improved Visual Perception”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1903.09766v3.
Vancouver
1. Islam MJ, Xia Y, Sattar J (2019) Fast Underwater Image Enhancement for Improved Visual Perception. arXiv

BibTeX

@article{islam2019fast,
  title = {Fast Underwater Image Enhancement for Improved Visual Perception},
  author = {Islam, Md Jahidul and Xia, Youya and Sattar, Junaed},
  year = {2019},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1903.09766v3},
  eprint = {1903.09766}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF