Towards Robust Event-guided Low-Light Image Enhancement: A Large-Scale Real-World Event-Image Dataset and Novel Approach

Guoqiang LiangKanghao ChenHangyu LiYunfan LuLin Wang

article2024CVPR78 citations

Presents a large-scale, robot-aligned real-world event-image dataset alongside EvLight, a framework that uses signal-to-noise ratio guidance to selectively combine event edges and frame features for low-light image restoration.

Listen

Capturing clear visual data in poor lighting conditions is a fundamental challenge for downstream artificial intelligence systems, such as nighttime autonomous navigation, surveillance, and facial recognition. While event cameras—sensors that record brightness changes at high speeds with high dynamic range—offer rich structural and edge information, combining them with standard color cameras has been hindered by a critical data shortage. Prior research lacked large-scale, real-world datasets that accurately align dynamic images and event streams in both space and time, forcing existing methods to rely on synthetic data, static scenes, or uncalibrated inputs.

The article addresses this gap with two main objectives: first, creating a high-precision, real-world dataset of paired low-light and normal-light image-event sequences under complex dynamic motion; and second, developing an event-guided low-light image enhancement framework, named EvLight, that selectively merges image and event features based on local signal quality.

To construct the dataset, the researchers deployed an event camera on a high-precision robotic arm following complex non-linear trajectories across indoor and outdoor environments. They gathered over 30,000 aligned pairs across 91 sequences by recording identical paths under normal lighting and low-light conditions using optical filters. To overcome mechanical and timing variations, they implemented a matching strategy that brought temporal alignment error below 0.01 seconds for 90% of the data and achieved spatial error margins under 0.03 millimeters. Leveraging this dataset, the team designed the EvLight framework. The system uses a signal-to-noise ratio map to extract clean color features from well-lit image regions while selectively drawing boundary and structural details from event streams in dark, high-noise areas, combining them through a holistic attention-based fusion network.

The findings demonstrate clear performance advantages. Across indoor and outdoor benchmarks on the newly collected dataset, EvLight consistently outperformed leading frame-based and event-guided enhancement models, delivering higher structural similarity and visual quality while avoiding common artifacts like over-saturation and color distortion. In evaluations on a benchmark dynamic dataset, EvLight outperformed the previous state-of-the-art frame-based method by 2.62 decibels in reconstruction accuracy (peak signal-to-noise ratio) and exceeded alternative event-guided approaches by around 0.9 decibels indoors. Ablation tests confirmed that the signal-to-noise ratio feature selection is essential, yielding a 0.86-decibel performance gain over basic fusion setups that suffer from noise accumulation.

These results demonstrate that event cameras can significantly enhance low-light computer vision when guided by region-specific noise awareness. For operational systems operating in dark or high-contrast environments, this approach offers a viable path toward robust visibility without severe computational distortion or sensor noise amplification. Moving forward, teams deploying low-light enhancement should consider incorporating event sensors alongside selective fusion mechanisms. Future development should focus on upgrading hardware to achieve automated hardware-level synchronization between cameras and robotic systems, eliminating manual capture overhead and addressing minor sensor artifacts such as chromatic aberrations.

arXiv: 2404.00834
  • Paper: SNR-Aware Low-light Image Enhancement, Xiaogang Xu et al. (2022). It introduces SNR-guided feature modulation and dual-branch processing for low-light enhancement, establishing the conceptual foundation for EvLight's SNR-aware image and event feature selection.
  • Paper: Event-Based Vision: A Survey, Guillermo Gallego et al. (2019). It provides a comprehensive overview of event cameras, event stream representations, and neuromorphic visual sensing principles essential for understanding event-guided visual enhancement.
  • Paper: Learning to See in the Dark, Chen Chen et al. (2018). It provides the seminal benchmark and deep learning methodology for extreme low-light raw image restoration, establishing core baseline problems in low-light noise and amplification.
  • Paper: Deep Retinex Decomposition for Low-Light Enhancement, Chen Wei et al. (2018). It establishes paired data collection paradigms and deep Retinex illumination decomposition strategies that serve as standard baselines for low-light restoration.
  • Paper: Zero-Reference Deep Curve Estimation for Low-Light Image Enhancement, Chunle Guo et al. (2020). It introduces non-reference curve estimation for dynamic range adjustment, providing key background on contemporary frame-based low-light enhancement techniques.
  • Paper: U2Fusion: A Unified Unsupervised Image Fusion Network, Han Xu et al. (2020). It outlines fundamental principles for multi-modal feature fusion and information measurement across complementary imaging sensors.
Cover for Towards Robust Event-guided Low-Light Image Enhancement: A Large-Scale Real-World Event-Image Dataset and Novel Approach

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Our SDE Dataset
  • 4. The Proposed EvLight Framework
  • 4.1. Preprocessing
  • 4.2. SNR-guided Regional Feature Selection
  • 4.3. Holistic-Regional Fusion Branch
  • 5. Experiments
  • 5.1. Comparison and Evaluation
  • 5.2. Ablation Studies and Analysis
  • 6. Conclusion
  • References

Knowls

  1. Knowl 1 — SDE real-world event-image dataset

    data/table

    The paper introduces the SDE dataset, a real-world low-light enhancement dataset containing 31,477 spatially and temporally aligned pairs of RGB images and event streams. It covers indoor and outdoor scenes, dynamic scenes with nonlinear camera trajectories, and both low-illumination and normal-illumination conditions. The data were captured with a DAVIS346 event camera, producing RGB frames and events at 346×260346\times260 resolution. The dataset contains 91 paired sequences—43 indoor and 48 outdoor—with 76 sequences used for training and 15 for testing.

    Spatial alignment is obtained by mounting the camera on a Universal UR5 robotic arm with a reported repeated positioning error of 0.03 mm0.03\,\mathrm{mm}. For each scene, normal-light sequences are recorded first, followed by low-light sequences using an ND8 filter while keeping camera parameters such as exposure time and frame intervals fixed. The robot follows the same predefined nonlinear trajectory in both illumination conditions.

  2. Knowl 2 — Matching strategy for temporal alignment

    algorithm

    The SDE temporal-alignment procedure receives three low-light and three normal-light image-event sequences captured for the same scene and outputs one temporally matched low-light/normal-light pair.

    Input: Three low-light sequences, three normal-light sequences, and the predefined robot trajectory start and end timestamps
    Output: One temporally matched low-light/normal-light sequence pair
    Trim every sequence to the trajectory start and end timestamps
    For every trimmed sequence, measure the interval from the trajectory start to the timestamp of its first captured frame
    For every low-light sequence and every normal-light sequence, compute the absolute difference between their measured intervals
    Select the low-light/normal-light pair with the smallest absolute interval difference
    Return the selected pair as the temporally aligned sequence pair

    The procedure compensates for the random delay between robot motion initiation and the first captured frame, which remains variable even when the camera and robot settings are unchanged. The resulting temporal alignment error is below 0.01 s0.01\,\mathrm{s} for 90% of the dataset; the maximum reported errors are 0.013 s0.013\,\mathrm{s} indoors and 0.027 s0.027\,\mathrm{s} outdoors.

  3. Knowl 3 — EvLight preprocessing and event representation

    model/method

    EvLight takes a low-light RGB image II and a stream of NN paired events {ek}k=1N\{e_k\}_{k=1}^{N}, and outputs an enhanced image IenI_{\mathrm{en}}. Each event is assigned according to its polarity to the two closest temporal voxels, producing an event voxel grid EE with 32 temporal bins in all experiments.

    The low-light image is first coarsely lightened. Let Iprior=max⁡c(I)I_{\mathrm{prior}}=\max_c(I) be the per-pixel maximum over the color channels of II, let FF be the illumination-estimation network, and let L=F(I,Iprior)L=F(I,I_{\mathrm{prior}}) be the estimated illumination map. The initial light-up image is

    Ilu=I⊙L,I_{\mathrm{lu}}=I\odot L,

    where ⊙\odot denotes pixelwise multiplication and IluI_{\mathrm{lu}} has the same spatial and color dimensions as II. A 3×33\times3 convolution initially extracts an image feature FimgF_{\mathrm{img}} from IluI_{\mathrm{lu}} and an event feature FevF_{\mathrm{ev}} from EE.

  4. Knowl 4 — SNR-guided regional feature selection

    model/method

    EvLight selects image and event features region by region using an estimated signal-to-noise-ratio map. The light-up image IluI_{\mathrm{lu}} is converted to grayscale Ig∈RH×WI_g\in\mathbb{R}^{H\times W}, and a mean-filtered version I~g\widetilde I_g is used to calculate

    Msnr=I~gabs⁡(Ig−I~g),M_{\mathrm{snr}}=\frac{\widetilde I_g}{\operatorname{abs}(I_g-\widetilde I_g)},

    where HH and WW are image height and width and abs⁡\operatorname{abs} is elementwise absolute value. Image and event features are downsampled to three scales indexed by i∈{0,1,2}i\in\{0,1,2\}, with

    Fimgi,Fevi∈RH22−i×W22−i×22−iC,Msnri∈RH22−i×W22−i,F^i_{\mathrm{img}},F^i_{\mathrm{ev}}\in\mathbb{R}^{\frac{H}{2^{2-i}}\times\frac{W}{2^{2-i}}\times 2^{2-i}C},\qquad M^i_{\mathrm{snr}}\in\mathbb{R}^{\frac{H}{2^{2-i}}\times\frac{W}{2^{2-i}}},

    where CC is the channel count of the full-resolution feature. The feature maps are downsampled using stride-2 4×44\times4 convolutions and the SNR maps using 2×22\times2 average pooling.

    The image-regional feature selection block processes FimgiF^i_{\mathrm{img}} with two residual blocks, each containing two 3×33\times3 convolutions and efficient channel attention. The SNR map is expanded over channels, normalized to [0,1][0,1], and thresholded to produce M^snri\widehat M^i_{\mathrm{snr}}. The selected image feature is

    Fsel−imgi=M^snri⊙F^imgi,F^i_{\mathrm{sel-img}}=\widehat M^i_{\mathrm{snr}}\odot\widehat F^i_{\mathrm{img}},

    where F^imgi\widehat F^i_{\mathrm{img}} is the processed image feature.

    The event-regional feature selection block applies the analogous feature processing to FeviF^i_{\mathrm{ev}}. It uses the complementary mask M‾snri=1−M^snri\overline M^i_{\mathrm{snr}}=1-\widehat M^i_{\mathrm{snr}} and computes

    Fsel−evi=M‾snri⊙F^evi.F^i_{\mathrm{sel-ev}}=\overline M^i_{\mathrm{snr}}\odot\widehat F^i_{\mathrm{ev}}.

    Thus, high-SNR regions preferentially contribute image features, while low-SNR and edge-rich regions preferentially contribute event features. The complementary event selection also suppresses event noise in well-lit smooth regions.

  5. Knowl 5 — Holistic-regional fusion architecture

    model/method

    EvLight combines long-range holistic information with the SNR-selected regional image and event features using a U-Net-like encoder-decoder with skip connections. The encoder contains two stages and the decoder contains three stages. The encoder repeatedly applies holistic feature extraction and stride-2 4×44\times4 convolutional downsampling. The decoder upsamples with stride-2 2×22\times2 deconvolutions and fuses the upsampled holistic feature with Fsel−imgiF^i_{\mathrm{sel-img}} and Fsel−eviF^i_{\mathrm{sel-ev}} at each scale.

    For a holistic feature Fhoi−1F^{i-1}_{\mathrm{ho}}, the holistic feature extraction block uses channel-wise multi-head self-attention, a residual connection, layer normalization, and a feed-forward network:

    F^midi−1=Attention⁡(Fhoi−1)+Fhoi−1,\widehat F^{i-1}_{\mathrm{mid}}=\operatorname{Attention}(F^{i-1}_{\mathrm{ho}})+F^{i-1}_{\mathrm{ho}}, F^hoi−1=FFN⁡ ⁣(LN⁡(F^midi−1))+F^midi−1.\widehat F^{i-1}_{\mathrm{ho}}=\operatorname{FFN}\!\left(\operatorname{LN}(\widehat F^{i-1}_{\mathrm{mid}})\right)+\widehat F^{i-1}_{\mathrm{mid}}.

    Here, Attention⁡\operatorname{Attention} is channel-wise self-attention, LN⁡\operatorname{LN} is layer normalization, and FFN⁡\operatorname{FFN} is a feed-forward network.

    At decoder scale ii, the holistic-regional fusion block concatenates the upsampled holistic feature with the selected image and event features to form FcatiF^i_{\mathrm{cat}}. Three convolutional operations F1F_1, F2F_2, and F3F_3 then produce

    Fhoi=F3 ⁣(σ ⁣(F1(Fcati))⊙F2(Fcati)+Fcati),F^i_{\mathrm{ho}}=F_3\!\left(\sigma\!\left(F_1(F^i_{\mathrm{cat}})\right)\odot F_2(F^i_{\mathrm{cat}})+F^i_{\mathrm{cat}}\right),

    where σ\sigma is the elementwise sigmoid function and ⊙\odot is elementwise multiplication. The decoder output is transformed into the enhanced RGB image IenI_{\mathrm{en}}.

  6. Knowl 6 — Training objective for EvLight

    equation

    EvLight is trained against a ground-truth normal-light image using a combination of a robust pixel reconstruction term and an AlexNet feature-space term. Let IenI_{\mathrm{en}} be the predicted enhanced image, IgtI_{\mathrm{gt}} the ground-truth image, Φ(⋅)\Phi(\cdot) the feature extractor from AlexNet, λ\lambda the weight of the feature loss, and ϵ=10−4\epsilon=10^{-4}. The training loss is

    L=∥Ien−Igt∥22+ϵ2+λ∥Φ(Ien)−Φ(Igt)∥1.\mathcal{L}=\sqrt{\lVert I_{\mathrm{en}}-I_{\mathrm{gt}}\rVert_2^2+\epsilon^2}+\lambda\left\lVert\Phi(I_{\mathrm{en}})-\Phi(I_{\mathrm{gt}})\right\rVert_1.

    The first term encourages pixel-level agreement while remaining numerically robust near zero residual error; the second term encourages perceptual feature agreement.

  7. Knowl 7 — Training and evaluation protocol

    experimental setup

    EvLight is optimized with Adam for 80 epochs using batch size 8 on an NVIDIA A30 GPU. The learning rate is 10−410^{-4} on the SDE dataset and 2×10−42\times10^{-4} on the SDSD dataset. Training augmentation consists of random 256×256256\times256 crops, horizontal flips, and rotations by 90∘90^\circ, 180∘180^\circ, or 270∘270^\circ.

    For SDSD evaluation, the dynamic version of the dataset is used: its original 1920×10801920\times1080 videos are resized to 346×260346\times260, and noisy event streams are synthesized with the default noise model of the v2e event simulator. The split contains 125 training sequences and 25 test sequences.

    Performance is measured with PSNR and SSIM. The paper also reports PSNR* to reduce the effect of global brightness mismatch. Let Gray⁡(⋅)\operatorname{Gray}(\cdot) convert an RGB image to grayscale, let Mean⁡(⋅)\operatorname{Mean}(\cdot) compute the mean grayscale intensity, and let PSNR⁡(⋅,⋅)\operatorname{PSNR}(\cdot,\cdot) be peak signal-to-noise ratio. Define

    Rgt−en=Mean⁡(Gray⁡(Igt))Mean⁡(Gray⁡(Ien)).R_{\mathrm{gt-en}}=\frac{\operatorname{Mean}(\operatorname{Gray}(I_{\mathrm{gt}}))}{\operatorname{Mean}(\operatorname{Gray}(I_{\mathrm{en}}))}.

    The brightness-adjusted metric is

    PSNR⁡∗=PSNR⁡(Ien×Rgt−en,Igt),\operatorname{PSNR}^{*}=\operatorname{PSNR}(I_{\mathrm{en}}\times R_{\mathrm{gt-en}},I_{\mathrm{gt}}),

    where the scalar Rgt−enR_{\mathrm{gt-en}} multiplies every pixel and channel of IenI_{\mathrm{en}}.

  8. Knowl 8 — Benchmark performance on real and synthetic-event data

    data/table

    EvLight is compared with an event-only method, four image-only methods, and three image-plus-event methods on the SDE and SDSD test sets. The in and out columns are the paper's indoor and outdoor evaluation splits. PSNR and PSNR* are in dB, and SSIM is dimensionless. E2VID+ is evaluated in grayscale because it reconstructs grayscale images only.

    The complete quantitative comparison is:

    Could not parse LaTeX table

    EvLight has the best value for every reported metric and split. Relative to the strongest competing method, its PSNR advantage is 0.65 dB0.65\,\mathrm{dB} on SDE-in and 0.29 dB0.29\,\mathrm{dB} on SDE-out; its PSNR* advantage is 0.93 dB0.93\,\mathrm{dB} and 1.21 dB1.21\,\mathrm{dB} on those splits. On SDSD, the PSNR advantage is 0.94 dB0.94\,\mathrm{dB} on SDSD-in and 0.59 dB0.59\,\mathrm{dB} on SDSD-out. Qualitative comparisons show that EvLight recovers clearer edges in dark regions while producing less color distortion and noise than the compared image-only and event-guided methods.

  9. Knowl 9 — Ablation evidence for events and regional selection

    empirical result

    Ablations on the SDE indoor test split show that both event input and SNR-guided regional selection are important. In the base EvLight architecture without regional selection, removing events gives PSNR 21.35 dB21.35\,\mathrm{dB} and SSIM 0.69850.6985; adding events improves PSNR by 0.23 dB0.23\,\mathrm{dB} and SSIM by 0.0020.002.

    Replacing the SNR map with an all-ones mask tests whether regional selection alone is useful, while removing the entire selection module defines the base model. The results are:

    Could not parse LaTeX table

    Relative to the base model, generic regional selection improves PSNR by 0.28 dB0.28\,\mathrm{dB}, whereas SNR-guided selection improves it by 0.86 dB0.86\,\mathrm{dB}. The separate contributions of image-regional feature selection (IRFS) and event-regional feature selection (ERFS) are:

    Could not parse LaTeX table

    IRFS alone raises PSNR by 0.34 dB0.34\,\mathrm{dB}, ERFS alone by 0.60 dB0.60\,\mathrm{dB}, and their combination by 0.86 dB0.86\,\mathrm{dB} relative to the base model. The visual ablations indicate that either block reduces color distortion, while using both produces the best structural and color quality.

  10. Knowl 10 — Dataset and method limitations

    limitation

    The SDE dataset is collected with a DAVIS346 event camera, whose RGB output can exhibit partial chromatic aberrations and moiré patterns. The collection system also requires repeated manual capture procedures; the authors identify synchronous triggering between the robotic arm and event camera as future hardware work intended to reduce this labor cost.

Coverage note — The paper's cross-dataset generalization experiments on CED, MVSEC, and real SDE events are not given numerical results in the paper text and are therefore not made into a separate knowl; the comparative survey of prior datasets is background rather than a contribution.

References

  1. 1.Tarik Arici, Salih Dikbas, and Yucel Altunbasak. A histogram modification framework and its application for image contrast enhancement. IEEE Transactions on image processing, 18(9):1921–1935, 2009. 3
  2. 2.Antoni Buades, Bartomeu Coll, and J-M Morel. A non-local algorithm for image denoising. In 2005 IEEE computer society conference on computer vision and pattern recognition (CVPR’05), pages 60–65. Ieee, 2005. 4
  3. 3.Vladimir Bychkovsky, Sylvain Paris, Eric Chan, and Fredo Durand. Learning photographic global tonal adjustment with a database of input/output image pairs. In CVPR 2011, pages 97–104. IEEE, 2011. 3
  4. 4.Yuanhao Cai, Hao Bian, Jing Lin, Haoqian Wang, Radu Timofte, and Yulun Zhang. Retinexformer: One-stage retinex-based transformer for low-light image enhancement. arXiv preprint arXiv:2303.06705, 2023. 1, 2, 3, 4, 6, 7
  5. 5.Chen Chen, Qifeng Chen, Jia Xu, and Vladlen Koltun. Learning to see in the dark. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 3291–3300, 2018. 3
  6. 6.Chen Chen, Qifeng Chen, Minh N Do, and Vladlen Koltun. Seeing motion in the dark. In Proceedings of the IEEE/CVF International conference on computer vision, pages 3185–3194, 2019. 2, 3
  7. 7.Wenhan Yang Jiaying Liu Chen Wei, Wenjing Wang. Deep retinex decomposition for low-light enhancement. In British Machine Vision Conference, 2018. 3
  8. 8.Kostadin Dabov, Alessandro Foi, Vladimir Katkovnik, and Karen Egiazarian. Image denoising with block-matching and 3d filtering. In Image processing: algorithms and systems, neural networks, and machine learning, pages 354–365. SPIE, 2006. 4
  9. 9.Huiyuan Fu, Wenkai Zheng, Xiangyu Meng, Xin Wang, Chuanming Wang, and Huadong Ma. You do not need additional priors or regularizers in retinex-based low-light image enhancement. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 18125–18134, 2023. 3
  10. 10.Huiyuan Fu, Wenkai Zheng, Xicong Wang, Jiaxuan Wang, Heng Zhang, and Huadong Ma. Dancing in the dark: A benchmark towards general low-light video enhancement. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 12877–12886, 2023. 2, 3
  11. 11.Xueyang Fu, Delu Zeng, Yue Huang, Xiao-Ping Zhang, and Xinghao Ding. A weighted variational model for simultaneous reflectance and illumination estimation. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2782–2790, 2016. 3
  12. 12.Xiaojie Guo, Yu Li, and Haibin Ling. Lime: Low-light image enhancement via illumination map estimation. IEEE Transactions on image processing, 26(2):982–993, 2016. 3
  13. 13.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016. 5
  14. 14.Alain Hore and Djemel Ziou. Image quality metrics: Psnr vs. ssim. In 2010 20th international conference on pattern recognition, pages 2366–2369. IEEE, 2010. 6
  15. 15.Yuhuang Hu, Shih-Chii Liu, and Tobi Delbruck. v2e: From video frames to realistic dvs events. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1312–1321, 2021. 2, 6, 7
  16. 16.Andrey Ignatov, Nikolay Kobyshev, Radu Timofte, Kenneth Vanhoey, and Luc Van Gool. Dslr-quality photos on mobile devices with deep convolutional networks. In Proceedings of the IEEE international conference on computer vision, pages 3277–3285, 2017. 3
  17. 17.Haiyang Jiang and Yinqiang Zheng. Learning to see moving objects in the dark. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 7324–7333, 2019. 3
  18. 18.Yu Jiang, Yuehang Wang, Siqi Li, Yongji Zhang, Minghao Zhao, and Yue Gao. Event-based low-illumination image enhancement. IEEE Transactions on Multimedia, 2023. 2, 3, 6, 7
  19. 19.Haiyan Jin, Qiaobin Wang, Haonan Su, and Zhaolin Xiao. Event-guided low light image enhancement via a dual branch gan. Journal of Visual Communication and Image Representation, 95:103887, 2023. 3
  20. 20.Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014. 6
  21. 21.Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. Imagenet classification with deep convolutional neural networks. Advances in neural information processing systems, 25, 2012. 6
  22. 22.Sohyun Lee, Jaesung Rim, Boseung Jeong, Geonu Kim, Byungju Woo, Haechan Lee, Sunghyun Cho, and Suha Kwak. Human pose estimation in extremely low-light conditions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 704–714, 2023. 3
  23. 23.Chongyi Li, Chunle Guo, Linghao Han, Jun Jiang, Ming-Ming Cheng, Jinwei Gu, and Chen Change Loy. Low-light image and video enhancement using deep learning: A survey. IEEE transactions on pattern analysis and machine intelligence, 44(12):9396–9416, 2021. 1
  24. 24.Jinxiu Liang, Yixin Yang, Boyu Li, Peiqi Duan, Yong Xu, and Boxin Shi. Coherent event guided low-light video enhancement. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 10615–10625, 2023. 2, 3
  25. 25.Lin Liu, Junfeng An, Jianzhuang Liu, Shanxin Yuan, Xiangyu Chen, Wengang Zhou, Houqiang Li, Yan Feng Wang, and Qi Tian. Low-light video enhancement with synthetic event guidance. In Proceedings of the AAAI Conference on Artificial Intelligence, pages 1692–1700, 2023. 2, 3, 6, 7
  26. 26.David G Lowe. Distinctive image features from scale-invariant keypoints. International journal of computer vision, 60:91–110, 2004. 3
  27. 27.Long Ma, Tengyu Ma, Risheng Liu, Xin Fan, and Zhongxuan Luo. Toward fast, flexible, and robust low-light image enhancement. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5637–5646, 2022. 1
  28. 28.Keita Nakai, Yoshikatsu Hoshi, and Akira Taguchi. Color image contrast enhacement method based on differential intensity/saturation gray-levels histograms. In 2013 International Symposium on Intelligent Signal Processing and Communication Systems, pages 445–449. IEEE, 2013. 3
  29. 29.Jingyi Pan, Sihang Li, Yucheng Chen, Jinjing Zhu, and Lin Wang. Towards dynamic and small objects refinement for unsupervised domain adaptative nighttime semantic segmentation. arXiv preprint arXiv:2310.04747, 2023. 1
  30. 30.Henri Rebecq, Rene Ranftl, Vladlen Koltun, and Davide Scaramuzza. Events-to-video: Bringing modern computer vision to event cameras. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3857–3866, 2019. 4
  31. 31.Jae­sung Rim, Haeyun Lee, Jucheol Won, and Sunghyun Cho. Real-world blur dataset for learning and benchmarking deblurring algorithms. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XXV 16, pages 184–201. Springer, 2020. 3
  32. 32.Olaf Ronneberger, Philipp Fischer, and Thomas Brox. U-net: Convolutional networks for biomedical image segmentation. In Medical Image Computing and Computer-Assisted Intervention–MICCAI 2015: 18th International Conference, Munich, Germany, October 5-9, 2015, Proceedings, Part III 18, pages 234–241. Springer, 2015. 5
  33. 33.Cedric Scheerlinck, Henri Rebecq, Timo Stoffregen, Nick Barnes, Robert Mahony, and Davide Scaramuzza. Ced: Color event camera dataset. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, pages 0–0, 2019. 1, 6, 8
  34. 34.Timo Stoffregen, Cedric Scheerlinck, Davide Scaramuzza, Tom Drummond, Nick Barnes, Lindsay Kleeman, and Robert Mahony. Reducing the sim-to-real gap for event cameras. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XXVII 16, pages 534–549. Springer, 2020. 6, 7
  35. 35.Gemma Taverni, Diederik Paul Moeys, Chenghan Li, Celso Cavaco, Vasyl Motsnyi, David San Segundo Bello, and Tobi Delbruck. Front and back illuminated dynamic and active pixel vision sensors comparison. IEEE Transactions on Circuits and Systems II: Express Briefs, 65(5):677–681, 2018. 2
  36. 36.Bishan Wang, Jingwei He, Lei Yu, Gui-Song Xia, and Wen Yang. Event enhanced high-quality image recovery. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XIII 16, pages 155–171. Springer, 2020. 6, 7
  37. 37.Qilong Wang, Banggu Wu, Pengfei Zhu, Peihua Li, Wangmeng Zuo, and Qinghua Hu. Eca-net: Efficient channel attention for deep convolutional neural networks. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 11534–11542, 2020. 5
  38. 38.Ruixing Wang, Qing Zhang, Chi-Wing Fu, Xiaoyong Shen, Wei-Shi Zheng, and Jiaya Jia. Underexposed photo enhancement using deep illumination estimation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 6849–6857, 2019. 1, 3
  39. 39.Ruixing Wang, Xiaogang Xu, Chi-Wing Fu, Jiangbo Lu, Bei Yu, and Jiaya Jia. Seeing dynamic scene in the dark: A high-quality video dataset with mechatronic alignment. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 9700–9709, 2021. 2, 3, 6, 7, 8
  40. 40.Wei Wang, Xin Chen, Cheng Yang, Xiang Li, Xuemei Hu, and Tao Yue. Enhancing low light videos by exploring high sensitivity camera noise. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 4111–4119, 2019. 3
  41. 41.Yinglong Wang, Zhen Liu, Jianzhuang Liu, Songcen Xu, and Shuaicheng Liu. Low-light image enhancement with illumination-aware gamma correction and complete image modelling network. arXiv preprint arXiv:2308.08220, 2023. 3, 4
  42. 42.Zhou Wang, Alan C Bovik, Hamid R Sheikh, and Eero P Simoncelli. Image quality assessment: from error visibility to structural similarity. IEEE transactions on image processing, 13(4):600–612, 2004. 6
  43. 43.Zhendong Wang, Xiaodong Cun, Jianmin Bao, Wengang Zhou, Jianzhuang Liu, and Houqiang Li. Uformer: A general u-shaped transformer for image restoration. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 17683–17693, 2022. 6, 7
  44. 44.Wenhui Wu, Jian Weng, Pingping Zhang, Xu Wang, Wenhan Yang, and Jianmin Jiang. Uretinex-net: Retinex-based deep unfolding network for low-light image enhancement. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 5901–5910, 2022. 3
  45. 45.Yuhui Wu, Chen Pan, Guoqing Wang, Yang Yang, Jiwei Wei, Chongyi Li, and Heng Tao Shen. Learning semantic-aware knowledge guidance for low-light image enhancement. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1662–1671, 2023. 3, 6, 7
  46. 46.Jun Xu, Yingkun Hou, Dongwei Ren, Li Liu, Fan Zhu, Mengyang Yu, Haoqian Wang, and Ling Shao. Star: A structure and texture aware retinex model. IEEE Transactions on Image Processing, 29:5022–5037, 2020. 3
  47. 47.Ke Xu, Xin Yang, Baocai Yin, and Rynson WH Lau. Learning to restore low-light images via decomposition-and-enhancement. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 2281–2290, 2020. 1
  48. 48.Xiaogang Xu, Ruixing Wang, Chi-Wing Fu, and Jiaya Jia. Snr-aware low-light image enhancement. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 17714–17724, 2022. 3, 4, 5, 6, 7
  49. 49.Xiaogang Xu, Ruixing Wang, and Jiangbo Lu. Low-light image enhancement via structure modeling and guidance. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9893–9903, 2023. 3, 4
  50. 50.Jun Yu, Xinlong Hao, and Peng He. Single-stage face detection under extremely low-light conditions. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 3523–3532, 2021. 1
  51. 51.Syed Waqas Zamir, Aditya Arora, Salman Khan, Munawar Hayat, Fahad Shahbaz Khan, and Ming-Hsuan Yang. Restormer: Efficient transformer for high-resolution image restoration. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 5728–5739, 2022. 6
  52. 52.Song Zhang, Yu Zhang, Zhe Jiang, Dongqing Zou, Jimmy Ren, and Bin Zhou. Learning to see in the dark with events. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XVIII 16, pages 666–682. Springer, 2020. 2, 3
  53. 53.Yonghua Zhang, Jiawan Zhang, and Xiaojie Guo. Kindling the darkness: A practical low-light image enhancer. In Proceedings of the 27th ACM international conference on multimedia, pages 1632–1640, 2019. 3, 6
  54. 54.Yonghua Zhang, Xiaojie Guo, Jiayi Ma, Wei Liu, and Jiawan Zhang. Beyond brightening low-light images. International Journal of Computer Vision, 129:1013–1037, 2021. 1, 3
  55. 55.Xu Zheng, Yexin Liu, Yunfan Lu, Tongyan Hua, Tianbo Pan, Weiming Zhang, Dacheng Tao, and Lin Wang. Deep learning for event-based vision: A comprehensive survey and benchmarks. arXiv preprint arXiv:2302.08890, 2023. 1, 3
  56. 56.Chu Zhou, Minggui Teng, Jin Han, Chao Xu, and Boxin Shi. Delieve-net: Deblurring low-light images with light streaks and local events. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 1155–1164, 2021. 3
  57. 57.Alex Zihao Zhu, Dinesh Thakur, Tolga Ozaslan, Bernd Pfrommer, Vijay Kumar, and Kostas Daniilidis. The multi-vehicle stereo event camera dataset: An event camera dataset for 3d perception. IEEE Robotics and Automation Letters, 3 (3):2032–2039, 2018. 8

Citation

MLA
Liang, G., et al. “Towards Robust Event-guided Low-Light Image Enhancement: A Large-Scale Real-World Event-Image Dataset and Novel Approach”. arXiv, 2024, http://arxiv.org/abs/2404.00834v1.
APA
Liang, G., Chen, K., Li, H., Lu, Y., & Wang, L. (2024). Towards Robust Event-guided Low-Light Image Enhancement: A Large-Scale Real-World Event-Image Dataset and Novel Approach. arXiv. http://arxiv.org/abs/2404.00834v1
Chicago
Liang, G., K. Chen, H. Li, Y. Lu, and L. Wang. 2024. “Towards Robust Event-guided Low-Light Image Enhancement: A Large-Scale Real-World Event-Image Dataset and Novel Approach”. arXiv. http://arxiv.org/abs/2404.00834v1.
Harvard
Liang, G. et al. (2024) “Towards Robust Event-guided Low-Light Image Enhancement: A Large-Scale Real-World Event-Image Dataset and Novel Approach”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2404.00834v1.
Vancouver
1. Liang G, Chen K, Li H, Lu Y, Wang L (2024) Towards Robust Event-guided Low-Light Image Enhancement: A Large-Scale Real-World Event-Image Dataset and Novel Approach. arXiv

BibTeX

@article{liang2024towards,
  title = {Towards Robust Event-guided Low-Light Image Enhancement: A Large-Scale Real-World Event-Image Dataset and Novel Approach},
  author = {Liang, Guoqiang and Chen, Kanghao and Li, Hangyu and Lu, Yunfan and Wang, Lin},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2404.00834v1},
  eprint = {2404.00834}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/