Deep Multi-scale Convolutional Neural Network for Dynamic Scene Deblurring
Seungjun NahTae Hyun KimKyoung Mu Lee
Presents an end-to-end multi-scale deep network paired with a realistic high-speed camera dataset, shifting dynamic scene deblurring from restrictive blur kernel estimations toward direct coarse-to-fine image restoration.
Motion blur from camera shake and fast-moving objects in dynamic scenes remains a persistent challenge in photography, as it produces complex, spatially varying artifacts that degrade image quality. Conventional deblurring methods struggle because they depend on simplified assumptions about blur kernels, such as uniform or locally linear motion, which fail at object boundaries, occlusions, and depth changes, often introducing ringing artifacts.
The article set out to develop an end-to-end method that restores sharp images directly from blurry inputs without estimating explicit blur kernels, while also creating a more realistic training dataset to support such learning. Researchers built a multi-scale convolutional neural network that processes images at multiple resolutions in a coarse-to-fine manner, trained it using a combination of multi-scale content loss and adversarial loss on pairs of blurry and sharp images, and generated a new dataset of 3,214 image pairs by averaging sequences of frames captured at 240 frames per second with a high-speed camera.
On the authors' GOPRO test set the approach achieved roughly 4–5 dB higher PSNR and substantially better SSIM scores than prior leading methods while running in a few seconds rather than minutes or hours; similar gains appeared on the Köhler and Lai datasets, with visibly cleaner object boundaries and fewer artifacts on both synthetic and real dynamic scenes. These results indicate that kernel-free, data-driven restoration can handle the full range of real-world motion blur sources more reliably than optimization-based techniques that rely on kernel estimation.
The work demonstrates that high-quality dynamic scene deblurring is now feasible in practical time frames, which could improve downstream tasks such as object recognition, surveillance, and consumer photo editing. Because the model learns directly from realistic blur examples, it avoids the artifacts that arise when kernels are misspecified.
Further gains may come from expanding the training data to additional camera types and lighting conditions, or from integrating the network into video pipelines. The main limitations are that performance still depends on the distribution of the training set and that quantitative metrics such as PSNR do not always match human perception of sharpness; readers should verify results on their own imagery before deployment at scale.
No sufficiently relevant recommendations were found.
- Paper: Scale-Recurrent Network for Deep Image Deblurring, Xin Tao et al. (2018). Building on the source’s coarse-to-fine deblurring setup and GoPro benchmark, SRN shares weights across scales and passes structural information between them recurrently.
- Paper: DeblurGAN: Blind Motion Deblurring Using Conditional Adversarial Networks, Orest Kupyn et al. (2017). DeblurGAN continues the source’s kernel-free, GoPro-based deblurring approach by adding adversarial training to favor sharper, more perceptually realistic restorations.
- Paper: Multi-Stage Progressive Image Restoration, Syed Waqas Zamir et al. (2021). MPRNet extends learned motion deblurring into a progressive restoration framework, combining multi-scale context with staged refinement and feature fusion.
- Paper: EDVR: Video Restoration With Enhanced Deformable Convolutional Networks, Xintao Wang et al. (2019). EDVR carries deblurring into video restoration, aligning and fusing neighboring frames to handle motion and blur beyond single-image inputs.
- Paper: Flow-Guided Sparse Transformer for Video Deblurring, Jing Lin et al. (2022). FGST continues deblurring on the GoPro benchmark in the video setting, using optical-flow-guided sparse attention to exploit information across frames.
