Dual-Domain Attention for Image Deblurring
Yuning CuiYi TaoWenqi RenAlois Knoll
Proposes a dual-domain attention network that pairs dynamic group convolution for localized spatial self-attention with a lightweight frequency-decoupling module, achieving state-of-the-art image deblurring quality with substantially faster inference speeds.
Motion blur caused by camera shake or moving objects degrades visual clarity across critical applications, including autonomous driving, medical imaging, remote sensing, and digital photography. While deep learning methods have significantly advanced blind image deblurring, existing approaches face major practical hurdles: transformer-based models achieve high restoration quality but suffer from excessive computational complexity and slow processing speeds, while conventional convolutional architectures struggle to capture essential spatial relationships and often neglect valuable frequency-domain information.
The article develops and evaluates the Dual-Domain Attention Network (DDANet), a framework designed to bridge the structural and detail gaps between blurry and sharp images across both spatial and frequency domains simultaneously while drastically reducing processing latency.
To achieve this, the authors designed two complementary, lightweight components integrated into a hierarchical, multi-scale network architecture. The spatial attention module formulates self-attention in the style of dynamic group convolution, restricting information exchange to local regions to cut computational overhead while applying a hyperbolic tangent activation function to actively suppress harmful or irrelevant pixel data. Complementing this, the frequency attention module uses multi-scale average pooling to cleanly separate low- and high-frequency components without requiring computationally heavy transforms, directly learning weights to accentuate informative fine details. The overall architecture was trained and tested on standard benchmark datasets, including GoPro (2,103 training pairs and 1,111 evaluation pairs), and further tested on synthetic and real-world benchmarks without task-specific fine-tuning.
Empirical evaluation demonstrates several key outcomes. First, DDANet achieves an inference speed roughly five times faster than leading Transformer-based models, processing high-definition images in approximately 0.25 seconds compared to over 1.2 seconds for competing models. Second, this speedup is achieved alongside higher restoration quality, outperforming state-of-the-art architectures on the primary benchmark with a peak signal-to-noise ratio of 33.07 dB. Third, the model demonstrated superior parameter efficiency, utilizing approximately 38% fewer parameters (16.18 million versus 26.13 million) compared to leading alternatives. Finally, DDANet generalized robustly to unseen synthetic and real-world datasets without retraining, consistently matching or exceeding specialized methods.
These findings indicate that image restoration networks do not require quadratic-complexity attention mechanisms or cumbersome frequency transformations to achieve high performance. By blending the efficiency of local convolutions with targeted attention, organizations can deploy high-quality deblurring systems in compute-constrained and time-sensitive operational environments, substantially lowering processing costs and hardware requirements for downstream computer vision systems.
Technical leaders and engineering teams should consider adopting dual-domain local attention strategies when architecting vision pipelines for real-time edge devices or automated platforms like autonomous vehicles. Prior to production rollout, teams should validate DDANet on domain-specific camera hardware and edge computing platforms to evaluate exact throughput under varying real-world lighting and motion conditions.
Confidence in these findings is high given the thorough ablation studies and cross-dataset evaluations; however, slight performance variations may occur when deploying across edge hardware architectures not tested in the study, and performance bounds on extreme out-of-distribution motion artifacts remain an area for ongoing testing.
- Paper: Deep Multi-scale Convolutional Neural Network for Dynamic Scene Deblurring, Seungjun Nah et al. (2017). This seminal paper introduces multi-scale deep learning for dynamic scene deblurring along with the standard GoPro benchmark dataset evaluated by DDANet.
- Paper: Scale-Recurrent Network for Deep Image Deblurring, Xin Tao et al. (2018). It establishes coarse-to-fine parameter-efficient scale-recurrent architectures for deblurring that motivate DDANet's multi-scale design.
- Paper: DeblurGAN: Blind Motion Deblurring Using Conditional Adversarial Networks, Orest Kupyn et al. (2017). This work pioneered blind motion deblurring without explicit kernel estimation, forming the foundation of modern deep restoration baselines.
- Paper: Restormer: Efficient Transformer for High-Resolution Image Restoration, Syed Waqas Zamir et al. (2022). It presents Restormer, the leading Transformer-based restoration baseline whose high computational complexity DDANet explicitly aims to surpass in speed and parameter efficiency.
- Paper: Simple Baselines for Image Restoration, Liangyu Chen et al. (2022). It introduces NAFNet, establishing a crucial lightweight baseline for image deblurring that investigates simple architectural replacements for complex attention.
- Paper: Multi-Stage Progressive Image Restoration, Syed Waqas Zamir et al. (2021). It details multi-stage progressive restoration and supervised attention mechanisms that provide key context for hierarchical deblurring network designs.
- Paper: Uformer: A General U-Shaped Transformer for Image Restoration, Zhendong Wang et al. (2021). It provides a core benchmark for locally windowed Transformer architectures in image restoration that DDANet seeks to optimize via local dynamic convolution.
- Paper: Dynamic Convolution: Attention Over Convolution Kernels, Yinpeng Chen et al. (2019). This paper establishes dynamic convolution via attention over kernels, which underpins the dynamic group convolution design in DDANet's spatial attention module.
- Paper: Fourier Priors-Guided Diffusion for Zero-Shot Joint Low-Light Enhancement and Deblurring, Xiaoqian Lv et al. (2024). It extends frequency-domain deblurring principles to a zero-shot diffusion framework that jointly tackles severe motion blur and low-light enhancement.
- Paper: Frequency-Adaptive Dilated Convolution for Semantic Segmentation, Linwei Chen et al. (2024). It generalizes spatial frequency decomposition concepts into frequency-adaptive convolutions for higher-level downstream vision tasks like semantic segmentation.
