Dual-Domain Attention for Image Deblurring

Yuning CuiYi TaoWenqi RenAlois Knoll

article2023AAAI81 citations

Proposes a dual-domain attention network that pairs dynamic group convolution for localized spatial self-attention with a lightweight frequency-decoupling module, achieving state-of-the-art image deblurring quality with substantially faster inference speeds.

Listen

Motion blur caused by camera shake or moving objects degrades visual clarity across critical applications, including autonomous driving, medical imaging, remote sensing, and digital photography. While deep learning methods have significantly advanced blind image deblurring, existing approaches face major practical hurdles: transformer-based models achieve high restoration quality but suffer from excessive computational complexity and slow processing speeds, while conventional convolutional architectures struggle to capture essential spatial relationships and often neglect valuable frequency-domain information.

The article develops and evaluates the Dual-Domain Attention Network (DDANet), a framework designed to bridge the structural and detail gaps between blurry and sharp images across both spatial and frequency domains simultaneously while drastically reducing processing latency.

To achieve this, the authors designed two complementary, lightweight components integrated into a hierarchical, multi-scale network architecture. The spatial attention module formulates self-attention in the style of dynamic group convolution, restricting information exchange to local regions to cut computational overhead while applying a hyperbolic tangent activation function to actively suppress harmful or irrelevant pixel data. Complementing this, the frequency attention module uses multi-scale average pooling to cleanly separate low- and high-frequency components without requiring computationally heavy transforms, directly learning weights to accentuate informative fine details. The overall architecture was trained and tested on standard benchmark datasets, including GoPro (2,103 training pairs and 1,111 evaluation pairs), and further tested on synthetic and real-world benchmarks without task-specific fine-tuning.

Empirical evaluation demonstrates several key outcomes. First, DDANet achieves an inference speed roughly five times faster than leading Transformer-based models, processing high-definition images in approximately 0.25 seconds compared to over 1.2 seconds for competing models. Second, this speedup is achieved alongside higher restoration quality, outperforming state-of-the-art architectures on the primary benchmark with a peak signal-to-noise ratio of 33.07 dB. Third, the model demonstrated superior parameter efficiency, utilizing approximately 38% fewer parameters (16.18 million versus 26.13 million) compared to leading alternatives. Finally, DDANet generalized robustly to unseen synthetic and real-world datasets without retraining, consistently matching or exceeding specialized methods.

These findings indicate that image restoration networks do not require quadratic-complexity attention mechanisms or cumbersome frequency transformations to achieve high performance. By blending the efficiency of local convolutions with targeted attention, organizations can deploy high-quality deblurring systems in compute-constrained and time-sensitive operational environments, substantially lowering processing costs and hardware requirements for downstream computer vision systems.

Technical leaders and engineering teams should consider adopting dual-domain local attention strategies when architecting vision pipelines for real-time edge devices or automated platforms like autonomous vehicles. Prior to production rollout, teams should validate DDANet on domain-specific camera hardware and edge computing platforms to evaluate exact throughput under varying real-world lighting and motion conditions.

Confidence in these findings is high given the thorough ablation studies and cross-dataset evaluations; however, slight performance variations may occur when deploying across edge hardware architectures not tested in the study, and performance bounds on extreme out-of-distribution motion artifacts remain an area for ongoing testing.

Cover for Dual-Domain Attention for Image Deblurring

Abstract

As a long-standing and challenging task, image deblurring aims to reconstruct the latent sharp image from its degraded counterpart. In this study, to bridge the gaps between degraded/sharp image pairs in the spatial and frequency domains simultaneously, we develop the dual-domain attention mechanism for image deblurring. Self-attention is widely used in vision tasks, however, due to the quadratic complexity, it is not applicable to image deblurring with high-resolution images. To alleviate this issue, we propose a novel spatial attention module by implementing self-attention in the style of dynamic group convolution for integrating information from the local region, enhancing the representation learning capability and reducing computational burden. Regarding frequency domain learning, many frequency-based deblurring approaches either treat the spectrum as a whole or decompose frequency components in a complicated manner. In this work, we devise a frequency attention module to compactly decouple the spectrum into distinct frequency parts and accentuate the informative part with extremely lightweight learnable parameters. Finally, we incorporate attention modules into a U-shaped network. Extensive comparisons with prior arts on the common benchmarks show that our model, named Dual-Domain Attention Network (DDANet), obtains comparable results with a significantly improved inference speed.

Table of Contents

  • Introduction
  • Related Work
  • Proposed Algorithm
  • Spatial Attention Module (SAM)
  • Frequency Attention Module (FAM)
  • Overall Pipeline
  • Loss Function
  • Experiments
  • Datasets and Implementation Details
  • Quantitative and Qualitative Evaluation
  • Ablation Studies
  • Effects of Individual Modules
  • Design Choices for SAM
  • Alternatives to SAM
  • Design Choices for FAM
  • Conclusion
  • References

Knowls

  1. Knowl 1 — Dual-domain U-shaped deblurring network

    model/method

    DDANet is a hierarchical U-shaped image-deblurring network that enhances feature representations in both spatial and frequency domains. An input blurry image first passes through a convolution layer for shallow-feature extraction, followed by three encoder scales and three decoder scales. Each scale contains 20 residual blocks. The encoder progressively reduces spatial resolution while doubling channel width, and the decoder reverses this process.

    Within every scale, the last eight residual blocks use the global branch of the frequency attention module (FAM), the last four use the local branch of FAM, and only the final residual block uses the spatial attention module (SAM). The attention modules are inserted between two 3×33\times3 convolutions in an altered residual block. Feature-level and image-level skip connections are used between corresponding encoder and decoder stages, and a global skip connection adds the degraded input to the network output to produce the final sharp image.

    DDANet uses multi-input and multi-output training, and its input/output arrangement follows the MIMO-UNet design. The final full model uses 16 groups in SAM, a 3×33\times3 SAM kernel, four local FAM placements per scale, and eight global FAM placements per scale.

  2. Knowl 2 — Context-aware local spatial attention module

    model/method

    SAM approximates self-attention with a dynamic group-convolution operation whose integration is local rather than global. For an input feature tensor X∈RC×H×WX\in\mathbb{R}^{C\times H\times W}, global average pooling followed by a convolution produces context-dependent attention kernels:

    W=tanh⁡ ⁣(W1∗GAP⁡(X)),W∈Rg×k2.W=\tanh\!\left(W_1*\operatorname{GAP}(X)\right),\qquad W\in\mathbb{R}^{g\times k^2}.

    Here, C,H,WC,H,W are the channel count, height, and width of XX; gg is the number of channel groups; kk is the local kernel size; W1W_1 is a learnable convolution; ∗* denotes convolution; and each row of WW contains the k2k^2 weights for one group. The hyperbolic tangent allows negative weights, enabling the module to suppress locally harmful features rather than restricting every weight to be positive.

    The input is split into gg channel groups XiX_i, and the same local kernel WiW_i is shared across all channels and spatial positions within group ii:

    Si=Wi∗Xi,i=1,…,g.S_i=W_i*X_i,\qquad i=1,\ldots,g.

    The group outputs are concatenated and passed through a learnable convolution W2W_2 to restore interactions between groups:

    SAM⁡(X)=Concat⁡(S1,S2,…,Sg)W2.\operatorname{SAM}(X)=\operatorname{Concat}(S_1,S_2,\ldots,S_g)W_2.

    The final DDANet configuration uses k=3k=3 and g=16g=16. Compared with full self-attention, SAM obtains weights from contextual information, aggregates only within a small spatial neighborhood, and shares weights within groups, thereby avoiding a global HW×HWHW\times HW attention map and reducing computation and parameters.

  3. Knowl 3 — Lightweight multi-branch frequency attention

    model/method

    FAM separates feature representations into global and local low- and high-frequency components using average pooling rather than Fourier or wavelet transforms. For an input feature tensor X∈RC×H×WX\in\mathbb{R}^{C\times H\times W}, global average pooling produces the lowest-frequency component, and subtraction produces its complementary high-frequency component:

    Xgl=GAP⁡(X),Xgh=X−Xgl.X_g^{l}=\operatorname{GAP}(X),\qquad X_g^{h}=X-X_g^{l}.

    A local low-frequency component is obtained with average pooling using a 3×33\times3 kernel, with its complementary high-frequency component defined as

    Xll=AP⁡3×3(X),Xlh=X−Xll.X_l^{l}=\operatorname{AP}_{3\times3}(X),\qquad X_l^{h}=X-X_l^{l}.

    The pooled components are broadcast to the spatial dimensions of XX. FAM then applies directly learned channel-wise scalars to the four components and adds them:

    FAM⁡(X)=WglXgl+WghXgh+WllXll+WlhXlh.\operatorname{FAM}(X)=W_g^{l}X_g^{l}+W_g^{h}X_g^{h}+W_l^{l}X_l^{l}+W_l^{h}X_l^{h}.

    The four WW terms are learnable channel-wise weights, multiplication is channel-wise, and ll and gg denote local and global branches while superscripts ll and hh denote low- and high-frequency components. No auxiliary attention subnetwork, FFT, wavelet inverse transform, or post-processing convolution is used. The final architecture applies global and local FAM branches at different residual-block depths so that distinct frequency components can be recalibrated independently.

  4. Knowl 4 — Joint spatial- and frequency-domain training objective

    equation

    DDANet trains its multi-scale outputs with an L1 loss in both image and Fourier domains. Let RR be the number of output scales, y^r\hat y_r the predicted image at scale rr, yry_r the corresponding sharp ground-truth image, SrS_r the number of elements in y^r\hat y_r, and F(⋅)\mathcal{F}(\cdot) the Fourier transform. The spatial and frequency losses are

    Lspa=∑r=1R1Sr∥y^r−yr∥1,L_{\mathrm{spa}}=\sum_{r=1}^{R}\frac{1}{S_r}\left\|\hat y_r-y_r\right\|_1, Lfre=∑r=1R1Sr∥F(y^r)−F(yr)∥1.L_{\mathrm{fre}}=\sum_{r=1}^{R}\frac{1}{S_r}\left\|\mathcal{F}(\hat y_r)-\mathcal{F}(y_r)\right\|_1.

    The total objective is

    L=Lspa+λLfre,L=L_{\mathrm{spa}}+\lambda L_{\mathrm{fre}},

    where the paper sets the frequency-loss weight to λ=0.1\lambda=0.1. The normalization by SrS_r makes losses from different output scales comparable.

  5. Knowl 5 — Training and evaluation protocol

    experimental setup

    DDANet is trained on the GoPro dataset, which contains 2,103 blurry/sharp image pairs for training and 1,111 pairs for evaluation. Generalization is tested by applying the GoPro-trained model without fine-tuning to the synthetic HIDE dataset and the real-world RealBlur-R dataset. PSNR and SSIM are the evaluation metrics.

    Training uses Adam with an initial learning rate of 1×10−41\times10^{-4}, cosine annealing to 1×10−61\times10^{-6}, 256×256256\times256 training patches, batch size 4, and 3,000 epochs. Horizontal flips are applied with probability 0.5. Test images are processed at full resolution. Experiments use an NVIDIA Tesla V100 GPU and an Intel Xeon Platinum 8255C CPU; FLOPs are measured on 256×256256\times256 patches.

  6. Knowl 6 — GoPro benchmark accuracy and efficiency trade-off

    data/table

    On the GoPro test set, DDANet achieves the highest reported PSNR and SSIM among the compared methods while using fewer parameters than Restormer and similar computational cost. DDANet reaches 33.07 dB PSNR and 0.962 SSIM, compared with Restormer's 32.92 dB and 0.961, while using 16.18 reported parameter units versus 26.13 for Restormer. The complete comparison is:

    Could not parse LaTeX table

    The table demonstrates that the dual-domain attention design improves restoration quality without incurring the parameter and computation costs of larger transformer-based models.

  7. Knowl 7 — Fast full-resolution inference

    data/table

    DDANet substantially reduces inference time on 720×1280720\times1280 GoPro test images while maintaining the best PSNR among the compared methods. Inference time is measured in seconds per image using released test code and pretrained models on the same hardware:

    Could not parse LaTeX table

    DDANet is almost five times faster than Restormer while improving PSNR by 0.15 dB. For a fair comparison, DeepRFT+ is evaluated without its patch-based testing strategy.

  8. Knowl 8 — Cross-dataset generalization from GoPro training

    data/table

    A DDANet model trained only on GoPro and applied directly without fine-tuning performs strongly on both synthetic and real-world blur benchmarks. On HIDE, DDANet obtains 30.64 dB PSNR and 0.937 SSIM, improving PSNR over MIMO-UNet+ by 0.65 dB. On RealBlur-R, DDANet obtains 35.81 dB and 0.951, improving PSNR over MAXIM-3S by 0.03 dB.

    Could not parse LaTeX table

    These results indicate that the learned dual-domain representation transfers beyond the blur distribution used for training.

  9. Knowl 9 — Individual contributions of SAM and FAM

    empirical result

    Ablation experiments on a smaller baseline network show that both attention modules independently improve deblurring quality with very small computational overhead. The ablation network has eight residual blocks per scale and is trained for 1,900 epochs on GoPro.

    Could not parse LaTeX table

    SAM raises PSNR by 0.40 dB over the baseline, while FAM raises it by 0.75 dB. Using both modules yields a 0.87 dB improvement with only 0.12 additional parameter units and 0.4 additional FLOPs units, demonstrating complementary spatial- and frequency-domain benefits.

  10. Knowl 10 — Ablations select SAM and FAM design settings

    empirical result

    Design ablations support the use of Tanh-normalized, context-aware group attention and four local FAM placements. In an SAM experiment with eight groups and no output convolution, the activation-function results were:

    Could not parse LaTeX table

    With the output convolution omitted, increasing the number of SAM groups improved PSNR but also increased parameters:

    Could not parse LaTeX table

    Because 16 and 32 groups have nearly identical PSNR, the final network uses 16. Replacing SAM with standard group convolution gives 31.52 dB, while SAM without and with its cross-group output convolution gives 31.80 and 31.82 dB, respectively. In a matched comparison of attention variants, MDTA, global self-attention, and SAM obtain 31.34, 31.50, and 31.55 dB, with 6.88, 8.99, and 6.83 parameter units and 67.45, 67.30, and 67.17 FLOPs units, respectively.

    The number of local FAM placements per scale also affects accuracy:

    Could not parse LaTeX table

    The final configuration selects four local filters as an accuracy–complexity compromise.

Coverage note — The qualitative image examples and the plotted layer-wise high-frequency proportions were not made separate knowls because they reinforce, rather than independently extend, the quantitative comparisons and FAM ablations.

References

  1. 1.Aittala, M.; and Durand, F. 2018. Burst image deblurring using permutation invariant convolutional neural networks. In Proceedings of the European conference on computer vision (ECCV), 731–747.
  2. 2.Ayers, G.; and Dainty, J. C. 1988. Iterative blind deconvolution method and its applications. Optics letters, 13(7): 547–549.
  3. 3.Bahat, Y.; Efrat, N.; and Irani, M. 2017. Non-uniform blind deblurring by reblurring. In Proceedings of the IEEE international conference on computer vision, 3286–3294.
  4. 4.Bertero, M.; Boccacci, P.; and De Mol, C. 2021. Introduction to inverse problems in imaging. CRC press.
  5. 5.Campisi, P.; and Egiazarian, K. 2017. Blind image deconvolution: theory and applications. CRC press.
  6. 6.Cao, J.; Li, Y.; Sun, M.; Chen, Y.; Lischinski, D.; Cohen-Or, D.; Chen, B.; and Tu, C. 2022. Do-conv: Depthwise over-parameterized convolutional layer. IEEE Transactions on Image Processing.
  7. 7.Chen, H.; Wang, Y.; Guo, T.; Xu, C.; Deng, Y.; Liu, Z.; Ma, S.; Xu, C.; Xu, C.; and Gao, W. 2021a. Pre-trained image processing transformer. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 12299–12310.
  8. 8.Chen, L.; Chu, X.; Zhang, X.; and Sun, J. 2022. Simple baselines for image restoration. arXiv preprint arXiv:2204.04676.
  9. 9.Chen, L.; Lu, X.; Zhang, J.; Chu, X.; and Chen, C. 2021b. HINet: Half instance normalization network for image restoration. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 182–192.
  10. 10.Chen, Y.; Fan, H.; Xu, B.; Yan, Z.; Kalantidis, Y.; Rohrbach, M.; Yan, S.; and Feng, J. 2019. Drop an octave: Reducing spatial redundancy in convolutional neural networks with octave convolution. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 3435–3444.
  11. 11.Cho, S.-J.; Ji, S.-W.; Hong, J.-P.; Jung, S.-W.; and Ko, S.-J. 2021. Rethinking coarse-to-fine approach in single image deblurring. In Proceedings of the IEEE/CVF international conference on computer vision, 4641–4650.
  12. 12.Cui, Y.; Tao, Y.; Bing, Z.; Ren, W.; Gao, X.; Cao, X.; Huang, K.; and Knoll, A. 2023. Selective Frequency Network for Image Restoration. In The Eleventh International Conference on Learning Representations.
  13. 13.Delbracio, M.; and Sapiro, G. 2015. Burst deblurring: Removing camera shake through fourier burst accumulation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2385–2393.
  14. 14.Keuper, M.; Schmidt, T.; Temerinac-Ott, M.; Padeken, J.; Heun, P.; Ronneberger, O.; and Brox, T. 2013. Blind deconvolution of widefield fluorescence microscopic data by regularization of the optical transfer function (OTF). In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2179–2186.
  15. 15.Kingma, D. P.; and Ba, J. 2014. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980.
  16. 16.Krizhevsky, A.; Sutskever, I.; and Hinton, G. E. 2012. Imagenet classification with deep convolutional neural networks. Advances in neural information processing systems, 25.
  17. 17.Kundur, D.; and Hatzinakos, D. 1996. Blind image deconvolution. IEEE signal processing magazine, 13(3): 43–64.
  18. 18.Lee, H.; Choi, H.; Sohn, K.; and Min, D. 2022. KNN Local Attention for Image Restoration. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2139–2149.
  19. 19.Li, J.; Xia, X.; Li, W.; Li, H.; Wang, X.; Xiao, X.; Wang, R.; Zheng, M.; and Pan, X. 2022a. Next-ViT: Next Generation Vision Transformer for Efficient Deployment in Realistic Industrial Scenarios. arXiv preprint arXiv:2207.05501.
  20. 20.Li, Y.; Wu, C.-Y.; Fan, H.; Mangalam, K.; Xiong, B.; Malik, J.; and Feichtenhofer, C. 2022b. MViTv2: Improved Multi-scale Vision Transformers for Classification and Detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 4804–4814.
  21. 21.Liang, J.; Cao, J.; Sun, G.; Zhang, K.; Van Gool, L.; and Timofte, R. 2021. Swinir: Image restoration using swin transformer. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 1833–1844.
  22. 22.Liu, K.-H.; Yeh, C.-H.; Chung, J.-W.; and Chang, C.-Y. 2020. A motion deblur method based on multi-scale high frequency residual image learning. IEEE Access, 8: 66025–66036.
  23. 23.Liu, Z.; Lin, Y.; Cao, Y.; Hu, H.; Wei, Y.; Zhang, Z.; Lin, S.; and Guo, B. 2021. Swin transformer: Hierarchical vision transformer using shifted windows. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 10012–10022.
  24. 24.Loshchilov, I.; and Hutter, F. 2016. Sgdr: Stochastic gradient descent with warm restarts. arXiv preprint arXiv:1608.03983.
  25. 25.Lucy, L. B. 1974. An iterative technique for the rectification of observed distributions. The astronomical journal, 79: 745.
  26. 26.Mao, X.; Liu, Y.; Shen, W.; Li, Q.; and Wang, Y. 2021. Deep residual fourier transformation for single image deblurring. arXiv preprint arXiv:2111.11745.
  27. 27.Michailovich, O. V.; and Adam, D. 2005. A novel approach to the 2-D blind deconvolution problem in medical ultrasound. IEEE transactions on medical imaging, 24(1): 86–104.
  28. 28.Nah, S.; Hyun Kim, T.; and Mu Lee, K. 2017. Deep multi-scale convolutional neural network for dynamic scene deblurring. In Proceedings of the IEEE conference on computer vision and pattern recognition, 3883–3891.
  29. 29.Purohit, K.; and Rajagopalan, A. 2020. Region-adaptive dense network for efficient motion deblurring. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, 11882–11889.
  30. 30.Purohit, K.; Suin, M.; Rajagopalan, A.; and Boddeti, V. N. 2021. Spatially-adaptive image restoration using distortion-guided networks. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 2309–2319.
  31. 31.Qin, Z.; Zhang, P.; Wu, F.; and Li, X. 2021. Fcanet: Frequency channel attention networks. In Proceedings of the IEEE/CVF international conference on computer vision, 783–792.
  32. 32.Richardson, W. H. 1972. Bayesian-based iterative method of image restoration. JoSA, 62(1): 55–59.
  33. 33.Rim, J.; Lee, H.; Won, J.; and Cho, S. 2020. Real-world blur dataset for learning and benchmarking deblurring algorithms. In European Conference on Computer Vision, 184–201. Springer.
  34. 34.Sayed, M.; and Brostow, G. 2021. Improved handling of motion blur in online object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 1706–1716.
  35. 35.Shen, Z.; Wang, W.; Lu, X.; Shen, J.; Ling, H.; Xu, T.; and Shao, L. 2019. Human-aware motion deblurring. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 5572–5581.
  36. 36.Sroubek, F.; and Flusser, J. 2005. Multichannel blind deconvolution of spatially misaligned images. IEEE Transactions on Image Processing, 14(7): 874–883.
  37. 37.Suin, M.; Purohit, K.; and Rajagopalan, A. 2020. Spatially-attentive patch-hierarchical network for adaptive motion deblurring. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 3606–3615.
  38. 38.Tsai, F.-J.; Peng, Y.-T.; Lin, Y.-Y.; Tsai, C.-C.; and Lin, C.-W. 2022. Stripformer: Strip Transformer for Fast Image Deblurring. arXiv preprint arXiv:2204.04627.
  39. 39.Tu, Z.; Talebi, H.; Zhang, H.; Yang, F.; Milanfar, P.; Bovik, A.; and Li, Y. 2022. Maxim: Multi-axis mlp for image processing. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 5769–5780.
  40. 40.Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017. Attention is all you need. Advances in neural information processing systems, 30.
  41. 41.Wang, R.; and Tao, D. 2014. Recent progress in image deblurring. arXiv preprint arXiv:1409.6838.
  42. 42.Wang, X.; Girshick, R.; Gupta, A.; and He, K. 2018. Non-local neural networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, 7794–7803.
  43. 43.Wang, Z.; Bovik, A. C.; Sheikh, H. R.; and Simoncelli, E. P. 2004. Image quality assessment: from error visibility to structural similarity. IEEE transactions on image processing, 13(4): 600–612.
  44. 44.Wang, Z.; Cun, X.; Bao, J.; Zhou, W.; Liu, J.; and Li, H. 2022. Uformer: A general u-shaped transformer for image restoration. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 17683–17693.
  45. 45.Wiener, N.; Wiener, N.; Mathematician, C.; Wiener, N.; Wiener, N.; and Mathematicien, C. 1949. ´ Extrapolation, interpolation, and smoothing of stationary time series: with engineering applications, volume 113. MIT press Cambridge, MA.
  46. 46.Wu, F.; Fan, A.; Baevski, A.; Dauphin, Y. N.; and Auli, M. 2019. Pay less attention with lightweight and dynamic convolutions. arXiv preprint arXiv:1901.10430.
  47. 47.Yuan, Y.; Su, W.; and Ma, D. 2020. Efficient dynamic scene deblurring using spatially variant deconvolution network with optical flow guided training. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 3555–3564.
  48. 48.Zamir, S. W.; Arora, A.; Khan, S.; Hayat, M.; Khan, F. S.; and Yang, M.-H. 2022. Restormer: Efficient transformer for high-resolution image restoration. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 5728–5739.
  49. 49.Zamir, S. W.; Arora, A.; Khan, S.; Hayat, M.; Khan, F. S.; Yang, M.-H.; and Shao, L. 2021. Multi-stage progressive image restoration. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 14821–14831.
  50. 50.Zhang, K.; Ren, W.; Luo, W.; Lai, W.-S.; Stenger, B.; Yang, M.-H.; and Li, H. 2022a. Deep image deblurring: A survey. International Journal of Computer Vision, 1–28.
  51. 51.Zhang, Y.; Li, Q.; Qi, M.; Liu, D.; Kong, J.; and Wang, J. 2022b. Multi-scale frequency separation network for image deblurring. arXiv preprint arXiv:2206.00798.
  52. 52.Zou, W.; Jiang, M.; Zhang, Y.; Chen, L.; Lu, Z.; and Wu, Y. 2021. Sdwnet: A straight dilated network with wavelet transformation for image deblurring. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 1895–1904.

Citation

MLA
Cui, Y., et al. “Dual-Domain Attention for Image Deblurring”. Proceedings of the AAAI Conference on Artificial Intelligence, vol. 37, no. 1, 2023, pp. 479–87, https://doi.org/10.1609/AAAI.V37I1.25122.
APA
Cui, Y., Tao, Y., Ren, W., & Knoll, A. (2023). Dual-Domain Attention for Image Deblurring. Proceedings of the AAAI Conference on Artificial Intelligence, 37(1), 479–487. https://doi.org/10.1609/AAAI.V37I1.25122
Chicago
Cui, Y., Y. Tao, W. Ren, and A. Knoll. 2023. “Dual-Domain Attention for Image Deblurring”. Proceedings of the AAAI Conference on Artificial Intelligence 37 (1): 479–87. https://doi.org/10.1609/AAAI.V37I1.25122.
Harvard
Cui, Y. et al. (2023) “Dual-Domain Attention for Image Deblurring”, Proceedings of the AAAI Conference on Artificial Intelligence, 37(1), pp. 479–487. Available at: https://doi.org/10.1609/AAAI.V37I1.25122.
Vancouver
1. Cui Y, Tao Y, Ren W, Knoll A (2023) Dual-Domain Attention for Image Deblurring. Proceedings of the AAAI Conference on Artificial Intelligence 37:479–487

BibTeX

@article{Cui_2023, title={Dual-Domain Attention for Image Deblurring}, volume={37}, ISSN={2159-5399}, url={http://dx.doi.org/10.1609/AAAI.V37I1.25122}, DOI={10.1609/aaai.v37i1.25122}, number={1}, journal={Proceedings of the AAAI Conference on Artificial Intelligence}, publisher={Association for the Advancement of Artificial Intelligence (AAAI)}, author={Cui, Yuning and Tao, Yi and Ren, Wenqi and Knoll, Alois}, year={2023}, month=June, pages={479–487} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF