Built independently by an author, for readers. Read the story and support ChapterPal

keyword

Spatial Attention Module

A spatial attention module is a neural network component in computer vision designed to identify and emphasize the most informative spatial regions within intermediate feature maps while suppressing irrelevant background areas. By computing a two-dimensional attention map across the height and width dimensions of visual representations, it dynamically weights spatial locations according to their task relevance. This mechanism allows deep learning models to selectively focus on key visual structures, such as target objects or salient localized features, often by performing pooling or convolutional operations across feature channels and applying the resulting spatial weights to the input features for adaptive refinement.

4 items

Self-Emphasizing Network for Continuous Sign Language Recognition

Self-Emphasizing Network for Continuous Sign Language Recognition

Lianyu Hu, Liqing Gao, Zekang Liu, Wei Feng

OrganizationsTianjin University

Why you should read this

Proposes a lightweight continuous sign language recognition framework that adaptively highlights informative hand, face, and temporal features without relying on computationally expensive pose estimation or extra supervision.

Hand and face play an important role in expressing sign language. Their features are usually especially leveraged to improve system performance. However, to effectively extract visual representations and capture trajectories for hands and face, previous methods always come at high computations with increased training complexity. They usually employ extra heavy pose-estimation networks to locate human body keypoints or rely on additional pre-extracted heatmaps for supervision. To relieve this problem, we propose a self-emphasizing network (SEN) to emphasize informative spatial regions in a self-motivated way, with few extra computations and without additional expensive supervision. Specifically, SEN first employs a lightweight subnetwork to incorporate local spatial-temporal features to identify informative regions, and then dynamically augment original features via attention maps. It's also observed that not all frames contribute equally to recognition. We present a temporal self-emphasizing module to adaptively emphasize those discriminative frames and suppress redundant ones. A comprehensive comparison with previous methods equipped with hand and face features demonstrates the superiority of our method, even though they always require huge computations and rely on expensive extra supervision. Remarkably, with few extra computations, SEN achieves new state-of-the-art accuracy on four large-scale datasets, PHOENIX14, PHOENIX14-T, CSL-Daily, and CSL. Visualizations verify the effects of SEN on emphasizing informative spatial and temporal features. Code is available at https://github.com/hulianyuyy/SEN_CSLR.

Added

2026-09-26

Dual-Domain Attention for Image Deblurring

Dual-Domain Attention for Image Deblurring

Yuning Cui, Yi Tao, Wenqi Ren, Alois Knoll

OrganizationsMassachusetts Institute of TechnologySun Yat-sen UniversityTechnical University of Munich

Why you should read this

Proposes a dual-domain attention network that pairs dynamic group convolution for localized spatial self-attention with a lightweight frequency-decoupling module, achieving state-of-the-art image deblurring quality with substantially faster inference speeds.

As a long-standing and challenging task, image deblurring aims to reconstruct the latent sharp image from its degraded counterpart. In this study, to bridge the gaps between degraded/sharp image pairs in the spatial and frequency domains simultaneously, we develop the dual-domain attention mechanism for image deblurring. Self-attention is widely used in vision tasks, however, due to the quadratic complexity, it is not applicable to image deblurring with high-resolution images. To alleviate this issue, we propose a novel spatial attention module by implementing self-attention in the style of dynamic group convolution for integrating information from the local region, enhancing the representation learning capability and reducing computational burden. Regarding frequency domain learning, many frequency-based deblurring approaches either treat the spectrum as a whole or decompose frequency components in a complicated manner. In this work, we devise a frequency attention module to compactly decouple the spectrum into distinct frequency parts and accentuate the informative part with extremely lightweight learnable parameters. Finally, we incorporate attention modules into a U-shaped network. Extensive comparisons with prior arts on the common benchmarks show that our model, named Dual-Domain Attention Network (DDANet), obtains comparable results with a significantly improved inference speed.

Added

2026-09-26