keyword
local self-attention
Local self-attention is a neural network attention mechanism where each element in an input sequence or feature map computes attention weights and aggregates information only from a localized neighborhood or bounded window around itself, rather than across the entire input. While standard global self-attention incurs quadratic computational and memory complexity relative to the total input size, restricting the receptive field to local regions reduces the complexity to scale linearly with the number of tokens. This design retains the dynamic, content-dependent weighting capability of self-attention while significantly reducing computational overhead, making it especially effective for high-resolution vision tasks and long-sequence processing where capturing fine-grained, localized context is essential.
2 items

Focal Modulation Networks
Jianwei Yang, Chunyuan Li, Xiyang Dai, Jianfeng Gao
Why you should read this
Proposes FocalNets, an attention-free architecture that replaces self-attention with multi-scale contextual modulation to outperform standard vision transformers across image classification, object detection, and semantic segmentation at comparable computational costs.
We propose focal modulation networks (FocalNets in short), where self-attention (SA) is completely replaced by a focal modulation mechanism for modeling token interactions in vision. Focal modulation comprises three components: (i) hierarchical contextualization, implemented using a stack of depth-wise convolutional layers, to encode visual contexts from short to long ranges, (ii) gated aggregation to selectively gather contexts for each query token based on its content, and (iii) element-wise modulation or affine transformation to inject the aggregated context into the query. Extensive experiments show FocalNets outperform the state-of-the-art SA counterparts (e.g., Swin and Focal Transformers) with similar computational costs on the tasks of image classification, object detection, and segmentation. Specifically, FocalNets with tiny and base size achieve 82.3% and 83.9% top-1 accuracy on ImageNet-1K. After pretrained on ImageNet-22K in 224 resolution, it attains 86.5% and 87.3% top-1 accuracy when finetuned with resolution 224 and 384, respectively. When transferred to downstream tasks, FocalNets exhibit clear superiority. For object detection with Mask R-CNN, FocalNet base trained with 1\times outperforms the Swin counterpart by 2.1 points and already surpasses Swin trained with 3\times schedule (49.0 v.s. 48.5). For semantic segmentation with UPerNet, FocalNet base at single-scale outperforms Swin by 2.4, and beats Swin at multi-scale (50.5 v.s. 49.7). Using large FocalNet and Mask2former, we achieve 58.5 mIoU for ADE20K semantic segmentation, and 57.9 PQ for COCO Panoptic Segmentation. Using huge FocalNet and DINO, we achieved 64.3 and 64.4 mAP on COCO minival and test-dev, respectively, establishing new SoTA on top of much larger attention-based models like Swinv2-G and BEIT-3. Code and checkpoints are available at this https URL.
Added
2026-09-26

Dual-Domain Attention for Image Deblurring
Yuning Cui, Yi Tao, Wenqi Ren, Alois Knoll
Why you should read this
Proposes a dual-domain attention network that pairs dynamic group convolution for localized spatial self-attention with a lightweight frequency-decoupling module, achieving state-of-the-art image deblurring quality with substantially faster inference speeds.
As a long-standing and challenging task, image deblurring aims to reconstruct the latent sharp image from its degraded counterpart. In this study, to bridge the gaps between degraded/sharp image pairs in the spatial and frequency domains simultaneously, we develop the dual-domain attention mechanism for image deblurring. Self-attention is widely used in vision tasks, however, due to the quadratic complexity, it is not applicable to image deblurring with high-resolution images. To alleviate this issue, we propose a novel spatial attention module by implementing self-attention in the style of dynamic group convolution for integrating information from the local region, enhancing the representation learning capability and reducing computational burden. Regarding frequency domain learning, many frequency-based deblurring approaches either treat the spectrum as a whole or decompose frequency components in a complicated manner. In this work, we devise a frequency attention module to compactly decouple the spectrum into distinct frequency parts and accentuate the informative part with extremely lightweight learnable parameters. Finally, we incorporate attention modules into a U-shaped network. Extensive comparisons with prior arts on the common benchmarks show that our model, named Dual-Domain Attention Network (DDANet), obtains comparable results with a significantly improved inference speed.
Added
2026-09-26
