Built independently by an author, for readers. Read the story and support ChapterPal

keyword

soft attention

Soft attention is a deep learning mechanism that computes a continuous, weighted average across all parts of an input, assigning fractional importance scores to different regions, tokens, or feature representations rather than making discrete, all-or-nothing selections. Because these attention weights and the resulting feature aggregations are entirely continuous and differentiable, the underlying neural network can be trained end-to-end using standard backpropagation and gradient descent algorithms. This approach contrasts with hard attention, which stochastically selects discrete locations and typically requires reinforcement learning techniques to optimize. Widely applied in computer vision and natural language processing, soft attention enables models to smoothly emphasize salient, task-relevant features while suppressing irrelevant information and background noise.

4 items

Attention Gated Networks: Learning to Leverage Salient Regions in Medical Images

Attention Gated Networks: Learning to Leverage Salient Regions in Medical Images

Jo Schlemper, Ozan Oktay, Michiel Schaap, Mattias Heinrich, Bernhard Kainz, Ben Glocker, Daniel Rueckert

OrganizationsHeartFlowImperial College LondonUniversity of Lübeck

Why you should read this

Introduces computationally efficient attention gates that integrate into standard convolutional architectures to automatically focus on target anatomical structures, eliminating the need for dedicated localization steps while improving medical image classification and 3D segmentation performance.

We propose a novel attention gate (AG) model for medical image analysis that automatically learns to focus on target structures of varying shapes and sizes. Models trained with AGs implicitly learn to suppress irrelevant regions in an input image while highlighting salient features useful for a specific task. This enables us to eliminate the necessity of using explicit external tissue/organ localisation modules when using convolutional neural networks (CNNs). AGs can be easily integrated into standard CNN models such as VGG or U-Net architectures with minimal computational overhead while increasing the model sensitivity and prediction accuracy. The proposed AG models are evaluated on a variety of tasks, including medical image classification and segmentation. For classification, we demonstrate the use case of AGs in scan plane detection for fetal ultrasound screening. We show that the proposed attention mechanism can provide efficient object localisation while improving the overall prediction performance by reducing false positives. For segmentation, the proposed architecture is evaluated on two large 3D CT abdominal datasets with manual annotations for multiple organs. Experimental results show that AG models consistently improve the prediction performance of the base architectures across different datasets and training sizes while preserving computational efficiency. Moreover, AGs guide the model activations to be focused around salient regions, which provides better insights into how model predictions are made. The source code for the proposed AG models is publicly available.

Added

2026-09-18

Attention U-Net: Learning Where to Look for the Pancreas

Attention U-Net: Learning Where to Look for the Pancreas

Ozan Oktay, Jo Schlemper, Loic Le Folgoc, Matthew Lee, Mattias Heinrich, Kazunari Misawa, Kensaku Mori, Steven McDonagh, Nils Y Hammerla, Bernhard Kainz, Ben Glocker, Daniel Rueckert

OrganizationsAichi Cancer CenterBabylon HealthHeartFlowImperial College LondonNagoya UniversityUniversity of Lübeck

Why you should read this

Proposes attention gates for the U-Net architecture that automatically focus on target structures of varying shapes and sizes while suppressing irrelevant background regions, eliminating the need for multi-stage localization pipelines in medical image segmentation.

We propose a novel attention gate (AG) model for medical imaging that automatically learns to focus on target structures of varying shapes and sizes. Models trained with AGs implicitly learn to suppress irrelevant regions in an input image while highlighting salient features useful for a specific task. This enables us to eliminate the necessity of using explicit external tissue/organ localisation modules of cascaded convolutional neural networks (CNNs). AGs can be easily integrated into standard CNN architectures such as the U-Net model with minimal computational overhead while increasing the model sensitivity and prediction accuracy. The proposed Attention U-Net architecture is evaluated on two large CT abdominal datasets for multi-class image segmentation. Experimental results show that AGs consistently improve the prediction performance of U-Net across different datasets and training sizes while preserving computational efficiency. The code for the proposed architecture is publicly available.

Added

2026-09-11

Residual Attention Network for Image Classification

Residual Attention Network for Image Classification

Fei Wang, Mengqing Jiang, Chen Qian, Shuo Yang, Cheng Li, Honggang Zhang, Xiaogang Wang, Xiaoou Tang

OrganizationsBeijing University of Posts and TelecommunicationsSenseTimeThe Chinese University of Hong KongTsinghua University

Why you should read this

Proposes attention residual learning to scale attention-aware convolutional neural networks to hundreds of layers, delivering state-of-the-art image classification accuracy while significantly reducing computational cost compared to standard deep residual networks.

In this work, we propose "Residual Attention Network", a convolutional neural network using attention mechanism which can incorporate with state-of-art feed forward network architecture in an end-to-end training fashion. Our Residual Attention Network is built by stacking Attention Modules which generate attention-aware features. The attention-aware features from different modules change adaptively as layers going deeper. Inside each Attention Module, bottom-up top-down feedforward structure is used to unfold the feedforward and feedback attention process into a single feedforward process. Importantly, we propose attention residual learning to train very deep Residual Attention Networks which can be easily scaled up to hundreds of layers. Extensive analyses are conducted on CIFAR-10 and CIFAR-100 datasets to verify the effectiveness of every module mentioned above. Our Residual Attention Network achieves state-of-the-art object recognition performance on three benchmark datasets including CIFAR-10 (3.90% error), CIFAR-100 (20.45% error) and ImageNet (4.8% single model and single crop, top-5 error). Note that, our method achieves 0.6% top-1 accuracy improvement with 46% trunk depth and 69% forward FLOPs comparing to ResNet-200. The experiment also demonstrates that our network is robust against noisy labels.

Added

2026-09-11

Effective Approaches to Attention-based Neural Machine Translation

Effective Approaches to Attention-based Neural Machine Translation

Minh-Thang Luong, Hieu Pham, Christopher D. Manning

OrganizationsStanford University

Why you should read this

Generalizes the attention mechanism by introducing the dot-product scoring function and distinguishing between global and local attention strategies.

An attentional mechanism has lately been used to improve neural machine translation (NMT) by selectively focusing on parts of the source sentence during translation.However, there has been little work exploring useful architectures for attention-based NMT.This paper examines two simple and effective classes of attentional mechanism: a global approach which always attends to all source words and a local one that only looks at a subset of source words at a time.We demonstrate the effectiveness of both approaches on the WMT translation tasks between English and German in both directions.With local attention, we achieve a significant gain of 5.0 BLEU points over non-attentional systems that already incorporate known techniques such as dropout.Our ensemble model using different attention architectures yields a new state-of-the-art result in the WMT'15 English to German translation task with 25.9 BLEU points, an improvement of 1.0 BLEU points over the existing best system backed by NMT and an n-gram reranker. 1

Added

2026-02-11