Attention Gated Networks: Learning to Leverage Salient Regions in Medical Images

Jo SchlemperOzan OktayMichiel SchaapMattias HeinrichBernhard KainzBen GlockerDaniel Rueckert

article2018Medical Image Anal.1,902 citations

Introduces computationally efficient attention gates that integrate into standard convolutional architectures to automatically focus on target anatomical structures, eliminating the need for dedicated localization steps while improving medical image classification and 3D segmentation performance.

Listen

Manual analysis and annotation of complex medical images are time-consuming and subject to human error. While deep learning networks have advanced automated medical diagnostics, standard models struggle to accurately isolate small organs or subtle anatomical views characterized by significant shape and size variations. Current systems routinely rely on multi-stage or cascaded frameworks that employ separate neural networks first to locate a region of interest and then to perform classification or segmentation. However, these multi-network setups lead to redundant computations, inflated parameter counts, and excessive training complexity.

The article demonstrates that incorporating soft-attention gates directly into standard single-stage convolutional networks improves sensitivity, precision, and efficiency across medical image classification and segmentation tasks. By learning to highlight salient target regions and suppress irrelevant background noise on the fly, this mechanism removes the need for separate, external organ localization models or manual bounding-box annotations.

The researchers developed a modular additive attention gate mechanism using grid-based contextual gating. They evaluated this framework across two challenging tasks: two-dimensional fetal ultrasound scan-plane classification using 2,694 patient examinations encompassing over 190,000 frames, and three-dimensional multi-organ abdominal computed tomography (CT) segmentation using two separate benchmarks of 150 and 82 patient scans. Performance was measured against standard single-stage networks and multi-stage cascaded baselines in terms of classification accuracy, precision, recall, Dice similarity coefficients (a measure of overlap accuracy), surface-to-surface error distances, parameter efficiency, and runtime.

Incorporating attention gates consistently improved performance while maintaining high computational efficiency. In 3D CT pancreas segmentation—a difficult organ due to low contrast and high anatomical variability—the Attention U-Net increased overlap accuracy from 0.814 to 0.840 and reduced boundary error distances from 2.36 mm to 1.92 mm with only an 8% increase in model parameters and negligible inference overhead (0.179 seconds versus 0.167 seconds per volume). In ultrasound plane detection, the attention-gated network improved overall classification precision from 0.878 to 0.916, showing up to a 5% precision gain on subtle structures like kidneys, profiles, and spine views by eliminating false positives. Crucially, single-stage attention models achieved segmentation performance competitive with complex, multi-model cascaded systems without requiring region cropping or multi-network training pipelines.

These findings indicate that soft-attention gates can replace cumbersome multi-stage computer vision pipelines with unified, end-to-end trainable models. In clinical deployment, this translates to faster processing, lower hardware and infrastructure costs, reduced engineering complexity, and fewer diagnostic false alarms. Furthermore, because attention gates generate visual spatial activation maps directly, they offer built-in model interpretability without additional computational overhead, helping clinicians understand and verify automated decisions.

Engineering and clinical teams developing medical imaging tools should adopt attention gating into single-network architectures rather than building complex, multi-stage cascading pipelines. Future development should explore deploying these 3D attention networks on higher-resolution, non-downsampled image batches as GPU hardware expands, while continuing to investigate training strategies that stabilize gradient flow across multi-scale attention layers.

A key operational limitation noted in the article is that CT volumes had to be downsampled to isotropic 2.00 mm resolution due to GPU memory constraints, whereas cascaded 2D models operate at original slice resolutions. Additionally, optimizing soft-attention parameters requires careful training strategies, such as deep supervision and two-stage fine-tuning, to prevent gradient saturation. Confidence in the reported results is high, as the attention mechanism demonstrated consistent, statistically significant gains across multiple clinical imaging modalities, diverse organ classes, and varying training dataset sizes.

Cover for Attention Gated Networks: Learning to Leverage Salient Regions in Medical Images

Abstract

We propose a novel attention gate (AG) model for medical image analysis that automatically learns to focus on target structures of varying shapes and sizes. Models trained with AGs implicitly learn to suppress irrelevant regions in an input image while highlighting salient features useful for a specific task. This enables us to eliminate the necessity of using explicit external tissue/organ localisation modules when using convolutional neural networks (CNNs). AGs can be easily integrated into standard CNN models such as VGG or U-Net architectures with minimal computational overhead while increasing the model sensitivity and prediction accuracy. The proposed AG models are evaluated on a variety of tasks, including medical image classification and segmentation. For classification, we demonstrate the use case of AGs in scan plane detection for fetal ultrasound screening. We show that the proposed attention mechanism can provide efficient object localisation while improving the overall prediction performance by reducing false positives. For segmentation, the proposed architecture is evaluated on two large 3D CT abdominal datasets with manual annotations for multiple organs. Experimental results show that AG models consistently improve the prediction performance of the base architectures across different datasets and training sizes while preserving computational efficiency. Moreover, AGs guide the model activations to be focused around salient regions, which provides better insights into how model predictions are made. The source code for the proposed AG models is publicly available.

Table of Contents

  • 1 Introduction
  • 1.1 Related Work
  • 1.2 Contributions
  • 2 Methodology
  • 2.1 Convolutional Neural Network
  • 2.2 Attention Gate Module
  • 2.2.1 Multi-dimensional Attention
  • 2.2.2 Gating Signal and Grid Attention
  • 2.2.3 Backward Pass through Attention Gates
  • 2.3 Attention Gates for Segmentation
  • 2.4 Attention Gates for Classification
  • 3 Experiments and Results
  • 3.1 Evaluation Datasets
  • 3.1.1 3D-CT Abdominal Image Datasets
  • 3.1.2 2D Fetal Ultrasound Image Dataset
  • 3.2 Model Training and Implementation Details
  • 3.2.1 Implementation Details:
  • 3.3 3D-CT Abdominal Image Segmentation Results
  • 3.3.1 Comparison to State-of-the-Art CT Abdominal Segmentation Frameworks
  • 3.4 2D Fetal Ultrasound Image Classification Results
  • 3.5 Attention Map Analysis
  • 3.5.1 Object Localisation using Attention Maps
  • 4 Discussion
  • 5 Conclusion
  • References

Knowls

  1. Knowl 1 — Additive Attention Gate with Grid-Based Gating

    model/method

    The Additive Attention Gate (AG) prunes feature activations in irrelevant background regions and highlights salient features without requiring explicit region proposal or bounding-box supervision. Let xl={xil}i=1nx^l = \{x_i^l\}_{i=1}^n represent the feature activation map at layer l∈{1,…,L}l \in \{1, \dots, L\}, where each xil∈RFlx_i^l \in \mathbb{R}^{F_l} is a spatial feature vector of channel depth FlF_l, and nn is the number of spatial locations. A contextual gating signal g∈RFgg \in \mathbb{R}^{F_g}, collected from a coarser spatial resolution, is used to disambiguate task-irrelevant information in xlx^l.

    Rather than collapsing gg into a global vector via flattening, grid attention retains the coarse-scale spatial feature map prior to flattening (e.g., Fg×Hx/2r×Wx/2rF_g \times H_x / 2^r \times W_x / 2^r after rr pooling operations). The coarse grid is resampled (e.g., via bilinear or trilinear interpolation) to match the spatial resolution of xlx^l, allowing spatial gating coefficients to be resolved on a local regional basis.

    The attention coefficients αl={αil}i=1n\alpha^l = \{\alpha_i^l\}_{i=1}^n with αil∈[0,1]\alpha_i^l \in [0, 1] are computed via additive compatibility:

    qatt,il=ψTσ1(WxTxil+WgTgi+bxg)+bψq_{att, i}^l = \psi^T \sigma_1\left(W_x^T x_i^l + W_g^T g_i + b_{xg}\right) + b_\psi

    αl=σ2(qattl(xl,g;Θatt))\alpha^l = \sigma_2\left(q_{att}^l(x^l, g; \Theta_{att})\right)

    where σ1(x)=max⁡(0,x)\sigma_1(x) = \max(0, x) is a Rectified Linear Unit (ReLU), σ2(x)\sigma_2(x) is a normalization function, Wx∈RFl×FintW_x \in \mathbb{R}^{F_l \times F_{int}} and Wg∈RFg×FintW_g \in \mathbb{R}^{F_g \times F_{int}} are linear transformations implemented as channel-wise 1×1×11 \times 1 \times 1 convolutions, ψ∈RFint×1\psi \in \mathbb{R}^{F_{int} \times 1} is a linear mapping, and bxg∈RFint,bψ∈Rb_{xg} \in \mathbb{R}^{F_{int}}, b_\psi \in \mathbb{R} are bias terms. The intermediate channel dimension FintF_{int} is set to match or reduce the parameter footprint.

    The gated output x^l={x^il}i=1n\hat{x}^l = \{\hat{x}_i^l\}_{i=1}^n scales each input activation vector element-wise by its spatial attention coefficient:

    x^il=αilxil\hat{x}_i^l = \alpha_i^l x_i^l

    The normalization σ2\sigma_2 is selected based on the task: the standard sigmoid function σ2(x)=1/(1+exp⁡(−x))\sigma_2(x) = 1 / (1 + \exp(-x)) is used for dense segmentation predictions to avoid over-sparse activations, whereas classification tasks use a shifted sum-normalization across spatial coordinates.

  2. Knowl 2 — Gradient Backpropagation Filtering in Attention Gates

    theoretical result

    In a feed-forward neural network incorporating Attention Gates, let xil=f(xil−1;Φl−1)x_i^l = f(x_i^{l-1}; \Phi^{l-1}) denote the feature map produced by layer l−1l-1 with parameters Φl−1\Phi^{l-1}, and let x^il=αilxil\hat{x}_i^l = \alpha_i^l x_i^l denote the attention-gated feature vector at spatial position ii. The gradient of the gated feature with respect to the convolution parameters Φl−1\Phi^{l-1} in layer l−1l-1 is:

    ∂x^il∂Φl−1=αil∂f(xil−1;Φl−1)∂Φl−1+∂αil∂Φl−1xil\frac{\partial \hat{x}_i^l}{\partial \Phi^{l-1}} = \alpha_i^l \frac{\partial f(x_i^{l-1}; \Phi^{l-1})}{\partial \Phi^{l-1}} + \frac{\partial \alpha_i^l}{\partial \Phi^{l-1}} x_i^l

    Because the first gradient term on the right-hand side is directly scaled by the attention coefficient αil∈[0,1]\alpha_i^l \in [0, 1], backward gradient updates originating from background and task-irrelevant regions (where αil≈0\alpha_i^l \approx 0) are down-weighted. As a consequence, parameter updates in shallower layers are driven predominantly by task-relevant foreground image regions.

  3. Knowl 3 — Multi-Dimensional Attention Gate Mechanism

    model/method

    When an input image contains multiple semantic target classes, Attention Gates can be extended to compute multi-dimensional attention coefficients. Instead of a single scalar coefficient per spatial location, the attention gate generates an mm-dimensional vector of coefficients αil=[αi,(1)l,αi,(2)l,…,αi,(m)l]T∈[0,1]m\alpha_i^l = \left[\alpha_{i, (1)}^l, \alpha_{i, (2)}^l, \dots, \alpha_{i, (m)}^l\right]^T \in [0, 1]^m for each spatial location ii.

    The gated output feature map x^l\hat{x}^l concatenates the feature responses filtered by each sub-gate:

    x^l=[α(1)l⊙xl,α(2)l⊙xl,…,α(m)l⊙xl]\hat{x}^l = \left[ \alpha_{(1)}^l \odot x^l, \alpha_{(2)}^l \odot x^l, \dots, \alpha_{(m)}^l \odot x^l \right]

    where ⊙\odot denotes element-wise multiplication. Each sub-attention gate α(k)l\alpha_{(k)}^l learns to focus on a distinct semantic structure or organ class, allowing complementary spatial representations to be extracted and merged along skip connections.

  4. Knowl 4 — Attention U-Net Architecture for 3D Medical Image Segmentation

    model/method

    The Attention U-Net integrates Attention Gates (AGs) into a 3D U-Net encoder-decoder architecture to segment anatomical structures without requiring cascaded localization models.

    In the encoding path, 3D input volumes are progressively filtered and downsampled by factors of 2 across successive scales. In the decoding path, feature maps are upsampled and merged with the corresponding encoder feature maps. AGs are placed directly on each skip connection prior to concatenation: the finer-scale encoder feature map xlx^l serves as the input to be gated, while the coarser-scale decoder feature map gg from the level below acts as the gating query signal.

    For dense segmentation, the gate normalization function is the sigmoid non-linearity σ2(x)=1/(1+exp⁡(−x))\sigma_2(x) = 1 / (1 + \exp(-x)), which avoids the extreme spatial sparsity caused by softmax normalizations. Deep supervision is incorporated at intermediate decoder resolutions, applying auxiliary segmentation loss heads at each scale. Training utilizes the Sørensen-Dice loss across all foreground classes to mitigate class imbalance between small target organs and the image background.

  5. Knowl 5 — Attention-Gated Sononet (AG-Sononet) for Image Classification

    model/method

    Attention-Gated Sononet (AG-Sononet) adapts a VGG-style classification network (Sononet) by incorporating spatial attention gates to focus on discriminative anatomical landmarks.

    The global gating signal gg is extracted from the coarsest convolutional activation map (the final layer of the feature extraction trunk). Attention gates are placed at intermediate depths (layers 11 and 14) immediately before spatial pooling operations. Earliest layers are not gated because low-level features lack semantic class discriminability.

    To prevent over-sparse attention coefficients that destabilize classification training, the attention maps at scale ll are normalized across spatial indices jj using a shifted sum-normalization:

    αil=αil−αmin⁡l∑j(αjl−αmin⁡l),where αmin⁡l=min⁡jαjl\alpha_i^l = \frac{\alpha_i^l - \alpha_{\min}^l}{\sum_{j} (\alpha_j^l - \alpha_{\min}^l)}, \quad \text{where } \alpha_{\min}^l = \min_j \alpha_j^l

    At each gated layer ll, a scale-specific feature vector x~l∈RFl\tilde{x}^l \in \mathbb{R}^{F_l} is obtained by computing the spatially weighted average x~l=∑i=1nαilxil\tilde{x}^l = \sum_{i=1}^n \alpha_i^l x_i^l. In addition, global average pooling is applied to the coarsest feature map x~coarse\tilde{x}^{coarse}. The scale-wise vectors are concatenated into a single representation [x~l1,x~l2,x~coarse][\tilde{x}^{l_1}, \tilde{x}^{l_2}, \tilde{x}^{coarse}] and classified via fully connected layers. Training proceeds by applying deep supervision losses to each scale individually before fine-tuning the final concatenated classification layers.

  6. Knowl 6 — Weakly Supervised Object Localisation via Attention Maps

    algorithm

    Attention maps produced during the standard forward pass of an attention-gated classification network (such as AG-Sononet) provide spatial localization of class-discriminative anatomical structures without requiring guided backpropagation or bounding-box labels at training time.

    Input: Multi-scale attention maps {αl\alpha^l}l=1L_{l=1}^L, binarization threshold τ\tau, Gaussian smoothing kernel κ\kappa
    Output: Bounding box B=(xmin⁡,ymin⁡,xmax⁡,ymax⁡)B = (x_{\min}, y_{\min}, x_{\max}, y_{\max})
    for each scale l∈{1,…,L}l \in \{1, \dots, L\} do
        α~l←Convolve(αl,κ)\tilde{\alpha}^l \leftarrow \text{Convolve}(\alpha^l, \kappa)
        Ml←BinaryMask(α~l≥τ)M^l \leftarrow \text{BinaryMask}(\tilde{\alpha}^l \ge \tau)
        {Ckl}←ConnectedComponents(Ml)\{C_k^l\} \leftarrow \text{ConnectedComponents}(M^l)
    end for
    C∗←SelectSpatiallyOverlappingComponent({{Ckl}}l=1L)C^* \leftarrow \text{SelectSpatiallyOverlappingComponent}(\{\{C_k^l\}\}_{l=1}^L)
    B←BoundingBoxAroundComponent(C∗)B \leftarrow \text{BoundingBoxAroundComponent}(C^*)
    return BB

    Because bounding boxes are generated directly from feed-forward attention maps via smoothing, thresholding, and connected-component selection across scales, the procedure incurs negligible computational cost during real-time inference.

  7. Knowl 7 — CT Abdominal Multi-Organ Segmentation Performance on CT-150

    data/table

    The 3D Attention U-Net (Att U-Net) was benchmarked against the standard 3D U-Net on the CT-150 dataset (150 contrast-enhanced 3D abdominal CT scans with manual segmentations of pancreas, spleen, and kidneys; downsampled to isotropic 2.00 mm resolution). Models were evaluated on 120/30 and 30/120 train/test splits, as well as against uniformly expanded U-Net baselines with equivalent or higher parameter capacities.

    Method U-Net Att U-Net U-Net Att U-Net
    Train/Test split 120/30 120/30 30/120 30/120
    Pancreas DSC 0.814±0.1160.814 \pm 0.116 0.840±0.087\mathbf{0.840 \pm 0.087} 0.741±0.1370.741 \pm 0.137 0.767±0.132\mathbf{0.767 \pm 0.132}
    Pancreas precision 0.848±0.1100.848 \pm 0.110 0.849±0.0980.849 \pm 0.098 0.789±0.1760.789 \pm 0.176 0.794±0.1500.794 \pm 0.150
    Pancreas recall 0.806±0.1260.806 \pm 0.126 0.841±0.092\mathbf{0.841 \pm 0.092} 0.743±0.1790.743 \pm 0.179 0.762±0.1450.762 \pm 0.145
    Pancreas S2S Dist (mm) 2.358±1.4642.358 \pm 1.464 1.920±1.284\mathbf{1.920 \pm 1.284} 3.765±3.4523.765 \pm 3.452 3.507±3.8143.507 \pm 3.814
    Spleen DSC 0.962±0.0130.962 \pm 0.013 0.965±0.0130.965 \pm 0.013 0.935±0.0950.935 \pm 0.095 0.943±0.0920.943 \pm 0.092
    Kidney DSC 0.963±0.0130.963 \pm 0.013 0.964±0.0160.964 \pm 0.016 0.951±0.0190.951 \pm 0.019 0.954±0.0210.954 \pm 0.021
    Number of params 5.88 M 6.40 M 5.88 M 6.40 M
    Inference time 0.167 s 0.179 s 0.167 s 0.179 s
    Method # Parameters DSC Precision Recall S2S Dist (mm) Run time
    U-Net 6.44 M 0.821±0.1190.821 \pm 0.119 0.849±0.1110.849 \pm 0.111 0.814±0.1250.814 \pm 0.125 2.383±1.9182.383 \pm 1.918 0.191 s
    U-Net 10.40 M 0.825±0.1040.825 \pm 0.104 0.861±0.0820.861 \pm 0.082 0.807±0.1210.807 \pm 0.121 2.202±1.1442.202 \pm 1.144 0.222 s
    Att U-Net 6.40 M 0.840±0.0870.840 \pm 0.087 0.849±0.0980.849 \pm 0.098 0.841±0.0920.841 \pm 0.092 1.920±1.2841.920 \pm 1.284 0.179 s

    Att U-Net improves pancreas Dice score by 2.6% (p=0.005p = 0.005) and reduces surface-to-surface distance from 2.358 mm to 1.920 mm on the 120/30 split, driven by a statistically significant increase in recall from 0.806 to 0.841. On the reduced training set (30/120 split), Att U-Net maintains a 2.6% Dice advantage (p=0.01p = 0.01). Allocating model capacity to attention gates yields higher accuracy than uniformly increasing standard U-Net capacity up to 10.40 M parameters (p=0.007p = 0.007).

  8. Knowl 8 — CT Pancreas Segmentation Benchmarks on TCIA CT-82

    data/table

    Pancreas segmentation performance was evaluated on the public NIH-TCIA Pancreas-CT benchmark (CT-82, comprising 82 contrast-enhanced 3D abdominal CT scans). Standard 3D U-Net and Attention U-Net were tested in three experimental regimes: zero-shot transfer before fine-tuning (BFT, trained on CT-150 and tested on CT-82), after fine-tuning (AFT, 61 train / 21 test), and trained from scratch (SCR, 61 train / 21 test).

    Regime Method Dice Score Precision Recall S2S Dist (mm)
    BFT U-Net 0.690±0.1320.690 \pm 0.132 0.680±0.1090.680 \pm 0.109 0.733±0.1900.733 \pm 0.190 6.389±3.9006.389 \pm 3.900
    BFT Attention U-Net 0.712±0.110\mathbf{0.712 \pm 0.110} 0.693±0.115\mathbf{0.693 \pm 0.115} 0.751±0.149\mathbf{0.751 \pm 0.149} 5.251±2.551\mathbf{5.251 \pm 2.551}
    AFT U-Net 0.820±0.0430.820 \pm 0.043 0.824±0.0700.824 \pm 0.070 0.828±0.0640.828 \pm 0.064 2.464±0.5292.464 \pm 0.529
    AFT Attention U-Net 0.831±0.038\mathbf{0.831 \pm 0.038} 0.825±0.0730.825 \pm 0.073 0.840±0.053\mathbf{0.840 \pm 0.053} 2.305±0.568\mathbf{2.305 \pm 0.568}
    SCR U-Net 0.815±0.0680.815 \pm 0.068 0.815±0.1050.815 \pm 0.105 0.826±0.0620.826 \pm 0.062 2.576±1.1802.576 \pm 1.180
    SCR Attention U-Net 0.821±0.057\mathbf{0.821 \pm 0.057} 0.815±0.0930.815 \pm 0.093 0.835±0.057\mathbf{0.835 \pm 0.057} 2.333±0.856\mathbf{2.333 \pm 0.856}
    Published Method Dataset Pancreas DSC Train/Test Validation
    Dense-Dilated FCN CT-82 66.0±10.066.0 \pm 10.0 63/9 5-fold CV
    2D U-Net CT-82 75.7±9.075.7 \pm 9.0 66/16 5-fold CV
    HN 2D FCN Stage-1 CT-82 76.8±11.176.8 \pm 11.1 62/20 4-fold CV
    HN 2D FCN Stage-2 (Cascaded) CT-82 81.2±7.381.2 \pm 7.3 62/20 4-fold CV
    2D FCN CT-82 80.3±9.080.3 \pm 9.0 62/20 4-fold CV
    2D FCN + RNN CT-82 82.3±6.782.3 \pm 6.7 62/20 4-fold CV
    Single Model 2D FCN CT-82 75.7±10.575.7 \pm 10.5 62/20 4-fold CV
    Multi-Model 2D FCN (Cascaded) CT-82 82.2±5.782.2 \pm 5.7 62/20 4-fold CV
    3D Attention U-Net CT-82 81.48±6.2381.48 \pm 6.23 65/17 5-fold CV

    Under 5-fold cross-validation on CT-82, single-model 3D Attention U-Net achieves 81.48±6.2381.48 \pm 6.23 DSC without requiring multi-stage region cropping or cascaded networks.

  9. Knowl 9 — Fetal Ultrasound Standard Scan Plane Classification Performance

    data/table

    AG-Sononet was benchmarked against the standard Sononet architecture on 2694 2D fetal ultrasound examinations (split by subject into 122,233 training, 30,553 validation, and 38,243 test frames) across 13 anatomical standard scan planes and background frames. Models were tested across three base channel capacities (N∈{8,16,32}N \in \{8, 16, 32\} initial filters).

    Method Accuracy F1 Precision Recall Fwd (ms) Bwd (ms) Parameters
    Sononet-8 0.969 0.899 0.878 0.922 1.36 2.60 0.16 M
    AG-Sononet-8 0.977 0.922 0.916 0.929 1.92 3.47 0.18 M
    Sononet-16 0.977 0.923 0.916 0.931 1.45 3.92 0.65 M
    AG-Sononet-16 0.978 0.929 0.924 0.934 1.94 5.13 0.70 M
    Sononet-32 0.979 0.931 0.924 0.938 2.40 6.72 2.58 M
    AG-Sononet-32 0.980 0.933 0.931 0.935 2.92 8.68 2.79 M
    Class AG-Sononet-8 Precision AG-Sononet-8 Recall AG-Sononet-8 F1
    Brain (Cb.) 0.988 (−0.002-0.002) 0.982 (−0.002-0.002) 0.985 (−0.002-0.002)
    Brain (Tv.) 0.980 (+0.003+0.003) 0.990 (+0.002+0.002) 0.985 (+0.003+0.003)
    Profile 0.953 (+0.055+0.055) 0.962 (+0.009+0.009) 0.958 (+0.033+0.033)
    Lips 0.976 (+0.029+0.029) 0.956 (−0.003-0.003) 0.966 (+0.013+0.013)
    Abdominal 0.963 (+0.011+0.011) 0.961 (+0.007+0.007) 0.962 (+0.009+0.009)
    Kidneys 0.863 (+0.054+0.054) 0.902 (+0.003+0.003) 0.882 (+0.030+0.030)
    Femur 0.975 (+0.019+0.019) 0.976 (−0.005-0.005) 0.975 (+0.007+0.007)
    Spine (Cor.) 0.935 (+0.049+0.049) 0.979 (+0.000+0.000) 0.957 (+0.026+0.026)
    Spine (Sag.) 0.936 (+0.055+0.055) 0.979 (−0.012-0.012) 0.957 (+0.024+0.024)
    4CH 0.943 (+0.035+0.035) 0.970 (+0.007+0.007) 0.956 (+0.022+0.022)
    3VV 0.694 (+0.050+0.050) 0.722 (−0.014-0.014) 0.708 (+0.021+0.021)
    RVOT 0.691 (+0.029+0.029) 0.705 (+0.044+0.044) 0.698 (+0.036+0.036)
    LVOT 0.925 (+0.022+0.022) 0.933 (+0.027+0.027) 0.929 (+0.024+0.024)
    Background 0.995 (−0.001-0.001) 0.992 (+0.007+0.007) 0.993 (+0.003+0.003)

    AG-Sononet matches the performance of the un-gated Sononet baseline with double the channel capacity (e.g., AG-Sononet-8 achieves 0.922 F1 and 0.916 precision with 0.18 M parameters, matching Sononet-16 at 0.65 M parameters). Improvements are most pronounced in precision (+5.5% for Profile and Spine Sagittal, +5.4% for Kidneys, +5.0% for 3-vessel view), demonstrating background noise suppression.

  10. Knowl 10 — Weakly Supervised Object Localisation Accuracy on Fetal Ultrasound

    data/table

    Bounding boxes extracted from AG-Sononet-16 feed-forward attention maps were evaluated against manual ground-truth bounding box annotations on the test set of 2D fetal ultrasound scan planes.

    Evaluation metrics include Mean Intersection over Union (IOU Mean ±\pm Std), Correctness (defined as IOU>0.5\text{IOU} > 0.5), and Relative Correctness (defined as IOU>0.5×max⁡(IOUclass)\text{IOU} > 0.5 \times \max(\text{IOU}_{\text{class}}) to account for the intrinsic bias where discriminant attention highlights sub-features rather than whole anatomical envelopes).

    Scan Plane Class IOU Mean (Std) Correctness (%) Relative Correctness (%)
    Brain (Cb.) 0.69±0.110.69 \pm 0.11 96% 96%
    Brain (Tv.) 0.68±0.120.68 \pm 0.12 96% 96%
    Profile 0.31±0.080.31 \pm 0.08 0% 80%
    Lips 0.42±0.180.42 \pm 0.18 36% 60%
    Abdominal 0.71±0.100.71 \pm 0.10 96% 96%
    Kidneys 0.73±0.130.73 \pm 0.13 92% 98%
    Femur 0.31±0.110.31 \pm 0.11 2% 58%
    Spine (Cor.) 0.53±0.130.53 \pm 0.13 56% 76%
    Spine (Sag.) 0.53±0.110.53 \pm 0.11 54% 94%
    4CH 0.61±0.140.61 \pm 0.14 76% 86%
    3VV 0.42±0.140.42 \pm 0.14 34% 62%
    RVOT 0.56±0.150.56 \pm 0.15 70% 76%
    LVOT 0.54±0.150.54 \pm 0.15 62% 80%

    High relative correctness scores (≥76%\ge 76\% across 10 of the 13 anatomical classes) confirm that the attention maps consistently localize salient anatomical landmarks in close spatial proximity to the ground-truth target structures without requiring gradient-based saliency mapping.

Coverage note — No substantial contributed material was omitted.

References

  1. 1.Anderson, P., He, X., Buehler, C., Teney, D., Johnson, M., Gould, S., Zhang, L., 2017. Bottom-up and top-down attention for image captioning and vqa. arXiv:1707.07998.
  2. 2.Bahdanau, D., Cho, K., Bengio, Y., 2014. Neural machine translation by jointly learning to align and translate. arXiv:1409.0473.
  3. 3.Bai, W., Sinclair, M., Tarroni, G., Oktay, O., Rajchl, M., Vaillant, G., Lee, A. M., Aung, N., Lukaschuk, E., Sanghvi, M. M., et al., 2017. Human-level cmr image analysis with deep fully convolutional networks. arXiv:1710.09289.
  4. 4.Baumgartner, C. F., Kamnitsas, K., Matthew, J., Fletcher, T. P., Smith, S., Koch, L. M., Kainz, B., Rueckert, D., 2016. Real-time detection and localisation of fetal standard scan planes in 2d freehand ultrasound. arXiv:1612.05601.
  5. 5.Britz, D., Goldie, A., Luong, M.-T., Le, Q., 2017. Massive exploration of neural machine translation architectures. arXiv:1703.03906.
  6. 6.Cai, J., Lu, L., Xie, Y., Xing, F., Yang, L., 2017. Improving deep pancreas segmentation in CT and MRI images via recurrent neural contextual learning and direct loss function. MICCAI.
  7. 7.Cerrolaza, J.J., Summers, R.M., Linguraru, M.G., 2016. Soft multi-organ shape models via generalized PCA: a general framework. In: MICCAI. Springer, pp. 219–228.
  8. 8.Chen, H., Dou, Q., Ni, D., Cheng, J.-Z., Qin, J., Li, S., Heng, P.-A., 2015. Automatic fetal ultrasound standard plane detection using knowledge transferred recurrent neural networks. In: International Conference on Medical Image Computing and Computer-Assisted Intervention. Springer, pp. 507–514.
  9. 9.Drozdzal, M., Vorontsov, E., Chartrand, G., Kadoury, S., Pal, C., 2016. The Importance of Skip Connections in Biomedical Image Segmentation. In: Deep Learning and Data Labeling for Medical Applications. Springer, pp. 179–187.
  10. 10.Esteva, A., Kuprel, B., Novoa, R.A., Ko, J., Swetter, S.M., Blau, H.M., Thrun, S., 2017. Dermatologist-level classification of skin cancer with deep neural networks. Nature 542 (7639), 115.
  11. 11.Gibson, E., Giganti, F., Hu, Y., Bonmati, E., Bandula, S., Gurusamy, K., Davidson, B.R., Pereira, S.P., Clarkson, M.J., Barratt, D.C., 2017. Towards image-guided pancreas and biliary endoscopy: Automatic multi-organ segmentation on abdominal ct with dense dilated networks. In: International Conference on Medical Image Computing and Computer-Assisted Intervention. Springer, pp. 728–736.
  12. 12.Greff, K., Srivastava, R. K., Schmidhuber, J., 2016. Highway and residual networks learn unrolled iterative estimation. arXiv:1612.07771.
  13. 13.Guan, Q., Huang, Y., Zhong, Z., Zheng, Z., Zheng, L., Yang, Y., 2018. Diagnose like a radiologist: attention guided convolutional neural network for thorax disease classification. arXiv:1801.09927.
  14. 14.Heinrich, M. P., Blendowski, M., Oktay, O., 2018. Ternarynet: faster deep model inference without GPUs for medical 3D segmentation using sparse and binary convolutions. arXiv:1801.09449.
  15. 15.Heinrich, M.P., Oktay, O., 2017. BRIEFnet: Deep pancreas segmentation using binary sparse convolutions. In: MICCAI. Springer, pp. 329–337.
  16. 16.Hu, J., Shen, L., Sun, G., 2017. Squeeze-and-excitation networks. arXiv:1709.01507.
  17. 17.Jetley, S., Lord, N.A., Lee, N., Torr, P., 2018. Learn to pay attention. In: International Conference on Learning Representations.
  18. 18.Kamnitsas, K., Bai, W., Ferrante, E., McDonagh, S., Sinclair, M., Pawlowski, N., Rajchl, M., Lee, M., Kainz, B., Rueckert, D., Glocker, B., 2018. Ensembles of multiple models and architectures for robust brain tumour segmentation. In: Brainlesion: Glioma, Multiple Sclerosis, Stroke and Traumatic Brain Injuries, pp. 450–462. Cham
  19. 19.Kamnitsas, K., Ledig, C., Newcombe, V.F., Simpson, J.P., Kane, A.D., Menon, D.K., Rueckert, D., Glocker, B., 2017. Efficient multi-scale 3D CNN with fully connected CRF for accurate brain lesion segmentation. Med. Image Anal. 36, 61–78.
  20. 20.Kawahara, J., Hamarneh, G., 2016. Multi-resolution-tract cnn with hybrid pretrained and skin-lesion trained layers. In: International Workshop on Machine Learning in Medical Imaging. Springer, pp. 164–171.
  21. 21.Khened, M., Kollerathu, V. A., Krishnamurthi, G., 2018. Fully convolutional multi-scale residual densenets for cardiac segmentation and automated cardiac diagnosis using ensemble of classifiers. arXiv:1801.05173.
  22. 22.Lee, C.-Y., Xie, S., Gallagher, P., Zhang, Z., Tu, Z., 2015. Deeply-supervised nets. In: Artificial Intelligence and Statistics, pp. 562–570.
  23. 23.Liao, F., Liang, M., Li, Z., Hu, X., Song, S., 2017. Evaluate the malignancy of pulmonary nodules using the 3D deep leaky noisy-or network. arXiv:1711.08324.
  24. 24.Litjens, G.J.S., Kooi, T., Bejnordi, B.E., Setio, A.A.A., Ciompi, F., Ghafoorian, M., van der Laak, J.A.W.M., van Ginneken, B., Sánchez, C.I., 2017. A survey on deep learning in medical image analysis. arXiv:1702.05747.
  25. 25.Liu, J., Wang, G., Hu, P., Duan, L.-Y., Kot, A.C., 2017. Global context-aware attention lstm networks for 3d action recognition. CVPR.
  26. 26.Long, J., Shelhamer, E., Darrell, T., 2015. Fully convolutional networks for semantic segmentation. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 3431–3440.
  27. 27.Lu, J., Xiong, C., Parikh, D., Socher, R., 2016. Knowing when to look: adaptive attention via a visual sentinel for image captioning. arXiv:1612.01887.
  28. 28.Luong, M.-T., Pham, H., Manning, C. D., 2015. Effective approaches to attention-based neural machine translation. arXiv:1508.04025.
  29. 29.Madani, A., Arnaout, R., Mofrad, M., Arnaout, R., 2018. Fast and accurate view classification of echocardiograms using deep learning. npj Digital Medicine 1 (1), 6.
  30. 30.Milletari, F., Navab, N., Ahmadi, S.-A., 2016. V-net: Fully convolutional neural networks for volumetric medical image segmentation. In: 3D Vision (3DV), 2016 Fourth International Conference on. IEEE, pp. 565–571.
  31. 31.Mnih, V., Heess, N., Graves, A., et al., 2014. Recurrent models of visual attention. In: Advances in neural information processing systems, pp. 2204–2212.
  32. 32.Nam, H., Ha, J., Kim, J., 2016. Dual attention networks for multimodal reasoning and matching. arXiv:1611.00471.
  33. 33.NHS Screening Programmes, 2015. Fetal anomaly screen programme handbook. NHS.
  34. 34.Oda, M., Shimizu, N., Roth, H.R., Karasawa, K., Kitasaka, T., Misawa, K., Fujiwara, M., Rueckert, D., Mori, K., 2017. 3D FCN Feature Driven Regression Forest-based Pancreas Localization and Segmentation. In: Deep Learning in Medical Image Analysis and Multimodal Learning for Clinical Decision Support. Springer, pp. 222–230.
  35. 35.Paszke, A., Gross, S., Chintala, S., Chanan, G., Yang, E., DeVito, Z., Lin, Z., Desmaison, A., Antiga, L., Lerer, A., 2017. Automatic differentiation in pytorch.
  36. 36.Payer, C., Štern, D., Bischof, H., Urschler, M., 2017. Multi-label whole heart segmentation using CNNs and anatomical label configurations. In: STACOM. Springer, pp. 190–198.
  37. 37.Pei, W., Baltrusaitis, T., Tax, D.M.J., Morency, L., 2016. Temporal attention-gated model for robust sequence classification. arXiv:1612.00385.
  38. 38.Pesce, E., Ypsilantis, P.-P., Withey, S., Bakewell, R., Goh, V., Montana, G., 2017. Learning to detect chest radiographs containing lung nodules using visual attention networks. arXiv:1712.00996.
  39. 39.Ren, M., Zemel, R. S., 2016. End-to-end instance segmentation and counting with recurrent attention. arXiv:1605.09410.
  40. 40.Ronneberger, O., Fischer, P., Brox, T., 2015. U-net: Convolutional networks for biomedical image segmentation. In: International Conference on Medical image computing and computer-assisted intervention. Springer, pp. 234–241.
  41. 41.Roth, H., Farag, A., Turkbey, E. B., Lu, L., Liu, J., Summers, R. M., 2016. Data from Pancreas-CT. The Cancer Imaging Archive.
  42. 42.Roth, H.R., Lu, L., Lay, N., Harrison, A.P., Farag, A., Sohn, A., Summers, R.M., 2018. Spatial aggregation of holistically-nested convolutional neural networks for automated pancreas localization and segmentation. Med Image Anal 45, 94–107.
  43. 43.Roth, H. R., Oda, H., Hayashi, Y., Oda, M., Shimizu, N., Fujiwara, M., Misawa, K., Mori, K., 2017. Hierarchical 3D fully convolutional networks for multi-organ segmentation. arXiv:1704.06382.
  44. 44.Saito, A., Nawano, S., Shimizu, A., 2016. Joint optimization of segmentation and shape prior from level-set-based statistical shape model, and its application to the automated segmentation of abdominal organs. Med. Image Anal. 28, 46–65.
  45. 45.Sarraf, S., DeSouza, D.D., Anderson, J., Tofighi, G., 2017. Deepad: Alzheimer’s disease classification via deep convolutional neural networks using mri and fmri. bioRxiv doi:10.1101/070441.
  46. 46.Shen, T., Zhou, T., Long, G., Jiang, J., Pan, S., Zhang, C., 2017. Disan: directional self-attention network for rnn/cnn-free language understanding. arXiv:1709.04696.
  47. 47.Simonyan, K., Zisserman, A., 2014. Very deep convolutional networks for large-scale image recognition. arXiv:1409.1556.
  48. 48.Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, Ł., Polosukhin, I., 2017. Attention is all you need. In: Advances in Neural Information Processing Systems, pp. 6000–6010.
  49. 49.Veličković, P., Cucurull, G., Casanova, A., Romero, A., Liò, P., Bengio, Y., 2017. Graph attention networks. arXiv:1710.10903.
  50. 50.Wang, F., Jiang, M., Qian, C., Yang, S., Li, C., Zhang, H., Wang, X., Tang, X., 2017a. Residual attention network for image classification. arXiv:1704.06904.
  51. 51.Wang, X., Girshick, R., Gupta, A., He, K., 2017b. Non-local neural networks. arXiv:1711.07971.
  52. 52.Wang, X., Peng, Y., Lu, L., Lu, Z., Summers, R.M., 2018. Tienet: text-image embedding network for common thorax disease classification and reporting in chest x-rays. arXiv:1801.04334.
  53. 53.Wolz, R., Chu, C., Misawa, K., Fujiwara, M., Mori, K., Rueckert, D., 2013. Automated abdominal multi-organ segmentation with subject-specific atlas generation. IEEE Trans. Med. Imag. 32 (9), 1723–1730.
  54. 54.Xie, S., Tu, Z., 2015. Holistically-nested edge detection. In: Proceedings of the IEEE international conference on computer vision, pp. 1395–1403.
  55. 55.Xu, K., Ba, J., Kiros, R., Cho, K., Courville, A., Salakhudinov, R., Zemel, R., Bengio, Y., 2015. Show, attend and tell: Neural image caption generation with visual attention. In: International Conference on Machine Learning, pp. 2048–2057.
  56. 56.Yang, Z., He, X., Gao, J., Deng, L., Smola, A.J., 2015. Stacked attention networks for image question answering. arXiv:1511.02274.
  57. 57.Yaqub, M., Kelly, B., Papageorghiou, A.T., Noble, J.A., 2015. Guided random forests for identification of key fetal anatomy and image categorization in ultrasound scans. In: International Conference on Medical Image Computing and Computer-Assisted Intervention. Springer, pp. 687–694.
  58. 58.Ypsilantis, P.-P., Montana, G., 2017. Learning what to look in chest x-rays with a recurrent visual attention model. arXiv:1701.06452.
  59. 59.Yu, Q., Xie, L., Wang, Y., Zhou, Y., Fishman, E. K., Yuille, A. L., 2017. Recurrent saliency transformation network: incorporating multi-stage visual cues for small organ segmentation. arXiv:1709.04518.
  60. 60.Zaharchuk, G., Gong, E., Wintermark, M., Rubin, D., Langlotz, C., 2018. Deep learning in neuroradiology. American Journal of Neuroradiology.
  61. 61.Zhang, Z., Chen, P., Sapkota, M., Yang, L., 2017. Tandemnet: Distilling knowledge from medical images using diagnostic reports as optional semantic references. In: International Conference on Medical Image Computing and Computer-Assisted Intervention. Springer, pp. 320–328.
  62. 62.Zhang, Z., Xie, Y., Xing, F., McGough, M., Yang, L., 2017. Mdnet: a semantically and visually interpretable medical image diagnosis network. arXiv:1707.02485.
  63. 63.Zhao, B., Feng, J., Wu, X., Yan, S., 2017. A survey on deep learning-based fine-grained object classification and semantic segmentation. Int. J. Autom. Comput. 14 (2), 119–135.
  64. 64.Zhou, Y., Xie, L., Shen, W., Wang, Y., Fishman, E.K., Yuille, A.L., 2017. A fixed–point model for pancreas segmentation in abdominal ct scans. In: International Conference on Medical Image Computing and Computer-Assisted Intervention. Springer, pp. 693–701.
  65. 65.Zhu, W., Liu, C., Fan, W., Xie, X., 2018. Deeplung: deep 3d dual path nets for automated pulmonary nodule detection and classification. arXiv:1801.09555.
  66. 66.Zografos, V., Valentinitsch, A., Rempfler, M., Tombari, F., Menze, B., 2015. Hierarchical multi-organ segmentation without registration in 3D abdominal CT images. In: International MICCAI Workshop on Medical Computer Vision. Springer, pp. 37–46.

Citation

MLA
Schlemper, J., et al. “Attention Gated Networks: Learning to Leverage Salient Regions in Medical Images”. arXiv, 2018, http://arxiv.org/abs/1808.08114v2.
APA
Schlemper, J., Oktay, O., Schaap, M., Heinrich, M., Kainz, B., Glocker, B., & Rueckert, D. (2018). Attention Gated Networks: Learning to Leverage Salient Regions in Medical Images. arXiv. http://arxiv.org/abs/1808.08114v2
Chicago
Schlemper, J., O. Oktay, M. Schaap, et al. 2018. “Attention Gated Networks: Learning to Leverage Salient Regions in Medical Images”. arXiv. http://arxiv.org/abs/1808.08114v2.
Harvard
Schlemper, J. et al. (2018) “Attention Gated Networks: Learning to Leverage Salient Regions in Medical Images”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1808.08114v2.
Vancouver
1. Schlemper J, Oktay O, Schaap M, Heinrich M, Kainz B, Glocker B, Rueckert D (2018) Attention Gated Networks: Learning to Leverage Salient Regions in Medical Images. arXiv

BibTeX

@article{schlemper2018attention,
  title = {Attention Gated Networks: Learning to Leverage Salient Regions in Medical Images},
  author = {Schlemper, Jo and Oktay, Ozan and Schaap, Michiel and Heinrich, Mattias and Kainz, Bernhard and Glocker, Ben and Rueckert, Daniel},
  year = {2018},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1808.08114v2},
  eprint = {1808.08114}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF