Medical Transformer: Gated Axial-Attention for Medical Image Segmentation

Jeya Maria Jose ValanarasuPoojan OzaIlker HacihalilogluVishal M. Patel

article2021International Conference on Medical Image Computing and Computer-Assisted Intervention1,470 citations

Introduces Medical Transformer (MedT), a gated axial-attention architecture paired with a local-global training strategy designed to overcome the data-scarcity bottleneck of vision transformers and outperform convolutional baselines in medical image segmentation.

Listen

Accurate automated medical image segmentation is critical for clinical decision support, surgical planning, and timely disease diagnosis. While standard deep convolutional neural networks remain widely used for these tasks, their localized receptive fields prevent them from effectively capturing broad image context and long-range dependencies, often leading to false positives in background regions. Conversely, self-attention transformer models excel at modeling full context but typically require massive labeled datasets to learn positional relationships accurately—an impractical requirement given the high cost and scarcity of annotated clinical images.

The article demonstrates that transformer architectures can be tailored specifically for medical imaging without large-scale pre-training by introducing a gated axial-attention mechanism and a two-branch training strategy. The authors evaluate this approach, termed Medical Transformer (MedT), to determine whether it can outperform standard convolutional and transformer baselines on small-sample medical datasets.

To address computational complexity and data scarcity, the authors decompose full image self-attention into height-wise and width-wise operations, known as axial attention, and introduce learnable control gates to dynamically weight relative positional embeddings. This gating allows the model to scale positional influence based on the data available, avoiding the errors caused by poorly trained positional biases on small datasets. Additionally, they implement a Local-Global training framework featuring a shallow global branch operating on the full image alongside a deeper local branch analyzing sixteen smaller image patches. The overall framework was evaluated on three distinct benchmarks: a neonatal brain ultrasound dataset of 1,629 scans, a colon gland histology dataset with 165 images, and a multi-organ cell nucleus dataset with 44 images.

The evaluations yielded several key findings. First, MedT achieved superior accuracy across all three benchmarks, delivering an F1 score of 88.84% on brain ultrasound, 81.02% on gland histology, and 79.55% on cell nucleus segmentation. Second, on the smaller histology datasets where standard transformers underperformed relative to convolutional baselines, MedT outperformed standard axial-attention transformers by 4.76 percentage points in F1 score on gland segmentation and 2.72 percentage points on nucleus segmentation. Third, MedT matched or exceeded leading convolutional baselines (such as Res-UNet) by 1.32, 2.19, and 0.06 percentage points on the three benchmarks, respectively, while noticeably eliminating false positives in surrounding background tissue and detecting subtle target structures. Finally, architectural analysis showed MedT achieves these gains efficiently with approximately 1.4 million parameters, which is lower than or comparable to standard convolutional networks like U-Net.

These results demonstrate that transformers can be trained from scratch on scarce, domain-specific clinical datasets without relying on large external training corpuses or heavy computational infrastructure. In clinical workflows, reducing false-positive segmentations directly improves the reliability of diagnostic tools and lowers procedural risk during image-guided interventions. The ability to learn global anatomical structures while capturing fine boundary details suggests transformer-based architectures can effectively replace conventional convolutional pipelines in medical computer vision.

Organizations developing medical imaging software should consider adopting gated attention mechanisms and dual-branch patch training when working with limited labeled data. Next steps should include deploying pilot studies to validate MedT across larger multi-center clinical trials, evaluating performance on whole-slide imaging or 3D volumetric modalities, and conducting runtime latency benchmarks in real-time surgical settings.

The findings are presented with strong confidence across the tested ultrasound and histology datasets; however, certain operational boundaries apply. The evaluations were performed on resized 2D image inputs using a specific hardware configuration, and the framework's effectiveness on full-resolution 3D volumetric scans or under severe class imbalance remains to be fully characterized.

Cover for Medical Transformer: Gated Axial-Attention for Medical Image Segmentation

Abstract

Over the past decade, Deep Convolutional Neural Networks have been widely adopted for medical image segmentation and shown to achieve adequate performance. However, due to the inherent inductive biases present in the convolutional architectures, they lack understanding of long-range dependencies in the image. Recently proposed Transformer-based architectures that leverage self-attention mechanism encode long-range dependencies and learn representations that are highly expressive. This motivates us to explore Transformer-based solutions and study the feasibility of using Transformer-based network architectures for medical image segmentation tasks. Majority of existing Transformer-based network architectures proposed for vision applications require large-scale datasets to train properly. However, compared to the datasets for vision applications, for medical imaging the number of data samples is relatively low, making it difficult to efficiently train transformers for medical applications. To this end, we propose a Gated Axial-Attention model which extends the existing architectures by introducing an additional control mechanism in the self-attention module. Furthermore, to train the model effectively on medical images, we propose a Local-Global training strategy (LoGo) which further improves the performance. Specifically, we operate on the whole image and patches to learn global and local features, respectively. The proposed Medical Transformer (MedT) is evaluated on three different medical image segmentation datasets and it is shown that it achieves better performance than the convolutional and other related transformer-based architectures. Code: this https URL

Table of Contents

  • 1 Introduction
  • 2 Medical Transformer (MedT)
  • 2.1 Self-Attention Overview
  • Axial-Attention
  • 2.2 Gated Axial-Attention
  • 2.3 Local-Global Training
  • 3 Experiments and Results
  • 3.1 Dataset details
  • 3.2 Implementation details
  • 3.3 Results
  • 4 Conclusion
  • References
  • 5 Dataset details
  • 5.1 Brain US Dataset
  • 5.2 GLAS Dataset
  • 5.3 MoNuSeg Dataset
  • 6 MedT details
  • 7 Training details
  • 8 Analysis
  • 8.1 Ablation Study
  • 8.2 Number of Parameters
  • 9 Results
  • 10 Concurrent works
  • References

Knowls

  1. Knowl 1 — Gated Axial-Attention Mechanism

    model/method

    The gated axial-attention mechanism factorizes standard 2D self-attention into two sequential 1D self-attention operations along the height and width axes, and introduces four learnable gating scalars to regulate the influence of relative positional encodings. This design prevents inaccurate positional encodings from degrading segmentation performance when models are trained on small datasets.

    Let x∈RCin×H×Wx \in \mathbb{R}^{C_{\text{in}} \times H \times W} be an input feature map with height HH, width WW, and channel dimension CinC_{\text{in}}. With learnable linear projections WQ,WK,WV∈RCin×CoutW_Q, W_K, W_V \in \mathbb{R}^{C_{\text{in}} \times C_{\text{out}}}, queries q=WQxq = W_Q x, keys k=WKxk = W_K x, and values v=WVxv = W_V x are computed. Along the width axis, the gated self-attention output yij∈RCouty_{ij} \in \mathbb{R}^{C_{\text{out}}} at spatial location (i,j)(i, j) (where i∈{1,…,H}i \in \{1, \dots, H\} and j∈{1,…,W}j \in \{1, \dots, W\}) is defined by:

    yij=∑w=1Wsoftmax(qijTkiw+GQqijTriwq+GKkiwTriwk)(GV1viw+GV2riwv)y_{ij} = \sum_{w=1}^W \text{softmax}\left( q_{ij}^T k_{iw} + G_Q q_{ij}^T r^q_{iw} + G_K k_{iw}^T r^k_{iw} \right) \left( G_{V1} v_{iw} + G_{V2} r^v_{iw} \right)

    where:

    • qij,kij,vij∈RCoutq_{ij}, k_{ij}, v_{ij} \in \mathbb{R}^{C_{\text{out}}} are query, key, and value vectors at coordinate (i,j)(i, j);
    • rq,rk,rv∈RW×Wr^q, r^k, r^v \in \mathbb{R}^{W \times W} are learnable relative positional encoding matrices along the width dimension for queries, keys, and values, respectively;
    • GQ,GK,GV1,GV2∈RG_Q, G_K, G_{V1}, G_{V2} \in \mathbb{R} are learnable scalar gates that dynamically weight the contributions of the positional bias.

    An equivalent 1D gated formulation is applied along the height axis HH. If the dataset is small and positional encodings cannot be reliably estimated, the gating parameters converge toward zero, suppressing noisy positional terms; if accurate positional encodings are learned, the gates scale them to higher values.

  2. Knowl 2 — Local-Global (LoGo) Multi-Branch Training Strategy

    model/method

    The Local-Global (LoGo) training strategy is a two-branch architecture paradigm designed to balance broad contextual understanding and fine-grained boundary localization in medical image segmentation:

    1. Global Branch: Operates on the full image at its original resolution (I×II \times I) using a shallow transformer network consisting of 2 encoder blocks and 2 decoder blocks. Because the initial layers of the transformer architecture suffice to capture global context and long-range pixel dependencies across the entire canvas, fewer layers are required.

    2. Local Branch: Subdivides the original image into 16 non-overlapping patches of spatial size (I/4)×(I/4)(I/4) \times (I/4). Each patch is processed through a deeper network consisting of 5 encoder blocks and 5 decoder blocks to extract fine local anatomical structures.

    3. Reassembly and Feature Fusion: The output feature maps of all 16 patches from the local branch are rearranged back to their original spatial positions to form a complete feature map of dimension I×II \times I. The global branch feature map and the reassembled local branch feature map are combined via element-wise addition, followed by a 1×11 \times 1 convolution layer to predict the final segmentation mask.

  3. Knowl 3 — Medical Transformer (MedT) Architecture

    model/method

    Medical Transformer (MedT) is an encoder-decoder architecture for medical image segmentation that combines gated axial-attention layers with the Local-Global (LoGo) dual-branch training scheme.

    • Initial Feature Extractor: Both the global and local branches take as input feature maps produced by a shared initial convolutional block consisting of three convolutional layers, each followed by Batch Normalization and ReLU activation.
    • Encoder Blocks: Constructed using gated axial transformer layers. Within each encoder block, features pass through a 1×11 \times 1 convolution layer, normalization, and two consecutive multi-head attention layers: one along the height axis and the other along the width axis. Each multi-head attention block contains 8 gated axial attention heads. The outputs of both axial attention blocks are concatenated, projected through another 1×11 \times 1 convolution, and added to the residual shortcut path.
    • Positional Encoding Usage: The global branch uses gated axial-attention with learnable relative positional encodings and gating scalars (GQ,GK,GV1,GV2G_Q, G_K, G_{V1}, G_{V2}). The local branch uses axial-attention without positional encodings.
    • Decoder Blocks and Skip Connections: Each decoder block consists of a convolutional layer followed by an upsampling operation and ReLU activation. Standard skip connections transfer feature maps directly from each encoder block to its corresponding decoder block in both branches.
    • Branch Depth Configuration: The global branch comprises 2 encoder blocks and 2 decoder blocks, whereas the local branch comprises 5 encoder blocks and 5 decoder blocks.
  4. Knowl 4 — MedT Optimization Objective and Gate Training Schedule

    experimental setup

    Medical Transformer (MedT) is trained by minimizing the binary cross-entropy (CE) loss between pixel predictions p^(x,y)∈[0,1]\hat{p}(x, y) \in [0, 1] and binary ground-truth labels p(x,y)∈{0,1}p(x, y) \in \{0, 1\} across all spatial coordinates of an image with dimensions w×hw \times h:

    LCE(p,p^)=−1wh∑x=0w−1∑y=0h−1[p(x,y)log⁡(p^(x,y))+(1−p(x,y))log⁡(1−p^(x,y))]\mathcal{L}_{\text{CE}}(p, \hat{p}) = -\frac{1}{wh} \sum_{x=0}^{w-1} \sum_{y=0}^{h-1} \left[ p(x, y) \log\left(\hat{p}(x, y)\right) + (1 - p(x, y)) \log\left(1 - \hat{p}(x, y)\right) \right]

    Training hyperparameters and scheduling constraints include:

    • Optimizer: Adam optimizer with an initial learning rate of 0.0010.001.
    • Batch Size: 4.
    • Epochs: 400 epochs total.
    • Gate Freezing / Warm-Up: During the first 10 training epochs, parameter updates to the learnable gates (GQ,GK,GV1,GV2G_Q, G_K, G_{V1}, G_{V2}) in the gated axial attention layers are disabled, allowing the base projection and feature extraction weights to stabilize before optimizing the positional gating parameters.
    • Hardware: Nvidia Quadro 8000 GPU.
  5. Knowl 5 — Medical Image Segmentation Evaluation Datasets

    experimental setup

    MedT and baseline methods are evaluated across three distinct biomedical segmentation benchmarks:

    1. Brain US Dataset: 2D cranial ultrasound scans obtained from 20 premature neonates (age <1< 1 year) for segmenting brain ventricles and septum pellucidum. Contains 1,629 annotated images, split into 1,300 images for training and 329 images for testing. Images are resized to 128×128128 \times 128 pixels.
    2. GLAnd Segmentation (GlaS) Dataset: Microscopic images of Hematoxylin and Eosin (H&E) stained colon histology slides for gland segmentation. Contains 165 images, partitioned into 85 training images and 80 testing images. Images are resized to 128×128128 \times 128 pixels.
    3. MoNuSeg Dataset: Microscopic H&E stained tissue images captured at 40×40\times magnification across multiple organs and patients for generalized cell nuclear boundary segmentation. Contains 30 training images (comprising ∼22,000\sim 22,000 nuclear boundaries) and 14 test images (comprising >7,000> 7,000 nuclear boundaries). Images are resized to 512×512512 \times 512 pixels.
  6. Knowl 6 — Segmentation Performance Comparison Across Datasets

    data/table

    The performance of MedT was compared against convolutional baselines (FCN, U-Net, U-Net++, Res-UNet), fully attention baselines (Axial Attention U-Net), and isolated components of the proposed method (Gated Axial Attention alone, LoGo alone) across three datasets using F1 score and Intersection-over-Union (IoU) metrics.

    On GlaS and MoNuSeg, standard Axial Attention U-Net performs worse than convolutional baselines due to sample scarcity making relative positional encodings hard to optimize. Gated axial attention and LoGo training alleviate this limitation, enabling MedT to achieve the highest F1 and IoU scores across all datasets.

    Type Network Brain US GlaS MoNuSeg
    F1 IoU F1 IoU F1 IoU
    Convolutional FCN 82.79 75.02 66.61 50.84 28.84 28.71
    Baselines U-Net 85.37 79.31 77.78 65.34 79.43 65.99
    U-Net++ 86.59 79.95 78.03 65.55 79.49 66.04
    Res-UNet 87.50 79.61 78.83 65.95 79.49 66.07
    Fully Attention Baseline Axial Attention U-Net 87.92 80.14 76.26 63.03 76.83 62.49
    Proposed Gated Axial Attn. 88.39 80.70 79.91 67.85 76.44 62.01
    LoGo 88.54 80.84 79.68 67.69 79.56 66.17
    MedT 88.84 81.34 81.02 69.61 79.55 66.17
  7. Knowl 7 — Ablation Study of Architecture Components on Brain US Dataset

    data/table

    An ablation study evaluated on the Brain US dataset quantifies the individual contribution of residual connections, axial attention layers, gating mechanisms, and dual-branch (global vs. local) configurations.

    Training solely on local patches (Local only) causes a severe performance drop to 77.55% F1 score due to the absence of inter-patch contextual modeling. Training on full images with a shallow global branch (Global only) achieves 87.67% F1 score. Combining both branches via LoGo yields 88.54% F1 score, and adding gated axial attention (MedT) yields the highest performance at 88.84% F1 score.

    Network Configuration F1 Score (%)
    U-Net 85.37
    Res-UNet 87.50
    Axial UNet 87.92
    Gated Axial UNet 88.39
    Global only 87.67
    Local only 77.55
    LoGo (Axial UNet without gating) 88.54
    MedT 88.84
  8. Knowl 8 — Model Parameter Efficiency and Capacity Comparison

    data/table

    The parameter count of MedT is compared with baseline networks on the Brain US dataset. To demonstrate that performance gains are architectural rather than driven by parameter count, modified baselines (mod) with reduced filter counts were evaluated.

    MedT contains 1.4 M parameters, outperforming models with significantly larger parameter footprints (e.g., FCN with 12.5 M, Res-UNet with 5.32 M, and U-Net with 3.13 M). Gated axial-attention introduces only 4 additional learnable scalar parameters per attention layer relative to standard axial-attention.

    Network Parameters F1 Score (%)
    FCN 12.5 M 82.79
    U-Net 3.13 M 87.71
    U-Net (mod) 1.3 M 85.37
    Res-UNet 5.32 M 87.73
    Res-UNet (mod) 1.34 M 87.50
    Axial UNet 1.3 M 87.92
    Gated Axial UNet 1.3 M 88.39
    MedT 1.4 M 88.84

Coverage note — Qualitative visual prediction comparisons illustrating specific false-positive reductions and small-structure delineations were omitted as standalone knowls because their core findings are quantitatively represented in the benchmark tables.

References

  1. 1.Badrinarayanan, V., Kendall, A., Cipolla, R.: Segnet: A deep convolutional encoder-decoder architecture for image segmentation. IEEE transactions on pattern analysis and machine intelligence 39(12), 2481–2495 (2017)
  2. 2.Chen, J., Lu, Y., Yu, Q., Luo, X., Adeli, E., Wang, Y., Lu, L., Yuille, A.L., Zhou, Y.: Transunet: Transformers make strong encoders for medical image segmentation. arXiv preprint arXiv:2102.04306 (2021)
  3. 3.Chen, L.C., Papandreou, G., Kokkinos, I., Murphy, K., Yuille, A.L.: Semantic image segmentation with deep convolutional nets and fully connected crfs. arXiv preprint arXiv:1412.7062 (2014)
  4. 4.Ç i̧cek, O., Abdulkadir, A., Lienkamp, S.S., Brox, T., Ronneberger, O.: 3d u-net: ¨ learning dense volumetric segmentation from sparse annotation. In: International conference on medical image computing and computer-assisted intervention. pp. 424–432. Springer (2016)
  5. 5.Devlin, J., Chang, M.W., Lee, K., Toutanova, K.: Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805 (2018)
  6. 6.Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al.: An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929 (2020)
  7. 7.Ho, J., Kalchbrenner, N., Weissenborn, D., Salimans, T.: Axial attention in multidimensional transformers. arXiv preprint arXiv:1912.12180 (2019)
  8. 8.Huang, H., Lin, L., Tong, R., Hu, H., Zhang, Q., Iwamoto, Y., Han, X., Chen, Y.W., Wu, J.: Unet 3+: A full-scale connected unet for medical image segmentation. In: ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). pp. 1055–1059. IEEE (2020)
  9. 9.Huang, Z., Wang, X., Huang, L., Huang, C., Wei, Y., Liu, W.: Ccnet: Criss-cross attention for semantic segmentation. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. pp. 603–612 (2019)
  10. 10.Kingma, D.P., Ba, J.: Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980 (2014)
  11. 11.Kumar, N., Verma, R., Anand, D., Zhou, Y., Onder, O.F., Tsougenis, E., Chen, H., Heng, P.A., Li, J., Hu, Z., et al.: A multi-organ nucleus segmentation challenge. IEEE transactions on medical imaging 39(5), 1380–1391 (2019)
  12. 12.Kumar, N., Verma, R., Sharma, S., Bhargava, S., Vahadane, A., Sethi, A.: A dataset and a technique for generalized nuclear segmentation for computational pathology. IEEE transactions on medical imaging 36(7), 1550–1560 (2017)
  13. 13.Li, X., Chen, H., Qi, X., Dou, Q., Fu, C.W., Heng, P.A.: H-denseunet: hybrid densely connected unet for liver and tumor segmentation from ct volumes. IEEE transactions on medical imaging 37(12), 2663–2674 (2018)
  14. 14.Mehta, S., Mercan, E., Bartlett, J., Weaver, D., Elmore, J.G., Shapiro, L.: Y-net: joint segmentation and classification for diagnosis of breast biopsy images. In: International Conference on Medical Image Computing and Computer-Assisted Intervention. pp. 893–901. Springer (2018)
  15. 15.Milletari, F., Navab, N., Ahmadi, S.A.: V-net: Fully convolutional neural networks for volumetric medical image segmentation. In: 2016 fourth international conference on 3D vision (3DV). pp. 565–571. IEEE (2016)
  16. 16.Oktay, O., Schlemper, J., Folgoc, L.L., Lee, M., Heinrich, M., Misawa, K., Mori, K., McDonagh, S., Hammerla, N.Y., Kainz, B., et al.: Attention u-net: Learning where to look for the pancreas. arXiv preprint arXiv:1804.03999 (2018)
  17. 17.Ronneberger, O., Fischer, P., Brox, T.: U-net: Convolutional networks for biomedical image segmentation. In: International Conference on Medical image computing and computer-assisted intervention. pp. 234–241. Springer (2015)
  18. 18.Shaw, P., Uszkoreit, J., Vaswani, A.: Self-attention with relative position representations. In: Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers). pp. 464–468 (2018)
  19. 19.Sirinukunwattana, K., Pluim, J.P., Chen, H., Qi, X., Heng, P.A., Guo, Y.B., Wang, L.Y., Matuszewski, B.J., Bruni, E., Sanchez, U., et al.: Gland segmentation in colon histology images: The glas challenge contest. Medical image analysis 35, 489–502 (2017)
  20. 20.Touvron, H., Cord, M., Douze, M., Massa, F., Sablayrolles, A., J´egou, H.: Training data-efficient image transformers & distillation through attention. arXiv preprint arXiv:2012.12877 (2020)
  21. 21.Valanarasu, J.M.J., Sindagi, V.A., Hacihaliloglu, I., Patel, V.M.: Kiu-net: Over-complete convolutional architectures for biomedical image and volumetric segmentation. arXiv preprint arXiv:2010.01663 (2020)
  22. 22.Valanarasu, J.M.J., Sindagi, V.A., Hacihaliloglu, I., Patel, V.M.: Kiu-net: Towards accurate segmentation of biomedical images using over-complete representations. In: International Conference on Medical Image Computing and Computer-Assisted Intervention. pp. 363–373. Springer (2020)
  23. 23.Valanarasu, J.M.J., Yasarla, R., Wang, P., Hacihaliloglu, I., Patel, V.M.: Learning to segment brain anatomy from 2d ultrasound with less data. IEEE Journal of Selected Topics in Signal Processing 14(6), 1221–1234 (2020)
  24. 24.Wang, H., Zhu, Y., Green, B., Adam, H., Yuille, A., Chen, L.C.: Axial-deeplab: Stand-alone axial-attention for panoptic segmentation. arXiv preprint arXiv:2003.07853 (2020)
  25. 25.Wang, P., Cuccolo, N.G., Tyagi, R., Hacihaliloglu, I., Patel, V.M.: Automatic real-time cnn-based neonatal brain ventricles segmentation. In: 2018 IEEE 15th International Symposium on Biomedical Imaging (ISBI 2018). pp. 716–719. IEEE (2018)
  26. 26.Wang, X., Han, S., Chen, Y., Gao, D., Vasconcelos, N.: Volumetric attention for 3d medical image segmentation and detection. In: International Conference on Medical Image Computing and Computer-Assisted Intervention. pp. 175–184. Springer (2019)
  27. 27.Xiao, X., Lian, S., Luo, Z., Li, S.: Weighted res-unet for high-quality retina vessel segmentation. In: 2018 9th international conference on information technology in medicine and education (ITME). pp. 327–331. IEEE (2018)
  28. 28.Zhang, Y., Liu, H., Hu, Q.: Transfuse: Fusing transformers and cnns for medical image segmentation. arXiv preprint arXiv:2102.08005 (2021)
  29. 29.Zhao, H., Shi, J., Qi, X., Wang, X., Jia, J.: Pyramid scene parsing network. In: Proceedings of the IEEE conference on computer vision and pattern recognition. pp. 2881–2890 (2017)
  30. 30.Zheng, S., Lu, J., Zhao, H., Zhu, X., Luo, Z., Wang, Y., Fu, Y., Feng, J., Xiang, T., Torr, P.H., et al.: Rethinking semantic segmentation from a sequence-to-sequence perspective with transformers. arXiv preprint arXiv:2012.15840 (2020)
  31. 31.Zhou, Z., Siddiquee, M.M.R., Tajbakhsh, N., Liang, J.: Unet++: A nested u-net architecture for medical image segmentation. In: Deep learning in medical image analysis and multimodal learning for clinical decision support, pp. 3–11. Springer (2018)

Citation

MLA
Valanarasu, J. M. J., et al. “Medical Transformer: Gated Axial-Attention for Medical Image Segmentation”. arXiv, 2021, http://arxiv.org/abs/2102.10662v2.
APA
Valanarasu, J. M. J., Oza, P., Hacihaliloglu, I., & Patel, V. M. (2021). Medical Transformer: Gated Axial-Attention for Medical Image Segmentation. arXiv. http://arxiv.org/abs/2102.10662v2
Chicago
Valanarasu, J. M. J., P. Oza, I. Hacihaliloglu, and V. M. Patel. 2021. “Medical Transformer: Gated Axial-Attention for Medical Image Segmentation”. arXiv. http://arxiv.org/abs/2102.10662v2.
Harvard
Valanarasu, J.M.J. et al. (2021) “Medical Transformer: Gated Axial-Attention for Medical Image Segmentation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2102.10662v2.
Vancouver
1. Valanarasu JMJ, Oza P, Hacihaliloglu I, Patel VM (2021) Medical Transformer: Gated Axial-Attention for Medical Image Segmentation. arXiv

BibTeX

@article{valanarasu2021medical,
  title = {Medical Transformer: Gated Axial-Attention for Medical Image Segmentation},
  author = {Valanarasu, Jeya Maria Jose and Oza, Poojan and Hacihaliloglu, Ilker and Patel, Vishal M.},
  year = {2021},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2102.10662v2},
  eprint = {2102.10662}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF