Medical Transformer: Gated Axial-Attention for Medical Image Segmentation
Jeya Maria Jose ValanarasuPoojan OzaIlker HacihalilogluVishal M. Patel
Introduces Medical Transformer (MedT), a gated axial-attention architecture paired with a local-global training strategy designed to overcome the data-scarcity bottleneck of vision transformers and outperform convolutional baselines in medical image segmentation.
Accurate automated medical image segmentation is critical for clinical decision support, surgical planning, and timely disease diagnosis. While standard deep convolutional neural networks remain widely used for these tasks, their localized receptive fields prevent them from effectively capturing broad image context and long-range dependencies, often leading to false positives in background regions. Conversely, self-attention transformer models excel at modeling full context but typically require massive labeled datasets to learn positional relationships accurately—an impractical requirement given the high cost and scarcity of annotated clinical images.
The article demonstrates that transformer architectures can be tailored specifically for medical imaging without large-scale pre-training by introducing a gated axial-attention mechanism and a two-branch training strategy. The authors evaluate this approach, termed Medical Transformer (MedT), to determine whether it can outperform standard convolutional and transformer baselines on small-sample medical datasets.
To address computational complexity and data scarcity, the authors decompose full image self-attention into height-wise and width-wise operations, known as axial attention, and introduce learnable control gates to dynamically weight relative positional embeddings. This gating allows the model to scale positional influence based on the data available, avoiding the errors caused by poorly trained positional biases on small datasets. Additionally, they implement a Local-Global training framework featuring a shallow global branch operating on the full image alongside a deeper local branch analyzing sixteen smaller image patches. The overall framework was evaluated on three distinct benchmarks: a neonatal brain ultrasound dataset of 1,629 scans, a colon gland histology dataset with 165 images, and a multi-organ cell nucleus dataset with 44 images.
The evaluations yielded several key findings. First, MedT achieved superior accuracy across all three benchmarks, delivering an F1 score of 88.84% on brain ultrasound, 81.02% on gland histology, and 79.55% on cell nucleus segmentation. Second, on the smaller histology datasets where standard transformers underperformed relative to convolutional baselines, MedT outperformed standard axial-attention transformers by 4.76 percentage points in F1 score on gland segmentation and 2.72 percentage points on nucleus segmentation. Third, MedT matched or exceeded leading convolutional baselines (such as Res-UNet) by 1.32, 2.19, and 0.06 percentage points on the three benchmarks, respectively, while noticeably eliminating false positives in surrounding background tissue and detecting subtle target structures. Finally, architectural analysis showed MedT achieves these gains efficiently with approximately 1.4 million parameters, which is lower than or comparable to standard convolutional networks like U-Net.
These results demonstrate that transformers can be trained from scratch on scarce, domain-specific clinical datasets without relying on large external training corpuses or heavy computational infrastructure. In clinical workflows, reducing false-positive segmentations directly improves the reliability of diagnostic tools and lowers procedural risk during image-guided interventions. The ability to learn global anatomical structures while capturing fine boundary details suggests transformer-based architectures can effectively replace conventional convolutional pipelines in medical computer vision.
Organizations developing medical imaging software should consider adopting gated attention mechanisms and dual-branch patch training when working with limited labeled data. Next steps should include deploying pilot studies to validate MedT across larger multi-center clinical trials, evaluating performance on whole-slide imaging or 3D volumetric modalities, and conducting runtime latency benchmarks in real-time surgical settings.
The findings are presented with strong confidence across the tested ultrasound and histology datasets; however, certain operational boundaries apply. The evaluations were performed on resized 2D image inputs using a specific hardware configuration, and the framework's effectiveness on full-resolution 3D volumetric scans or under severe class imbalance remains to be fully characterized.
- Paper: U-Net: Convolutional Networks for Biomedical Image Segmentation, Olaf Ronneberger et al. (2015). It introduces the foundational U-Net architecture for biomedical image segmentation, which MedT adapts and benchmarks against using self-attention mechanisms.
- Paper: Attention U-Net: Learning Where to Look for the Pancreas, Ozan Oktay et al. (2018). It establishes the concept of incorporating attention gates directly into medical segmentation pipelines to control feature saliency, inspiring MedT's gated attention formulation.
- Paper: Attention Gated Networks: Learning to Leverage Salient Regions in Medical Images, Jo Schlemper et al. (2018). It formalizes grid-based contextual attention gating for medical imaging, providing prerequisite principles for designing gated self-attention mechanisms.
- Paper: Rethinking Semantic Segmentation from a Sequence-to-Sequence Perspective with Transformers, Sixiao Zheng et al. (2020). It formulates semantic segmentation as a sequence-to-sequence prediction problem using Transformers, laying the conceptual groundwork for pure attention-based dense prediction.
- Paper: Image Transformer, Niki Parmar et al. (2018). It pioneers 1D and 2D restricted axial/local self-attention for images, which MedT adopts and modifies to manage computational complexity on 2D medical scans.
- Paper: Image Segmentation Using Deep Learning: A Survey, Shervin Minaee et al. (2020). It provides a comprehensive survey of deep learning paradigms and encoder-decoder baselines in image segmentation prior to the widespread shift to Transformer backbones.
- Paper: Swin-Unet: Unet-like Pure Transformer for Medical Image Segmentation, Hu Cao et al. (2021). It extends pure Transformer medical image segmentation by building a complete U-shaped encoder-decoder network using shifted-window self-attention.
- Paper: TransFuse: Fusing Transformers and CNNs for Medical Image Segmentation, Yundong Zhang et al. (2021). It builds on medical Transformer architectures by designing parallel CNN and Transformer branches with a BiFusion mechanism to balance global context and local boundary precision.
- Paper: UNETR: Transformers for 3D Medical Image Segmentation, Ali Hatamizadeh et al. (2021). It generalizes Transformer-based medical image segmentation into volumetric 3D sequence-to-sequence modeling with multi-scale convolutional decoding.
- Paper: Swin UNETR: Swin Transformers for Semantic Segmentation of Brain Tumors in MRI Images, Ali Hatamizadeh et al. (2022). It further develops 3D medical segmentation by pairing hierarchical shifted-window Transformer encoders with multi-resolution decoding for volumetric MRI scans.
- Paper: Segment anything in medical images, Jun Ma et al. (2023). It scales Transformer-based medical image segmentation into a universal, prompt-driven foundation model across diverse clinical imaging modalities.
- Paper: U-KAN Makes Strong Backbone for Medical Image Segmentation and Generation, Chenxin Li et al. (2025). It investigates alternative non-linear tokenized backbones beyond standard Transformer attention for medical image segmentation and generation.
