Unleashing the Potential of SAM for Medical Adaptation via Hierarchical Decoding
Zhiheng ChengQingyue WeiHongru ZhuYan WangLiangqiong QuWei ShaoYuyin Zhou
Proposes H-SAM, a prompt-free adaptation framework that equips the Segment Anything Model with a two-stage hierarchical mask decoder to achieve superior few-shot medical image segmentation performance using only minimal labeled data.
Precise medical image segmentation is critical for clinical diagnosis, treatment planning, and biomedical research, yet conventional deep learning models require vast amounts of costly, expert-annotated data. While large-scale vision foundation models like the Segment Anything Model offer powerful segmentation capabilities, their direct zero-shot performance drops significantly on specialized medical images. Adapting these models has typically required either expensive full fine-tuning or manual, expert-provided visual prompts such as bounding boxes or points during testing, which introduces clinical friction, latency, and potential human error.
The article demonstrates an efficient, prompt-free adaptation framework called H-SAM that tailors the Segment Anything Model for medical imaging using limited training data. By freezing the original image encoder and applying parameter-efficient tuning alongside a two-stage hierarchical decoding process, the model integrates learned medical priors to generate precise multi-class segmentations without needing manual prompts or massive annotated datasets.
To evaluate this framework, the authors conducted experiments across three benchmark medical imaging datasets: multi-organ abdominal CT scans from the Synapse dataset, cardiac MRI scans from the Left Atrial dataset, and prostate MRI scans from the PROMISE12 dataset. The architecture pairs parameter-efficient low-rank adaptation layers in the encoder with a two-stage mask decoder. The first decoding stage produces an initial coarse probabilistic mask, which then guides a second, refined decoding stage featuring class-balanced self-attention, learnable mask cross-attention, and a multi-scale pixel decoder with skip connections to capture fine anatomical details.
The experimental findings show substantial improvements across both few-shot and fully supervised settings. On multi-organ CT segmentation using only 10% of training slices, the proposed model achieved an 80.35% mean Dice score (an overlap accuracy metric where higher is better), outperforming competing prompt-free adaptation methods by nearly 5 percentage points and significantly lowering boundary errors. Under full supervision on the same dataset, it achieved an 86.49% mean Dice score, outperforming dedicated state-of-the-art medical segmentation architectures. Furthermore, in few-shot cardiac and prostate MRI segmentation using only 4 and 3 labeled training cases respectively, the model reached 89.22% and 87.27% accuracy, notably outperforming leading semi-supervised frameworks despite those methods relying on dozens of additional unlabeled scans.
These results indicate that foundation models can be effectively deployed in specialized clinical environments without the operational bottleneck of real-time expert prompting or the heavy computational overhead of full-model retraining. The ability to surpass complex semi-supervised approaches using only a minimal set of labeled scans highlights substantial potential to reduce data curation timelines, lower computational expenses, and mitigate label scarcity risks in healthcare artificial intelligence development.
Organizations developing clinical imaging pipelines should consider adopting two-stage, prior-guided fine-tuning strategies over prompt-dependent or data-heavy semi-supervised baselines. Future work should focus on validating the framework through clinical pilots across diverse hospital sites, evaluating its robustness on rare pathologies, and exploring extensions to native 3D volumetric architectures.
While the reported performance is strong, confidence should be framed within the context of the evaluation scope, which relied on standard public benchmark datasets and 2D slice processing rather than fully integrated prospective clinical workflows. External validation across broader multi-scanner cohorts is recommended prior to production deployment.
- Paper: Segment Anything, Alexander M. Kirillov et al. (2023). Introduces the Segment Anything Model (SAM) architecture and promptable segmentation paradigm that the source model freezes and adapts.
- Paper: Segment anything in medical images, Jun Ma et al. (2023). Establishes prompt-based adaptation of SAM for medical imaging (MedSAM), highlighting the specific limitations of manual prompt dependency that the source solves.
- Paper: Swin-Unet: Unet-like Pure Transformer for Medical Image Segmentation, Hu Cao et al. (2021). Presents a transformer-based encoder-decoder architecture for medical image segmentation that serves as a core baseline and benchmark in the field.
- Paper: U-Net: Convolutional Networks for Biomedical Image Segmentation, Olaf Ronneberger et al. (2015). Introduces the classic multi-scale encoder-decoder segmentation framework and skip-connection concepts utilized by the source pixel decoder.
- Paper: SurgicalSAM: Efficient Class Promptable Surgical Instrument Segmentation, Wenxi Yue et al. (2024). Extends the concept of prompt-free, parameter-efficient SAM adaptation to specialized clinical video environments for surgical instrument segmentation.
