Unleashing the Potential of SAM for Medical Adaptation via Hierarchical Decoding

Zhiheng ChengQingyue WeiHongru ZhuYan WangLiangqiong QuWei ShaoYuyin Zhou

article2024CVPR98 citations

Proposes H-SAM, a prompt-free adaptation framework that equips the Segment Anything Model with a two-stage hierarchical mask decoder to achieve superior few-shot medical image segmentation performance using only minimal labeled data.

Listen

Precise medical image segmentation is critical for clinical diagnosis, treatment planning, and biomedical research, yet conventional deep learning models require vast amounts of costly, expert-annotated data. While large-scale vision foundation models like the Segment Anything Model offer powerful segmentation capabilities, their direct zero-shot performance drops significantly on specialized medical images. Adapting these models has typically required either expensive full fine-tuning or manual, expert-provided visual prompts such as bounding boxes or points during testing, which introduces clinical friction, latency, and potential human error.

The article demonstrates an efficient, prompt-free adaptation framework called H-SAM that tailors the Segment Anything Model for medical imaging using limited training data. By freezing the original image encoder and applying parameter-efficient tuning alongside a two-stage hierarchical decoding process, the model integrates learned medical priors to generate precise multi-class segmentations without needing manual prompts or massive annotated datasets.

To evaluate this framework, the authors conducted experiments across three benchmark medical imaging datasets: multi-organ abdominal CT scans from the Synapse dataset, cardiac MRI scans from the Left Atrial dataset, and prostate MRI scans from the PROMISE12 dataset. The architecture pairs parameter-efficient low-rank adaptation layers in the encoder with a two-stage mask decoder. The first decoding stage produces an initial coarse probabilistic mask, which then guides a second, refined decoding stage featuring class-balanced self-attention, learnable mask cross-attention, and a multi-scale pixel decoder with skip connections to capture fine anatomical details.

The experimental findings show substantial improvements across both few-shot and fully supervised settings. On multi-organ CT segmentation using only 10% of training slices, the proposed model achieved an 80.35% mean Dice score (an overlap accuracy metric where higher is better), outperforming competing prompt-free adaptation methods by nearly 5 percentage points and significantly lowering boundary errors. Under full supervision on the same dataset, it achieved an 86.49% mean Dice score, outperforming dedicated state-of-the-art medical segmentation architectures. Furthermore, in few-shot cardiac and prostate MRI segmentation using only 4 and 3 labeled training cases respectively, the model reached 89.22% and 87.27% accuracy, notably outperforming leading semi-supervised frameworks despite those methods relying on dozens of additional unlabeled scans.

These results indicate that foundation models can be effectively deployed in specialized clinical environments without the operational bottleneck of real-time expert prompting or the heavy computational overhead of full-model retraining. The ability to surpass complex semi-supervised approaches using only a minimal set of labeled scans highlights substantial potential to reduce data curation timelines, lower computational expenses, and mitigate label scarcity risks in healthcare artificial intelligence development.

Organizations developing clinical imaging pipelines should consider adopting two-stage, prior-guided fine-tuning strategies over prompt-dependent or data-heavy semi-supervised baselines. Future work should focus on validating the framework through clinical pilots across diverse hospital sites, evaluating its robustness on rare pathologies, and exploring extensions to native 3D volumetric architectures.

While the reported performance is strong, confidence should be framed within the context of the evaluation scope, which relied on standard public benchmark datasets and 2D slice processing rather than fully integrated prospective clinical workflows. External validation across broader multi-scanner cohorts is recommended prior to production deployment.

  • Paper: Segment Anything, Alexander M. Kirillov et al. (2023). Introduces the Segment Anything Model (SAM) architecture and promptable segmentation paradigm that the source model freezes and adapts.
  • Paper: Segment anything in medical images, Jun Ma et al. (2023). Establishes prompt-based adaptation of SAM for medical imaging (MedSAM), highlighting the specific limitations of manual prompt dependency that the source solves.
  • Paper: Swin-Unet: Unet-like Pure Transformer for Medical Image Segmentation, Hu Cao et al. (2021). Presents a transformer-based encoder-decoder architecture for medical image segmentation that serves as a core baseline and benchmark in the field.
  • Paper: U-Net: Convolutional Networks for Biomedical Image Segmentation, Olaf Ronneberger et al. (2015). Introduces the classic multi-scale encoder-decoder segmentation framework and skip-connection concepts utilized by the source pixel decoder.
Cover for Unleashing the Potential of SAM for Medical Adaptation via Hierarchical Decoding

Abstract

The Segment Anything Model (SAM) has garnered significant attention for its versatile segmentation abilities and intuitive prompt-based interface. However, its application in medical imaging presents challenges, requiring either substantial training costs and extensive medical datasets for full model fine-tuning or high-quality prompts for optimal performance. This paper introduces H-SAM: a prompt-free adaptation of SAM tailored for efficient fine-tuning of medical images via a two-stage hierarchical decoding procedure. In the initial stage, H-SAM employs SAM’s original decoder to generate a prior probabilistic mask, guiding a more intricate decoding process in the second stage. Specifically, we propose two key designs: 1) A class-balanced, mask-guided self-attention mechanism addressing the unbalanced label distribution, enhancing image embedding; 2) A learnable mask cross-attention mechanism spatially modulating the interplay among different image regions based on the prior mask. Moreover, the inclusion of a hierarchical pixel decoder in H-SAM enhances its proficiency in capturing fine-grained and localized details. This approach enables SAM to effectively integrate learned medical priors, facilitating enhanced adaptation for medical image segmentation with limited samples. Our H-SAM demonstrates a 4.78% improvement in average Dice compared to existing prompt-free SAM variants for multi-organ segmentation using only 10% of 2D slices. Notably, without using any unlabeled data, H-SAM even outperforms state-of-the-art semi-supervised models relying on extensive unlabeled training data across various medical datasets. Our code is available at https://github.com/Cccccczh404/H-SAM.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Methodology
  • 3.1. H-SAM Overview
  • 3.2. Enhanced Image Embedding via ClassBalanced Mask-Guided Self-Attention
  • 3.3. Learnable Mask Cross Attention
  • 3.4. Hierarchical pixel decoder
  • 4. Experiments
  • 4.1. Dataset and Evaluation
  • 4.2. Implementation details
  • 4.3. Results
  • 4.4. Ablation study
  • 4.5. Efficiency Analysis
  • 4.6. Qualitative Results
  • 5. Conclusion
  • References

Knowls

  1. Knowl 1 — H-SAM Two-Stage Hierarchical Decoding Framework

    model/method

    The Hierarchical Segment Anything Model (H-SAM) is a prompt-free medical adaptation framework for the Segment Anything Model (SAM). It adapts the frozen SAM vision transformer (ViT) image encoder by inserting Low-Rank Adaptation (LoRA) bypass matrices with rank r=4r=4 into transformer blocks while fine-tuning a default prompt embedding in place of manual point or bounding-box prompts.

    Rather than relying on SAM's standard single-pass mask decoder, H-SAM introduces a two-stage hierarchical decoder:

    1. Stage 1 (Prior Generation): Employs SAM's original lightweight mask decoder (consisting of a Transformer decoder and a pixel decoder) to predict an initial probabilistic prior mask MM at a spatial resolution of H/4×W/4H/4 \times W/4.
    2. Stage 2 (Refined Decoding): Re-decodes image features guided by the stage-1 prior. It incorporates Class-Balanced Mask-Guided Self-Attention (CMAttn) to enrich the image embeddings, Learnable Mask Cross-Attention within the transformer decoder to restrict attention to relevant organ boundaries, and a Hierarchical Pixel Decoder with skip connections to reconstruct high-resolution predictions at full resolution H×WH \times W.

    The final segmentation prediction is obtained by ensembling the two stages, averaging their predicted category probabilities.

  2. Knowl 2 — Learnable Mask Cross-Attention Mechanism

    equation

    To exploit spatial priors from the initial stage without suffering from gradient vanishing or unweighted foreground regions, H-SAM replaces standard mask-attention with a continuous, learnable mask cross-attention formulation inside the second-stage Transformer decoder:

    X=M⊙softmax(KQT)V+XX = M \odot \text{softmax}\left(K Q^T\right) V + X

    where:

    • X∈RL×CX \in \mathbb{R}^{L \times C} denotes the input query feature of the transformer block with sequence length LL and channel dimension CC.
    • Q,K,VQ, K, V denote the query, key, and value matrices computed during cross-attention.
    • MM is the untransformed, continuous probabilistic mask generated by the stage-1 decoder, bilinearly resized and flattened to match the spatial resolution of the cross-attention saliency map softmax(KQT)\text{softmax}(K Q^T).
    • ⊙\odot represents the element-wise Hadamard product.

    Unlike traditional mask-attention (e.g., in Mask2Former) that adds a hard threshold penalty t(M)∈{−∞,0}t(M) \in \{-\infty, 0\} to cross-attention logits (which zeroes out gradients through binarization and weights all foreground pixels equally), this formulation allows full gradient backpropagation through MM and softly modulates foreground regions based on predicted confidence.

  3. Knowl 3 — Class-Balanced Mask-Guided Self-Attention (CMAttn)

    model/method

    Class-Balanced Mask-Guided Self-Attention (CMAttn) is a module designed to recalibrate image embeddings entering the second-stage transformer decoder, specifically mitigating class imbalance between frequent head organ categories and rare tail organ categories.

    A mask feature is computed by directly taking the element-wise product of the image embedding and the un-upsampled stage-1 probabilistic mask. Prior to passing through self-attention, a class-balanced feature perturbation is applied:

    P(gt==i)+=N(0,var(i))P(\text{gt} == i) \mathrel{+}= \mathcal{N}(0, \text{var}(i))

    where P∈RN×C×H×WP \in \mathbb{R}^{N \times C \times H \times W} is the normalized mask feature, gt\text{gt} is the ground truth mask resized to the feature resolution, and N(0,var(i))\mathcal{N}(0, \text{var}(i)) is zero-mean Gaussian noise whose variance var(i)\text{var}(i) is inversely proportional to the sample frequency of category ii calculated offline across the dataset.

    The augmented mask feature passes through a self-attention layer and a feed-forward network (FFN). A linear projection then reduces its channel dimension, and the resulting feature is merged into the original image embedding via an element-wise Hadamard product combined with a residual skip connection.

  4. Knowl 4 — Hierarchical Pixel Decoder with Skip Connections

    model/method

    Standard SAM pixel decoders upsample the transformer output feature directly to a downscaled mask of resolution H/4×W/4H/4 \times W/4, limiting the delineation of fine structures and small-scale medical targets. H-SAM implements a hierarchical pixel decoder composed of two successive pixel decoders arranged with U-Net-style skip connections:

    • The first-stage pixel decoder processes stage-1 transformer outputs to generate low-resolution feature representations at H/4×W/4H/4 \times W/4.
    • The second-stage pixel decoder incorporates multiscale intermediate localized features from the first pixel decoder via skip connections and progressively upsamples the enriched object queries from the second transformer decoder to the original full image resolution H×WH \times W.

    This structure supplies localized spatial details to complement the global object queries without imposing the high computational cost of full-resolution attention across all transformer blocks.

  5. Knowl 5 — Deep Supervision Loss with Exponential Decay Weighting

    equation

    H-SAM applies multi-stage deep supervision across both decoding stages using a combination of binary cross-entropy loss Lce\mathcal{L}_{ce} and Dice loss Ldice\mathcal{L}_{dice}:

    L=λceLce+λdiceLdice\mathcal{L} = \lambda_{ce}\mathcal{L}_{ce} + \lambda_{dice}\mathcal{L}_{dice}

    The total training loss Ltotal\mathcal{L}_{total} balances stage-1 and stage-2 outputs:

    Ltotal=λwLstage1+(1−λw)Lstage2\mathcal{L}_{total} = \lambda_w \mathcal{L}_{stage1} + (1 - \lambda_w) \mathcal{L}_{stage2}

    where:

    • Lstage1\mathcal{L}_{stage1} supervises the stage-1 mask decoder output against ground-truth labels downsampled to H/4×W/4H/4 \times W/4.
    • Lstage2\mathcal{L}_{stage2} supervises the stage-2 mask decoder output against the original full-resolution ground truth (H×WH \times W).
    • λw\lambda_w is a time-varying weighting parameter initialized at 0.80.8 that decreases exponentially during training according to λw(t)=0.8⋅e−0.005⋅t\lambda_w(t) = 0.8 \cdot e^{-0.005 \cdot t}, gradually shifting the gradient focus from initial prior generation to fine-grained stage-2 segmentation.
  6. Knowl 6 — Multi-Organ CT Segmentation Performance on Synapse Dataset

    data/table

    Evaluation on the Synapse Multi-Organ CT dataset (18 training cases, 12 testing cases; 8 abdominal organs: Spleen, Right Kidney, Left Kidney, Gallbladder, Liver, Stomach, Aorta, Pancreas) compares H-SAM to state-of-the-art medical segmentation models and prompt-free SAM adaptation variants under few-shot (10% slice budget, 512×512512 \times 512 resolution) and fully supervised (100% budget, 224×224224 \times 224 resolution) settings.

    Setting / Method Spleen R.Kid L.Kid Gall. Liver Stom. Aorta Panc. Mean Dice (%) ↑\uparrow HD ↓\downarrow
    10% Data
    AutoSAM 68.80 77.44 76.53 24.87 88.06 52.70 75.19 34.58 55.69 31.67
    SAM Adapter 72.42 68.38 66.77 22.38 89.69 53.15 66.74 26.76 58.28 54.22
    SAMed 85.82 82.25 82.62 63.15 92.72 67.20 78.72 52.12 75.57 23.02
    H-SAM (Ours) 90.21 84.16 85.65 70.70 94.29 76.10 85.54 56.17 80.35 15.54
    Fully Supervised
    TransUnet 87.23 63.13 81.87 77.02 94.08 55.86 85.08 75.62 77.48 31.69
    SwinUnet 85.47 66.53 83.28 79.61 94.29 56.58 90.66 76.60 79.13 21.55
    TransDeepLab 86.04 69.16 84.08 79.88 93.53 61.19 89.00 78.40 80.16 21.25
    DAE-Former 88.96 72.30 86.08 80.88 94.98 65.12 91.94 79.19 82.43 17.46
    MERIT 92.01 84.85 87.79 74.40 95.26 85.38 87.71 71.81 84.90 13.22
    AutoSAM 80.54 80.02 79.60 41.37 89.24 61.14 82.56 44.22 62.08 27.56
    SAM Adapter 83.68 79.00 79.02 57.49 92.67 69.48 77.93 43.07 72.80 33.08
    SAMed 87.77 69.11 80.45 79.95 94.80 72.17 88.72 82.06 81.88 20.64
    H-SAM (Ours) 93.34 89.93 91.88 73.49 95.72 87.10 89.38 71.11 86.49 8.18

    In the 10% few-shot setting, H-SAM achieves 80.35% Mean Dice, outperforming SAMed by 4.78% and reducing the Hausdorff distance (HD) from 23.02 to 15.54. In the fully supervised setting, H-SAM reaches 86.49% Mean Dice (HD 8.18), surpassing both prompt-free SAM baselines and dedicated medical transformer architectures (e.g., MERIT at 84.90%).

  7. Knowl 7 — Few-Shot Segmentation Performance on LA and PROMISE12 Benchmarks

    data/table

    H-SAM was evaluated on the Left Atrial (LA) MRI dataset (trained using only 4 labeled scans / 5% data, 0 unlabeled scans) and the PROMISE12 Prostate MRI dataset (trained using only 3 labeled scans / 7.5% data, 0 unlabeled scans). Results are compared against semi-supervised methods that utilized the full remaining pool of unlabeled scans (76 and 37 scans, respectively) as well as supervised baselines.

    Left Atrial (LA) Dataset PROMISE12 Prostate Dataset
    Method Labeled Unlabeled Mean Dice (%) ↑\uparrow Method Labeled Unlabeled Mean Dice (%) ↑\uparrow
    UA-MT 4 (5%) 76 (95%) 82.26 UA-MT 3 (7.5%) 37 (92.5%) 65.05
    SASSNet 4 (5%) 76 (95%) 81.60 DTC 3 (7.5%) 37 (92.5%) 63.44
    DTC 4 (5%) 76 (95%) 81.25 SASSNet 3 (7.5%) 37 (92.5%) 73.43
    URPC 4 (5%) 76 (95%) 82.48 MC-Net 3 (7.5%) 37 (92.5%) 72.66
    MC-Net 4 (5%) 76 (95%) 83.59 SS-Net 3 (7.5%) 37 (92.5%) 73.19
    SS-Net 4 (5%) 76 (95%) 86.33 Self-Paced 3 (7.5%) 37 (92.5%) 74.02
    BCP 4 (5%) 76 (95%) 88.02 MLB-Seg 3 (7.5%) 37 (92.5%) 78.27
    nnU-Net 4 (5%) 0 (0%) 64.02 nnU-Net 3 (7.5%) 0 (0%) 84.22
    AutoSAM 4 (5%) 0 (0%) 74.73 AutoSAM 3 (7.5%) 0 (0%) 68.40
    SAM Adapter 4 (5%) 0 (0%) 82.79 SAM Adapter 3 (7.5%) 0 (0%) 75.45
    SAMed 4 (5%) 0 (0%) 87.72 SAMed 3 (7.5%) 0 (0%) 86.00
    H-SAM (Ours) 4 (5%) 0 (0%) 89.22 H-SAM (Ours) 3 (7.5%) 0 (0%) 87.27

    Without leveraging any unlabeled training scans, H-SAM achieves 89.22% Dice on LA and 87.27% Dice on PROMISE12, outperforming state-of-the-art semi-supervised frameworks (BCP at 88.02% and MLB-Seg at 78.27%) as well as prompt-free SAM adaptations (SAMed at 87.72% and 86.00%).

  8. Knowl 8 — Ablation of Hierarchical Decoding Modules in H-SAM

    data/table

    Ablation experiments on the 10% few-shot split of the Synapse dataset quantify the contribution of Learnable Mask Cross-Attention, Hierarchical Pixel Decoder, and Class-Balanced Mask-Guided Self-Attention (CMAttn).

    Learnable Mask-Attention Hierarchical Pixel Decoder CM Self-Attention Mean Dice (%) ↑\uparrow
    ✗ ✗ ✗ 75.57
    ✓ ✗ ✗ 77.68
    ✓ ✓ ✗ 78.58
    ✗ ✓ ✗ 77.05
    ✗ ✓ ✓ 79.03
    ✗ ✗ ✓ 77.71
    ✓ ✗ ✓ 78.76
    ✓ ✓ ✓ 80.35

    Direct comparison of mask-attention formulations shows that the unlearnable binary mask-attention formulation from Mask2Former yields virtually no gain over baseline cross-attention due to gradient vanishing, whereas learnable mask cross-attention provides an isolated gain of +2.11%:

    Attention Mechanism Mean Dice (%) ↑\uparrow
    Without mask-attention (Baseline) 75.57
    Original mask-attention (Mask2Former) 75.61
    Learnable mask-attention (H-SAM) 77.68

    When all three modules are combined, H-SAM achieves an overall improvement of 4.78% (from 75.57% to 80.35% Mean Dice).

  9. Knowl 9 — Parameter and Computational Efficiency of H-SAM

    data/table

    Efficiency analysis compares H-SAM against SAMed configured with varying decoder depths (2, 4, and 6 transformer layers) and SAM Adapter on the 10% Synapse few-shot task.

    Method Transformer Layers Total Parameters Mean Dice (%) ↑\uparrow
    SAMed 2 108.8M 75.57
    SAMed 4 112.5M 76.80
    SAMed 6 116.2M 78.05
    SAM Adapter 2 131.5M 72.80
    H-SAM (Ours) 4 112.3M 80.35

    While scaling standard SAMed decoder layers from 2 to 6 increases parameters to 116.2M and improves Dice to 78.05%, H-SAM achieves 80.35% Mean Dice with 4 total transformer layers and 112.3M parameters. Furthermore, H-SAM exceeds SAM Adapter by 7.55% Mean Dice while requiring 19.2M fewer parameters, demonstrating that the performance gain stems from the hierarchical prior-guided design rather than mere parameter expansion.

Coverage note — None was omitted; all key architectural components, equations, training setups, dataset evaluations, ablations, and efficiency comparisons are covered.

References

  1. 1.Reza Azad, Moein Heidari, Moein Shariatnia, Ehsan Khodapanah Aghdam, Sanaz Karimijafarbigloo, Ehsan Adeli, and Dorit Merhof. Transdeeplab: Convolution-free transformer-based deeplab v3+ for medical image segmentation. In International Workshop on PRedictive Intelligence In MEdicine, pages 91–102. Springer, 2022.
  2. 2.Reza Azad, Rene Arimond, Ehsan Khodapanah Aghdam, Amirhossein Kazerouni, and Dorit Merhof. Dae-former: Dual attention-guided efficient transformer for medical image segmentation. In International Workshop on PRedictive Intelligence In MEdicine, pages 83–95. Springer, 2023.
  3. 3.Yunhao Bai, Duowen Chen, Qingli Li, Wei Shen, and Yan Wang. Bidirectional copy-paste for semi-supervised medical image segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 11514–11524, 2023.
  4. 4.Hu Cao, Yueyue Wang, Joy Chen, Dongsheng Jiang, Xiaopeng Zhang, Qi Tian, and Manning Wang. Swin-unet: Unet-like pure transformer for medical image segmentation. In European conference on computer vision, pages 205–218. Springer, 2022.
  5. 5.Shurong Chai, Rahul Kumar Jain, Shiyu Teng, Jiaqing Liu, Yinhao Li, Tomoko Tateyama, and Yen-wei Chen. Ladder fine-tuning approach for sam integrating complementary network. arXiv preprint arXiv:2306.12737, 2023.
  6. 6.Chen Chen, Wenjia Bai, and Daniel Rueckert. Multi-task learning for left atrial segmentation on ge-mri. In Statistical Atlases and Computational Models of the Heart. Atrial Segmentation and LV Quantification Challenges: 9th International Workshop, STACOM 2018, Held in Conjunction with MICCAI 2018, Granada, Spain, September 16, 2018, Revised Selected Papers 9, pages 292–301. Springer, 2019.
  7. 7.Jieneng Chen, Yongyi Lu, Qihang Yu, Xiangde Luo, Ehsan Adeli, Yan Wang, Le Lu, Alan L Yuille, and Yuyin Zhou. Transunet: Transformers make strong encoders for medical image segmentation. arXiv preprint arXiv:2102.04306, 2021.
  8. 8.Jieneng Chen, Jieru Mei, Xianhang Li, Yongyi Lu, Qihang Yu, Qingyue Wei, Xiangde Luo, Yutong Xie, Ehsan Adeli, Yan Wang, et al. 3d transunet: Advancing medical image segmentation through vision transformers. arXiv preprint arXiv:2310.07781, 2023.
  9. 9.Tianrun Chen, Lanyun Zhu, Chaotao Ding, Runlong Cao, Shangzhan Zhang, Yan Wang, Zejian Li, Lingyun Sun, Papa Mao, and Ying Zang. Sam fails to segment anything? – sam-adapter: Adapting sam in underperformed scenes: Camouflage, shadow, and more, 2023.
  10. 10.Zhe Chen, Yuchen Duan, Wenhai Wang, Junjun He, Tong Lu, Jifeng Dai, and Yu Qiao. Vision transformer adapter for dense predictions. arXiv preprint arXiv:2205.08534, 2022.
  11. 11.Bowen Cheng, Ishan Misra, Alexander G Schwing, Alexander Kirillov, and Rohit Girdhar. Masked-attention mask transformer for universal image segmentation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 1290–1299, 2022.
  12. 12.Dongjie Cheng, Ziyuan Qin, Zekun Jiang, Shaoting Zhang, Qicheng Lao, and Kang Li. Sam on medical images: A comprehensive study on three prompt modes. arXiv preprint arXiv:2305.00035, 2023.
  13. 13.Junlong Cheng, Shengwei Tian, Long Yu, Chengrui Gao, Xiaojing Kang, Xiang Ma, Weidong Wu, Shijia Liu, and Hongchun Lu. Resganet: Residual group attention network for medical image classification and segmentation. Medical Image Analysis, 76:102313, 2022.
  14. 14.Can Cui, Ruining Deng, Quan Liu, Tianyuan Yao, Shunxing Bao, Lucas W Remedios, Yucheng Tang, and Yuankai Huo. All-in-sam: from weak annotation to pixel-wise nuclei segmentation with prompt-based finetuning. arXiv preprint arXiv:2307.00290, 2023.
  15. 15.Jeffrey De Fauw, Joseph R Ledsam, Bernardino Romera-Paredes, Stanislav Nikolov, Nenad Tomasev, Sam Blackwell, Harry Askham, Xavier Glorot, Brendan O’Donoghue, Daniel Visentin, et al. Clinically applicable deep learning for diagnosis and referral in retinal disease. Nature medicine, 24(9):1342–1350, 2018.
  16. 16.Guoyao Deng, Ke Zou, Kai Ren, Meng Wang, Xuedong Yuan, Sancong Ying, and Huazhu Fu. Sam-u: Multi-box prompts triggered uncertainty estimation for reliable sam in medical image. arXiv preprint arXiv:2307.04973, 2023.
  17. 17.Ruining Deng, Can Cui, Quan Liu, Tianyuan Yao, Lucas W Remedios, Shunxing Bao, Bennett A Landman, Lee E Wheless, Lori A Coburn, Keith T Wilson, et al. Segment anything model (sam) for digital pathology: Assess zero-shot segmentation on whole slide imaging. arXiv preprint arXiv:2304.04155, 2023.
  18. 18.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018.
  19. 19.Foivos I Diakogiannis, Francois Waldner, Peter Caccetta, and Chen Wu. Resunet-a: A deep learning framework for semantic segmentation of remotely sensed data. ISPRS Journal of Photogrammetry and Remote Sensing, 162:94–114, 2020.
  20. 20.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929, 2020.
  21. 21.Weijia Feng, Lingting Zhu, and Lequan Yu. Cheap lunch for medical image segmentation by fine-tuning sam on few exemplars. arXiv preprint arXiv:2308.14133, 2023.
  22. 22.Yabo Fu, Yang Lei, Tonghe Wang, Walter J Curran, Tian Liu, and Xiaofeng Yang. A review of deep learning based methods for medical image multi-organ segmentation. Physica Medica, 85:107–122, 2021.
  23. 23.Yifan Gao, Wei Xia, Dingdu Hu, and Xin Gao. Desam: Decoupling segment anything model for generalizable medical image segmentation. arXiv preprint arXiv:2306.00499, 2023.
  24. 24.Shizhan Gong, Yuan Zhong, Wenao Ma, Jinpeng Li, Zhao Wang, Jingyang Zhang, Pheng-Ann Heng, and Qi Dou. 3dsam-adapter: Holistic adaptation of sam from 2d to 3d for promptable medical image segmentation. arXiv preprint arXiv:2306.13465, 2023.
  25. 25.Sheng He, Rina Bao, Jingpeng Li, P Ellen Grant, and Yangming Ou. Accuracy of segment-anything model (sam) in medical image segmentation tasks. arXiv preprint arXiv:2304.09324, 2023.
  26. 26.Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin De Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. Parameter-efficient transfer learning for nlp. In International Conference on Machine Learning, pages 2790–2799. PMLR, 2019.
  27. 27.Chuanfei Hu and Xinde Li. When sam meets medical images: An investigation of segment anything model (sam) on multi-phase liver tumor segmentation. arXiv preprint arXiv:2304.08506, 2023.
  28. 28.Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. Lora: Low-rank adaptation of large language models. arXiv preprint arXiv:2106.09685, 2021.
  29. 29.Mingzhe Hu, Yuheng Li, and Xiaofeng Yang. Skinsam: Empowering skin cancer segmentation with segment anything model. arXiv preprint arXiv:2304.13973, 2023.
  30. 30.Xinrong Hu, Xiaowei Xu, and Yiyu Shi. How to efficiently adapt large segmentation model (sam) to medical images. arXiv preprint arXiv:2306.13731, 2023.
  31. 31.Xiaohong Huang, Zhifang Deng, Dandan Li, and Xueguang Yuan. Missformer: An effective medical image segmentation transformer. arXiv preprint arXiv:2109.07162, 2021.
  32. 32.Yuhao Huang, Xin Yang, Lian Liu, Han Zhou, Ao Chang, Xinrui Zhou, Rusi Chen, Junxuan Yu, Jiongquan Chen, Chaoyu Chen, et al. Segment anything model for medical images? arXiv preprint arXiv:2304.14660, 2023.
  33. 33.Fabian Isensee, Jens Petersen, Andre Klein, David Zimmerer, Paul F Jaeger, Simon Kohl, Jakob Wasserthal, Gregor Koehler, Tobias Norajitra, Sebastian Wirkert, et al. nnu-net: Self-adapting framework for u-net-based medical image segmentation. arXiv preprint arXiv:1809.10486, 2018.
  34. 34.Fabian Isensee, Paul F Jaeger, Simon AA Kohl, Jens Petersen, and Klaus H Maier-Hein. nnu-net: a self-configuring method for deep learning-based biomedical image segmentation. Nature methods, 18(2):203–211, 2021.
  35. 35.Ge-Peng Ji, Deng-Ping Fan, Peng Xu, Ming-Ming Cheng, Bowen Zhou, and Luc Van Gool. Sam struggles in concealed scenes–empirical study on’’ segment anything’’. arXiv preprint arXiv:2304.06022, 2023.
  36. 36.Yuanfeng Ji, Haotian Bai, Chongjian Ge, Jie Yang, Ye Zhu, Ruimao Zhang, Zhen Li, Lingyan Zhanng, Wanling Ma, Xiang Wan, et al. Amos: A large-scale abdominal multi-organ benchmark for versatile medical image segmentation. Advances in Neural Information Processing Systems, 35:36722–36732, 2022.
  37. 37.Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C Berg, Wan-Yen Lo, et al. Segment anything. arXiv preprint arXiv:2304.02643, 2023.
  38. 38.Bennett Landman, Zhoubing Xu, J Igelsias, Martin Styner, T Langerak, and Arno Klein. Miccai multi-atlas labeling beyond the cranial vault–workshop and challenge. In Proc. MICCAI Multi-Atlas Labeling Beyond Cranial Vault—Workshop Challenge, page 12, 2015.
  39. 39.Wenhui Lei, Xu Wei, Xiaofan Zhang, Kang Li, and Shaoting Zhang. Medlsam: Localize and segment anything model for 3d medical images. arXiv preprint arXiv:2306.14752, 2023.
  40. 40.Chengyin Li, Prashant Khanduri, Yao Qiang, Rafi Ibn Sultan, Indrin Chetty, and Dongxiao Zhu. Auto-prompting sam for mobile friendly 3d medical image segmentation. arXiv preprint arXiv:2308.14936, 2023.
  41. 41.Mengke Li, Yiu-ming Cheung, and Yang Lu. Long-tailed visual recognition via gaussian clouded logit adjustment. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6929–6938, 2022.
  42. 42.Shuailin Li, Chuyu Zhang, and Xuming He. Shape-aware semi-supervised 3d semantic segmentation for medical images. In Medical Image Computing and Computer Assisted Intervention–MICCAI 2020: 23rd International Conference, Lima, Peru, October 4–8, 2020, Proceedings, Part I 23, pages 552–561. Springer, 2020.
  43. 43.Yuheng Li, Mingzhe Hu, and Xiaofeng Yang. Polypsam: Transfer sam for polyp segmentation. arXiv preprint arXiv:2305.00293, 2023.
  44. 44.Yi Lin, Yufan Chen, Kwang-Ting Cheng, and Hao Chen. Few shot medical image segmentation with cross attention transformer. arXiv preprint arXiv:2303.13867, 2023.
  45. 45.Geert Litjens, Robert Toth, Wendy Van De Ven, Caroline Hoeks, Sjoerd Kerkstra, Bram Van Ginneken, Graham Vincent, Gwenael Guillard, Neil Birbeck, Jindang Zhang, et al. Evaluation of prostate segmentation algorithms for mri: the promise12 challenge. Medical image analysis, 18(2):359–373, 2014.
  46. 46.Yihao Liu, Jiaming Zhang, Zhangcong She, Amir Kheradmand, and Mehran Armand. Samm (segment any medical model): A 3d slicer integration to sam. arXiv preprint arXiv:2304.05622, 2023.
  47. 47.Xiangde Luo, Jieneng Chen, Tao Song, and Guotai Wang. Semi-supervised medical image segmentation through dual-task consistency. In Proceedings of the AAAI conference on artificial intelligence, pages 8801–8809, 2021.
  48. 48.Xiangde Luo, Wenjun Liao, Jieneng Chen, Tao Song, Yinan Chen, Shichuan Zhang, Nianyong Chen, Guotai Wang, and Shaoting Zhang. Efficient semi-supervised gross target volume of nasopharyngeal carcinoma segmentation via uncertainty rectified pyramid consistency. In Medical Image Computing and Computer Assisted Intervention–MICCAI 2021: 24th International Conference, Strasbourg, France, September 27–October 1, 2021, Proceedings, Part II 24, pages 318–329. Springer, 2021.
  49. 49.Jun Ma and Bo Wang. Segment anything in medical images. arXiv preprint arXiv:2304.12306, 2023.
  50. 50.Maciej A Mazurowski, Haoyu Dong, Hanxue Gu, Jichen Yang, Nicholas Konz, and Yixin Zhang. Segment anything model for medical image analysis: an experimental study. Medical Image Analysis, 89:102918, 2023.
  51. 51.Fausto Milletari, Nassir Navab, and Seyed-Ahmad Ahmadi. V-net: Fully convolutional neural networks for volumetric medical image segmentation. In 2016 fourth international conference on 3D vision (3DV), pages 565–571. Ieee, 2016.
  52. 52.S Mohapatra, A Gosai, and G Schlaug. Sam vs bet: A comparative study for brain extraction and segmentation of magnetic resonance images using deep learning. arXiv preprint arXiv:2304.04738, 2:4, 2023.
  53. 53.David Ouyang, Bryan He, Amirata Ghorbani, Neal Yuan, Joseph Ebinger, Curtis P Langlotz, Paul A Heidenreich, Robert A Harrington, David H Liang, Euan A Ashley, et al. Video-based ai for beat-to-beat assessment of cardiac function. Nature, 580(7802):252–256, 2020.
  54. 54.Jizong Peng, Ping Wang, Christian Desrosiers, and Marco Pedersoli. Self-paced contrastive learning for semi-supervised medical image segmentation with meta-labels. Advances in Neural Information Processing Systems, 34:16686–16699, 2021.
  55. 55.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In International conference on machine learning, pages 8748–8763. PMLR, 2021.
  56. 56.Md Mostafijur Rahman and Radu Marculescu. G-cascade: Efficient cascaded graph convolutional decoding for 2d medical image segmentation. arXiv preprint arXiv:2310.16175, 2023.
  57. 57.Md Mostafijur Rahman and Radu Marculescu. Multi-scale hierarchical vision transformer with cascaded attention decoding for medical image segmentation. arXiv preprint arXiv:2303.16892, 2023.
  58. 58.Olaf Ronneberger, Philipp Fischer, and Thomas Brox. U-net: Convolutional networks for biomedical image segmentation. In Medical Image Computing and Computer-Assisted Intervention–MICCAI 2015: 18th International Conference, Munich, Germany, October 5-9, 2015, Proceedings, Part III 18, pages 234–241. Springer, 2015.
  59. 59.Hao Tang, Xingwei Liu, Shanlin Sun, Xiangyi Yan, and Xiaohui Xie. Recurrent mask refinement for few-shot medical image segmentation. In Proceedings of the IEEE/CVF international conference on computer vision, pages 3918–3928, 2021.
  60. 60.Tassilo Wald, Saikat Roy, Gregor Koehler, Nico Disch, Maximilian Rouven Rokuss, Julius Holzschuh, David Zimmerer, and Klaus Maier-Hein. Sam. md: Zero-shot medical image segmentation capabilities of the segment anything model. In Medical Imaging with Deep Learning, short paper track, 2023.
  61. 61.An Wang, Mobarakol Islam, Mengya Xu, Yang Zhang, and Hongliang Ren. Sam meets robotic surgery: An empirical study in robustness perspective. arXiv preprint arXiv:2304.14674, 2023.
  62. 62.Yan Wang, Yuyin Zhou, Wei Shen, Seyoun Park, Elliot K Fishman, and Alan L Yuille. Abdominal multi-organ segmentation with organ-attention networks and statistical fusion. Medical image analysis, 55:88–102, 2019.
  63. 63.Qingyue Wei, Lequan Yu, Xianhang Li, Wei Shao, Cihang Xie, Lei Xing, and Yuyin Zhou. Consistency-guided meta-learning for bootstrapping semi-supervised medical image segmentation. In International Conference on Medical Image Computing and Computer-Assisted Intervention, pages 183–193. Springer, 2023.
  64. 64.Junde Wu, Rao Fu, Huihui Fang, Yuanpei Liu, Zhaowei Wang, Yanwu Xu, Yueming Jin, and Tal Arbel. Medical sam adapter: Adapting segment anything model for medical image segmentation. arXiv preprint arXiv:2304.12620, 2023.
  65. 65.Yicheng Wu, Minfeng Xu, Zongyuan Ge, Jianfei Cai, and Lei Zhang. Semi-supervised left atrium segmentation with mutual consistency training. In Medical Image Computing and Computer Assisted Intervention–MICCAI 2021: 24th International Conference, Strasbourg, France, September 27–October 1, 2021, Proceedings, Part II 24, pages 297–306. Springer, 2021.
  66. 66.Yicheng Wu, Zhonghua Wu, Qianyi Wu, Zongyuan Ge, and Jianfei Cai. Exploring smoothness and class-separation for semi-supervised medical image segmentation. In International Conference on Medical Image Computing and Computer-Assisted Intervention, pages 34–43. Springer, 2022.
  67. 67.Lequan Yu, Shujun Wang, Xiaomeng Li, Chi-Wing Fu, and Pheng-Ann Heng. Uncertainty-aware self-ensembling model for semi-supervised 3d left atrium segmentation. In Medical Image Computing and Computer Assisted Intervention–MICCAI 2019: 22nd International Conference, Shenzhen, China, October 13–17, 2019, Proceedings, Part II 22, pages 605–613. Springer, 2019.
  68. 68.Chenyi Zeng, Lin Gu, Zhenzhong Liu, and Shen Zhao. Review of deep learning approaches for the segmentation of multiple sclerosis lesions on brain mri. Frontiers in Neuroinformatics, 14:610967, 2020.
  69. 69.Jingwei Zhang, Ke Ma, Saarthak Kapse, Joel Saltz, Maria Vakalopoulou, Prateek Prasanna, and Dimitris Samaras. Sam-path: A segment anything model for semantic segmentation in digital pathology. arXiv preprint arXiv:2307.09570, 2023.
  70. 70.Kaidong Zhang and Dong Liu. Customized segment anything model for medical image segmentation. arXiv preprint arXiv:2304.13785, 2023.
  71. 71.Lian Zhang, Zhengliang Liu, Lu Zhang, Zihao Wu, Xiaowei Yu, Jason Holmes, Hongying Feng, Haixing Dai, Xiang Li, Quanzheng Li, et al. Segment anything model (sam) for radiation oncology. arXiv preprint arXiv:2306.11730, 2023.
  72. 72.Yiming Zhang, Tianang Leng, Kun Han, and Xiaohui Xie. Self-sampling meta sam: Enhancing few-shot medical image segmentation with meta-learning. arXiv preprint arXiv:2308.16466, 2023.
  73. 73.Yizhe Zhang, Tao Zhou, Peixian Liang, and Danny Z Chen. Input augmentation with sam: Boosting medical image segmentation with segmentation foundation model. arXiv preprint arXiv:2304.11332, 2023.
  74. 74.Yuyin Zhou, Zhe Li, Song Bai, Chong Wang, Xinlei Chen, Mei Han, Elliot Fishman, and Alan L Yuille. Prior-aware neural network for partially-supervised multi-organ segmentation. In Proceedings of the IEEE/CVF international conference on computer vision, pages 10672–10681, 2019.
  75. 75.Yuyin Zhou, Yan Wang, Peng Tang, Song Bai, Wei Shen, Elliot Fishman, and Alan Yuille. Semi-supervised 3d abdominal multi-organ segmentation via deep multi-planar co-training. In 2019 IEEE Winter Conference on Applications of Computer Vision (WACV), pages 121–140. IEEE, 2019.
  76. 76.Zongwei Zhou, Md Mahfuzur Rahman Siddiquee, Nima Tajbakhsh, and Jianming Liang. Unet++: A nested u-net architecture for medical image segmentation. In Deep Learning in Medical Image Analysis and Multimodal Learning for Clinical Decision Support: 4th International Workshop, DLMIA 2018, and 8th International Workshop, ML-CDS 2018, Held in Conjunction with MICCAI 2018, Granada, Spain, September 20, 2018, Proceedings 4, pages 3–11. Springer, 2018.

Citation

MLA
Cheng, Z., et al. “Unleashing the Potential of SAM for Medical Adaptation via Hierarchical Decoding”. arXiv, 2024, http://arxiv.org/abs/2403.18271v1.
APA
Cheng, Z., Wei, Q., Zhu, H., Wang, Y., Qu, L., Shao, W., & Zhou, Y. (2024). Unleashing the Potential of SAM for Medical Adaptation via Hierarchical Decoding. arXiv. http://arxiv.org/abs/2403.18271v1
Chicago
Cheng, Z., Q. Wei, H. Zhu, et al. 2024. “Unleashing the Potential of SAM for Medical Adaptation via Hierarchical Decoding”. arXiv. http://arxiv.org/abs/2403.18271v1.
Harvard
Cheng, Z. et al. (2024) “Unleashing the Potential of SAM for Medical Adaptation via Hierarchical Decoding”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2403.18271v1.
Vancouver
1. Cheng Z, Wei Q, Zhu H, Wang Y, Qu L, Shao W, Zhou Y (2024) Unleashing the Potential of SAM for Medical Adaptation via Hierarchical Decoding. arXiv

BibTeX

@article{cheng2024unleashing,
  title = {Unleashing the Potential of SAM for Medical Adaptation via Hierarchical Decoding},
  author = {Cheng, Zhiheng and Wei, Qingyue and Zhu, Hongru and Wang, Yan and Qu, Liangqiong and Shao, Wei and Zhou, Yuyin},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2403.18271v1},
  eprint = {2403.18271}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE