SurgicalSAM: Efficient Class Promptable Surgical Instrument Segmentation

Wenxi YueJing ZhangKun HuYong XiaJiebo LuoZhiyong Wang

article2024AAAI134 citations

Proposes an end-to-end framework that efficiently adapts the Segment Anything Model for surgical instrument segmentation by replacing fragile manual bounding-box prompts with learned class prototypes and contrastive learning.

Listen

Accurate identification and segmentation of surgical instruments in video feeds is essential for advancing computer-assisted surgery and improving operating room safety. While large artificial intelligence foundation models like the Segment Anything Model offer strong general segmentation capabilities, directly applying them to surgery has proven difficult. Surgical instruments differ substantially from everyday objects, share high visual similarities across categories, and standard foundation models rely heavily on precise manual points or bounding boxes that are impractical during live clinical workflows.

To overcome these limitations, the article evaluates and demonstrates SurgicalSAM, a framework designed to efficiently adapt foundation segmentation models to surgical environments without requiring manual spatial prompts. The objective is to achieve state-of-the-art segmentation accuracy across instrument categories while keeping computational demands low and streamlining deployment into a single, automated step.

The authors developed an end-to-end tuning method that uses category labels as prompts instead of manual bounding boxes. By keeping the massive image-processing backbone frozen and only training a lightweight prompt encoder and mask decoder, the model learns category-specific visual prototypes. A contrastive learning technique ensures these prototypes clearly distinguish between visually similar tools. The system was validated against existing specialist architectures and foundation model baselines across benchmark surgical video datasets, specifically EndoVis2017 and EndoVis2018.

The evaluation revealed several key findings. First, SurgicalSAM matched or outperformed leading specialist models across both datasets while tuning only 4.65 million parameters, compared to nearly 69 million in competing state-of-the-art models. Second, it demonstrated superior generalisation when trained on one dataset and tested on another, achieving an 11.43 percentage-point gain in cross-dataset accuracy over top specialist architectures. Third, training efficiency increased dramatically: training was more than 10 times faster than comparable specialized models, utilizing less than one-sixth of the graphic processor memory. Finally, eliminating the need for precise spatial coordinate prompts made the system significantly more robust against noise compared to zero-shot foundation models.

These results demonstrate that adapting foundation models using lightweight, class-based tuning is a highly viable path for medical artificial intelligence. The approach drastically reduces development and compute costs while eliminating the latency and failure points associated with multi-stage detection systems. Enhanced cross-dataset reliability suggests these models can better adapt to variations across hospital setups and surgical tools.

Engineering and clinical teams should consider shifting from training large custom models from scratch toward adapting general foundation models using class-prototype strategies. Before broad clinical integration, next steps should include testing on larger, multicenter clinical datasets with greater procedural variety to establish performance across diverse operating conditions.

The findings are supported by strong benchmark results, but confidence should be framed around the dataset scale, as the benchmarks encompass limited video quantities and fixed instrument categories. Continued validation in real-time robotic surgery environments will be necessary before operational deployment.

  • Paper: Segment Anything, Alexander M. Kirillov et al. (2023). Introduces the Segment Anything Model (SAM) foundation architecture and promptable segmentation formulation that SurgicalSAM directly adapts for surgical instrument segmentation.
  • Paper: Segment anything in medical images, Jun Ma et al. (2023). Demonstrates the foundational adaptation of SAM to medical imaging via bounding box prompts and decoder tuning, establishing the domain-specific fine-tuning paradigm that SurgicalSAM builds upon.
  • Paper: PANet: Few-Shot Image Semantic Segmentation With Prototype Alignment, Kaixin Wang et al. (2019). Establishes prototype-based metric learning and class alignment representations for segmentation, which directly inform SurgicalSAM's category-level visual prototype formulation.
Cover for SurgicalSAM: Efficient Class Promptable Surgical Instrument Segmentation

Abstract

The Segment Anything Model (SAM) is a powerful foundation model that has revolutionised image segmentation. To apply SAM to surgical instrument segmentation, a common approach is to locate precise points or boxes of instruments and then use them as prompts for SAM in a zero-shot manner. However, we observe two problems with this naive pipeline: (1) the domain gap between natural objects and surgical instruments leads to inferior generalisation of SAM; and (2) SAM relies on precise point or box locations for accurate segmentation, requiring either extensive manual guidance or a well-performing specialist detector for prompt preparation, which leads to a complex multi-stage pipeline. To address these problems, we introduce SurgicalSAM, a novel end-to-end efficient-tuning approach for SAM to effectively integrate surgical-specific information with SAM's pre-trained knowledge for improved generalisation. Specifically, we propose a lightweight prototype-based class prompt encoder for tuning, which directly generates prompt embeddings from class prototypes and eliminates the use of explicit prompts for improved robustness and a simpler pipeline. In addition, to address the low inter-class variance among surgical instrument categories, we propose contrastive prototype learning, further enhancing the discrimination of the class prototypes for more accurate class prompting. The results of extensive experiments on both EndoVis2018 and EndoVis2017 datasets demonstrate that SurgicalSAM achieves state-of-the-art performance while only requiring a small number of tunable parameters. The source code is available at https://github.com/wenxi-yue/SurgicalSAM.

Table of Contents

  • Introduction
  • Related Work
  • Surgical Instrument Segmentation
  • Segment Anything Model
  • Methodology
  • Overview
  • Prototype-based Class Prompt Encoder
  • Contrastive Prototype Learning
  • Efficient Tuning
  • Experiments and Discussion Datasets and Evaluation
  • Implementation Details
  • Main Results
  • Ablation Study
  • Cross-Dataset Generalisation
  • Complexity Analysis
  • Conclusion
  • Acknowledgements
  • References

Knowls

  1. Knowl 1 — SurgicalSAM Framework Architecture

    model/method

    SurgicalSAM adapts the Segment Anything Model (SAM) for class-promptable surgical instrument segmentation. Given an input surgical image I∈RH×W×3I \in \mathbb{R}^{H \times W \times 3} and a target instrument category c∈{1,2,…,C}c \in \{1, 2, \dots, C\}, SurgicalSAM predicts the class-specific binary segmentation mask M(c)M^{(c)}.

    The framework consists of three main modules:

    1. Image Encoder (EIE_I): A frozen Vision Transformer (ViT-H) from SAM that extracts spatial image features: FI=EI(I)∈Rh×w×dF_I = E_I(I) \in \mathbb{R}^{h \times w \times d} where h×wh \times w represents the spatial dimensions of the feature map and dd is the channel dimension.

    2. Prototype-based Class Prompt Encoder (ECPE_{CP}): Takes the image feature FIF_I, a learnable prototype bank B∈RC×dB \in \mathbb{R}^{C \times d}, and the requested prompt class index cc to generate dense prompt embeddings TD(c)∈Rh×w×dT_D^{(c)} \in \mathbb{R}^{h \times w \times d} and sparse prompt embeddings TS(c)∈RC×n×dT_S^{(c)} \in \mathbb{R}^{C \times n \times d}: TD(c),TS(c)=ECP(FI,B,c)T_D^{(c)}, T_S^{(c)} = E_{CP}(F_I, B, c)

    3. Mask Decoder (DMD_M): A tuned decoder adapted from SAM that consumes the image feature FIF_I, the generated prompt embeddings, and SAM's learnable output tokens TOT_O to output the target mask: M(c)=DM(FI,[TD(c),TS(c),TO])M^{(c)} = D_M(F_I, [T_D^{(c)}, T_S^{(c)}, T_O])

    By generating latent prompt embeddings directly from learned class prototypes, SurgicalSAM operates end-to-end without requiring manual visual prompts (points or bounding boxes) or separate object detection models.

  2. Knowl 2 — Prototype-Based Class Prompt Encoder

    model/method

    The Prototype-Based Class Prompt Encoder (ECPE_{CP}) synthesizes dense and sparse prompt embeddings directly from a prototype bank B=concat({B(k)}k=1C)∈RC×dB = \text{concat}(\{B^{(k)}\}_{k=1}^C) \in \mathbb{R}^{C \times d}, where each B(k)∈RdB^{(k)} \in \mathbb{R}^d is a representative vector for instrument class kk.

    1. Class-Specific Spatial Activation: The spatial similarity map S(k)∈Rh×wS^{(k)} \in \mathbb{R}^{h \times w} for each class kk is obtained via dot product between the image feature FI∈Rh×w×dF_I \in \mathbb{R}^{h \times w \times d} and the prototype B(k)B^{(k)}: S(k)=FI×B(k),for k∈{1,…,C}S^{(k)} = F_I \times B^{(k)}, \quad \text{for } k \in \{1, \dots, C\} Each class feature is then activated by using S(k)S^{(k)} as a spatial attention map: FI(k)=FI∘S(k)+FIF_I^{(k)} = F_I \circ S^{(k)} + F_I where ∘\circ and ++ denote element-wise multiplication and addition. Stacking these yields the full class-activated feature tensor FIC=concat({FI(k)}k=1C)∈RC×h×w×dF_I^C = \text{concat}(\{F_I^{(k)}\}_{k=1}^C) \in \mathbb{R}^{C \times h \times w \times d}.

    2. Dense Prompt Embeddings: Dense prompt embeddings encode foreground spatial cues for the target prompted class cc. They are generated from the activated feature FI(c)F_I^{(c)} through a two-layer Multi-Layer Perceptron (MLP) with intermediate projection dimension rDr_D: TD(c)=gD(ReLU(fD(FI(c))))∈Rh×w×dT_D^{(c)} = g_D(\text{ReLU}(f_D(F_I^{(c)}))) \in \mathbb{R}^{h \times w \times d}

    3. Positivity-Aware Sparse Prompt Embeddings: Sparse prompt embeddings incorporate features from all classes (prompted positive class and unprompted negative classes). First, class-activated features FICF_I^C pass through an MLP with intermediate dimension rSr_S to yield positivity-agnostic sparse prompt tokens T^SC=concat({T^S(k)}k=1C)∈RC×n×d\hat{T}_S^C = \text{concat}(\{\hat{T}_S^{(k)}\}_{k=1}^C) \in \mathbb{R}^{C \times n \times d}, where nn is the number of tokens per class: T^SC=gS(ReLU(fS(FIC)))\hat{T}_S^C = g_S(\text{ReLU}(f_S(F_I^C))) Then, a frozen positive vector λ+∈Rd\lambda^+ \in \mathbb{R}^d and a frozen negative vector λ−∈Rd\lambda^- \in \mathbb{R}^d (derived from SAM's pre-trained point/box embeddings) are added to distinguish class cc from all other classes k≠ck \neq c: TS(c)=concat({T^S(k)+1(k=c)λ++(1−1(k=c))λ−}k=1C)∈RC×n×dT_S^{(c)} = \text{concat}\left(\left\{\hat{T}_S^{(k)} + \mathbf{1}(k = c)\lambda^+ + (1 - \mathbf{1}(k = c))\lambda^-\right\}_{k=1}^C\right) \in \mathbb{R}^{C \times n \times d} The sparse embeddings are reshaped to (C⋅n)×d(C \cdot n) \times d before being fed into the mask decoder.

  3. Knowl 3 — Contrastive Prototype Learning for Class Prototypes

    model/method

    To distinguish fine-grained surgical instrument categories with high visual similarity, SurgicalSAM trains the class prototypes B(k)∈RdB^{(k)} \in \mathbb{R}^d (k=1,…,Ck=1,\dots,C) using a prototype contrastive loss LPCL\mathcal{L}_{PCL}.

    During training, class prototype B(k)B^{(k)} acts as an anchor, and SAM-based foreground feature averages extracted from training images act as representation samples. For an image with ground-truth binary mask G(c)G^{(c)} of class cc downsampled to h×wh \times w, the image-level class embedding v(c)∈Rdv^{(c)} \in \mathbb{R}^d is computed via average pooling over the foreground: v(c)=∑i=1hw(FI∘G(c))i∑i=1hwGi(c)v^{(c)} = \frac{\sum_{i=1}^{hw} (F_I \circ G^{(c)})_i}{\sum_{i=1}^{hw} G^{(c)}_i} where FI∈Rh×w×dF_I \in \mathbb{R}^{h \times w \times d} is the image feature.

    The prototype contrastive loss is formulated as an InfoNCE loss with temperature parameter τ\tau: LPCL=−1C∑k=1Clog⁡exp⁡(B(k)⋅v(k)/τ)∑q=1Cexp⁡(B(k)⋅v(q)/τ)\mathcal{L}_{PCL} = -\frac{1}{C} \sum_{k=1}^C \log \frac{\exp\left(B^{(k)} \cdot v^{(k)} / \tau\right)}{\sum_{q=1}^C \exp\left(B^{(k)} \cdot v^{(q)} / \tau\right)} This objective pulls the learned prototype B(k)B^{(k)} closer to image features of instrument class kk while pushing it away from image features of all other instrument classes q≠kq \neq k.

  4. Knowl 4 — Efficient Tuning Protocol and Objective of SurgicalSAM

    experimental setup

    SurgicalSAM performs parameter-efficient end-to-end tuning on pre-trained SAM weights (ViT-H backbone):

    • Frozen Parameters: The large Image Encoder EIE_I and the positive/negative prompt token embeddings λ+,λ−\lambda^+, \lambda^- are fixed.
    • Tunable Parameters: Only the Prototype-based Class Prompt Encoder (ECPE_{CP}) and the Mask Decoder (DMD_M) are trained, totaling 4.65M tunable parameters.
    • Total Loss Function: Supervised jointly by Dice loss and Prototype Contrastive loss: L=LDICE+LPCL\mathcal{L} = \mathcal{L}_{DICE} + \mathcal{L}_{PCL} where LDICE=1−2∑imigi∑imi2+∑igi2\mathcal{L}_{DICE} = 1 - \frac{2 \sum_i m_i g_i}{\sum_i m_i^2 + \sum_i g_i^2}, with mim_i representing predicted mask logits and gig_i representing binary ground-truth pixels.
    • Hyperparameters: Temperature τ=0.07\tau = 0.07; MLP intermediate hidden dimensions rD=128,rS=128r_D = 128, r_S = 128; number of sparse tokens per class n=2n = 2 for EndoVis2018 and n=4n = 4 for EndoVis2017; Adam optimizer with batch size 32; learning rate 1×10−31 \times 10^{-3} (EndoVis2018) and 1×10−41 \times 10^{-4} (EndoVis2017).
    • Acceleration: Pre-computed image embeddings are cached during training to eliminate repeated forward passes through the heavy ViT-H encoder.
  5. Knowl 5 — Surgical Instrument Segmentation on the EndoVis2018 Benchmark

    data/table

    Performance of SurgicalSAM compared to specialist architectures and SAM-based zero-shot pipelines on the EndoVis2018 dataset across 7 instrument categories: Bipolar Forceps (BF), Prograsp Forceps (PF), Large Needle Driver (LND), Suction Instrument (SI), Clip Applier (CA), Monopolar Curved Scissors (MCS), and Ultrasound Probe (UP).

    Method Challenge IoU IoU mc IoU BF PF LND SI CA MCS UP #Params
    TernausNet 46.22 39.87 14.19 44.20 4.67 0.00 0.00 0.00 50.44 0.00 32.20M
    MF-TAPNet 67.87 39.14 24.68 69.23 6.10 11.68 14.00 0.91 70.24 0.57 37.73M
    Dual-MF 70.40 - 35.09 74.10 6.80 46.00 30.10 7.60 80.90 0.10 203.80M
    ISINet 73.03 70.94 40.21 73.83 48.61 30.98 37.68 0.00 88.16 2.16 162.52M
    TraSeTr 76.20 - 47.71 76.30 53.30 46.50 40.60 13.90 86.20 17.15 -
    S3Net 75.81 74.02 42.58 77.22 50.87 19.83 50.59 0.00 92.12 7.44 68.41M
    MATIS Frame 82.37 77.01 48.65 83.35 38.82 40.19 64.49 4.32 93.18 16.17 68.72M
    MT-RCNN + SAM 78.49 78.49 56.07 79.83 74.86 43.12 62.88 16.74 91.62 23.45 57.67M
    Mask2Former + SAM 78.72 78.72 52.50 85.95 82.31 44.08 0.00 49.80 92.17 13.18 68.72M
    TrackAnything (1 Point) 40.36 38.38 20.62 30.20 12.87 24.46 9.17 0.19 55.03 12.41 -
    TrackAnything (5 Points) 65.72 60.88 38.60 72.90 31.07 64.73 10.24 12.28 61.05 17.93 -
    PerSAM 49.21 49.21 34.55 51.26 34.40 46.75 16.45 15.07 52.28 25.62 -
    PerSAM (Fine-Tune) 52.21 52.21 37.24 57.19 36.13 53.86 14.34 25.94 54.66 18.57 2
    SurgicalSAM (Ours) 80.33 80.33 58.87 83.66 65.63 58.75 54.48 39.78 88.56 21.23 4.65M
    GT Centroid + SAM 60.26 60.26 63.34 44.35 65.92 30.99 87.14 69.69 80.04 65.26 -
    GT Bbox + SAM 88.04 88.04 84.23 87.10 86.81 72.23 91.21 75.91 93.08 83.24 -

    SurgicalSAM attains an IoU of 80.33% and a mean class IoU (mc IoU) of 58.87%, outperforming all zero-shot SAM approaches and surpassing full-parameter specialist models on mc IoU while using only 4.65M tunable parameters.

  6. Knowl 6 — Surgical Instrument Segmentation on the EndoVis2017 Benchmark

    data/table

    Performance of SurgicalSAM evaluated under 4-fold cross-validation on the EndoVis2017 dataset across 7 instrument categories: Bipolar Forceps (BF), Prograsp Forceps (PF), Large Needle Driver (LND), Vessel Sealer (VS), Grasping Retractor (GR), Monopolar Curved Scissors (MCS), and Ultrasound Probe (UP).

    Method Challenge IoU IoU mc IoU BF PF LND VS GR MCS UP
    TernausNet 35.27 12.67 10.17 13.45 12.39 20.51 5.97 1.08 1.00 16.76
    MF-TAPNet 37.25 13.49 10.77 16.39 14.11 19.01 8.11 0.31 4.09 13.40
    Dual-MF 45.80 - 26.40 34.40 21.50 64.30 24.10 0.80 17.90 21.80
    ISINet 55.62 52.20 28.96 38.70 38.50 50.09 27.43 2.10 28.72 12.56
    TraSeTr 60.40 - 32.56 45.20 56.70 55.80 38.90 11.40 31.30 18.20
    S3Net 72.54 71.99 46.55 75.08 54.32 61.84 35.50 27.47 43.23 28.38
    MATIS Frame 68.79 62.74 37.30 66.18 50.99 52.23 32.84 15.71 19.27 23.90
    Mask2Former + SAM 66.21 66.21 55.26 66.84 55.36 83.29 73.52 26.24 36.26 45.34
    TrackAnything (1 Point) 54.90 52.46 55.35 47.59 28.71 43.27 82.75 63.10 66.46 55.54
    TrackAnything (5 Points) 67.41 64.50 62.97 55.42 44.46 62.43 83.68 62.59 67.03 65.17
    PerSAM 42.47 42.47 41.80 53.99 25.89 50.17 52.87 24.24 47.33 38.16
    PerSAM (Fine-Tune) 41.90 41.90 39.78 46.21 28.22 53.12 57.98 12.76 41.19 38.99
    SurgicalSAM (Ours) 69.94 69.94 67.03 68.30 51.77 75.52 68.24 57.63 86.95 60.80
    GT Centroid + SAM 44.42 44.42 54.41 63.42 36.03 22.57 54.21 75.18 70.17 59.25
    GT Bbox + SAM 76.31 76.31 81.18 89.36 73.44 67.67 90.04 87.79 94.03 65.91

    SurgicalSAM achieves 69.94% Challenge IoU and 67.03% mc IoU on EndoVis2017, demonstrating consistent gains over both zero-shot pipelines and specialist segmentation architectures across the 7 categories.

  7. Knowl 7 — Ablation on Contrastive Prototype Learning and Token Count

    data/table

    Ablation study on EndoVis2018 analyzing the effect of enabling the prototype contrastive loss LPCL\mathcal{L}_{PCL} versus using fixed class prototypes (computed by averaging class embeddings across the training set without contrastive tuning), evaluated over different numbers of sparse prompt tokens per class (n∈{2,4,6,8}n \in \{2, 4, 6, 8\}).

    Without LPCL\mathcal{L}_{PCL} With LPCL\mathcal{L}_{PCL}
    nn Challenge IoU mc IoU Challenge IoU mc IoU
    2 76.38 53.95 80.33 58.87
    4 78.26 56.54 79.46 58.40
    6 77.28 53.71 79.67 56.97
    8 76.98 53.94 80.10 58.30

    Adding LPCL\mathcal{L}_{PCL} yields consistent improvements (up to +4.92% mc IoU at n=2n=2), confirming that contrastive optimization prevents prototype representations of fine-grained categories from collapsing. Additionally, performance remains stable across varying token counts nn (80.33%80.33\% to 79.46%79.46\% Challenge IoU), demonstrating robustness to the prompt token hyperparameter.

  8. Knowl 8 — Cross-Dataset Generalization Between EndoVis2018 and EndoVis2017

    data/table

    Cross-dataset generalization evaluated by training on one dataset and evaluating on the other across the four common instrument categories: Bipolar Forceps (BF), Prograsp Forceps (PF), Large Needle Driver (LND), and Monopolar Curved Scissors (MCS).

    Training Validation Method BF PF LND MCS Mean IoU
    EndoVis2018 EndoVis2017 MATIS Frame 45.57 32.62 44.98 58.84 45.50
    EndoVis2018 EndoVis2017 SurgicalSAM 70.95 35.21 45.46 76.08 56.93
    EndoVis2017 EndoVis2018 MATIS Frame 65.55 13.89 38.25 65.58 45.81
    EndoVis2017 EndoVis2018 SurgicalSAM 44.50 27.17 50.76 62.94 46.34

    When trained on EndoVis2018 and tested on EndoVis2017, SurgicalSAM achieves a Mean IoU of 56.93%, outperforming the specialist model MATIS Frame (45.50%) by 11.43%. When trained on EndoVis2017 and tested on EndoVis2018, SurgicalSAM achieves 46.34% vs. 45.81% for MATIS Frame.

  9. Knowl 9 — Computational Efficiency and Throughput Analysis

    data/table

    Training speed (frames per second, fps), training GPU memory consumption (GB) across batch sizes (bz∈{2,16,32}bz \in \{2, 16, 32\}), and inference speed (fps) tested on an Nvidia Tesla V100 16GB GPU.

    Training Speed (fps) Training Memory (GB)
    Method bz=2bz=2 bz=16bz=16 bz=32bz=32 bz=2bz=2 bz=16bz=16 bz=32bz=32
    MATIS Frame 3.1 - - 13.1 - -
    MT-RCNN + SAM 8.2 12.8 - 3.2 13.9 -
    SurgicalSAM 40.1 57.4 59.8 1.9 5.9 9.6
    Inference Speed (fps)
    Method Online Feature Offline Feature
    MT-RCNN + SAM 1.6 14.3
    SurgicalSAM 1.7 91.7

    At bz=2bz=2, SurgicalSAM trains at 40.1 fps using 1.9 GB GPU memory, compared to 3.1 fps and 13.1 GB for MATIS Frame (over 10×10\times faster with less than 1/61/6 the memory). With pre-computed offline image features, SurgicalSAM achieves an inference speed of 91.7 fps compared to 14.3 fps for the multi-stage MT-RCNN + SAM pipeline.

Coverage note — None was omitted; all key architectural components, objectives, tuning setups, experiments, ablations, and complexity metrics are fully captured.

References

  1. 1.Allan, M.; Kondo, S.; Bodenstedt, S.; Leger, S.; Kadkhodamohammadi, R.; Luengo, I.; Fuentes, F.; Flouty, E.; Mohammed, A.; Pedersen, M.; Kori, A.; Alex, V.; Krishnamurthi, G.; Rauber, D.; Mendel, R.; Palm, C.; Bano, S.; Saibro, G.; Shih, C.-S.; Chiang, H.-A.; Zhuang, J.; Yang, J.; Iglovikov, V.; Dobrenkii, A.; Reddiboina, M.; Reddy, A.; Liu, X.; Gao, C.; Unberath, M.; Kim, M.; Kim, C.; Kim, C.; Kim, H.; Lee, G.; Ullah, I.; Luna, M.; Park, S. H.; Azizian, M.; Stoyanov, D.; Maier-Hein, L.; and Speidel, S. 2020. 2018 Robotic Scene Segmentation Challenge. arXiv:2001.11190.
  2. 2.Allan, M.; Shvets, A.; Kurmann, T.; Zhang, Z.; Duggal, R.; Su, Y.-H.; Rieke, N.; Laina, I.; Kalavakonda, N.; Bodenstedt, S.; Herrera, L.; Li, W.; Iglovikov, V.; Luo, H.; Yang, J.; Stoyanov, D.; Maier-Hein, L.; Speidel, S.; and Azizian, M. 2019. 2017 Robotic Instrument Segmentation Challenge. arXiv:1902.06426.
  3. 3.Ayobi, N.; Perez-Rondon, A.; Rodrıguez, S.; and Arbelaez, P. 2023. MATIS: Masked-Attention Transformers for Surgical Instrument Segmentation. In ISBI, 1–5.
  4. 4.Baby, B.; Thapar, D.; Chasmai, M.; Banerjee, T.; Dargan, K.; Suri, A.; Banerjee, S.; and Arora, C. 2023. From Forks to Forceps: A New Framework for Instance Segmentation of Surgical Instruments. In WACV, 6180–6190. IEEE.
  5. 5.Chen, T.; Zhu, L.; Deng, C.; Cao, R.; Wang, Y.; Zhang, S.; Li, Z.; Sun, L.; Zang, Y.; and Mao, P. 2023. SAM-Adapter: Adapting Segment Anything in Underperformed Scenes. In ICCV Workshops, 3367–3375.
  6. 6.Cheng, B.; Misra, I.; Schwing, A. G.; Kirillov, A.; and Girdhar, R. 2022. Masked-attention Mask Transformer for Universal Image Segmentation. In CVPR, 1290–1299.
  7. 7.Cheng, D.; Qin, Z.; Jiang, Z.; Zhang, S.; Lao, Q.; and Li, K. 2023. SAM on Medical Images: A Comprehensive Study on Three Prompt Modes. arXiv:2305.00035.
  8. 8.Deng, R.; Cui, C.; Liu, Q.; Yao, T.; Remedios, L. W.; Bao, S.; Landman, B. A.; Tang, Y.; Wheless, L. E.; Coburn, L. A.; Wilson, K. T.; Wang, Y.; Fogo, A. B.; Yang, H.; and Huo, Y. 2023. Segment Anything Model (SAM) for Digital Pathology: Assess Zero-shot Segmentation on Whole Slide Imaging. In Medical Imaging with Deep Learning, short paper track.
  9. 9.Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; et al. 2020. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. In ICLR.
  10. 10.Gonzalez, C.; Bravo-Sanchez, L.; and Arbelaez, P. 2020. ISINet: An Instance-Based Approach for Surgical Instrument Segmentation. In MICCAI, 595–605. Springer.
  11. 11.He, K.; Gkioxari, G.; Dollar, P.; and Girshick, R. 2017. Mask R-CNN. In ICCV, 2961–2969.
  12. 12.He, S.; Bao, R.; Li, J.; Stout, J.; Bjornerud, A.; Grant, P. E.; and Ou, Y. 2023. Computer-Vision Benchmark SegmentAnything Model (SAM) in Medical Images: Accuracy in 12 Datasets. arXiv:2304.09324.
  13. 13.Huang, Y.; Yang, X.; Liu, L.; Zhou, H.; Chang, A.; Zhou, X.; Chen, R.; Yu, J.; Chen, J.; Chen, C.; et al. 2023. Segment Anything Model for Medical Images? Medical Image Analysis, 103061.
  14. 14.Jian, Z.; Yue, W.; Wu, Q.; Li, W.; Wang, Z.; and Lam, V. 2020. Multitask Learning for Video-based Surgical Skill Assessment. In DICTA, 1–8.
  15. 15.Jin, Y.; Cheng, K.; Dou, Q.; and Heng, P.-A. 2019. Incorporating Temporal Prior from Motion Flow for Instrument Segmentation in Minimally Invasive Surgery Video. In MICCAI, 440–448. Springer.
  16. 16.Jin, Y.; Long, Y.; Chen, C.; Zhao, Z.; Dou, Q.; and Heng, P.A. 2021. Temporal Memory Relation Network for Workflow Recognition From Surgical Video. IEEE Transactions on Medical Imaging, 40(7): 1911–1923.
  17. 17.Kirillov, A.; Mintun, E.; Ravi, N.; Mao, H.; Rolland, C.; Gustafson, L.; Xiao, T.; Whitehead, S.; Berg, A. C.; Lo, W.Y.; Dollar, P.; and Girshick, R. 2023. Segment Anything. In ICCV, 4015–4026.
  18. 18.Li, Y.; Zhang, J.; Teng, X.; and Lan, L. 2023. RefSAM: Efficiently Adapting Segmenting Anything Model for Referring Video Object Segmentation. arXiv:2307.00997.
  19. 19.Liu, D.; Li, Q.; Jiang, T.; Wang, Y.; Miao, R.; Shan, F.; and Li, Z. 2021. Towards Unified Surgical Skill Assessment. In CVPR, 9522–9531.
  20. 20.Ma, J.; He, Y.; Li, F.; Han, L.; You, C.; and Wang, B. 2023. Segment Anything in Medical Images. arXiv:2304.12306.
  21. 21.Mazurowski, M. A.; Dong, H.; Gu, H.; Yang, J.; Konz, N.; and Zhang, Y. 2023. Segment Anything Model for Medical Image Analysis: An Experimental Study. Medical Image Analysis, 102918.
  22. 22.Milletari, F.; Navab, N.; and Ahmadi, S.-A. 2016. V-Net: Fully Convolutional Neural Networks for Volumetric Medical Image Segmentation. In 3DV, 565–571. IEEE.
  23. 23.Ni, Z.-L.; Bian, G.-B.; Wang, G.-A.; Zhou, X.-H.; Hou, Z.G.; Chen, H.-B.; and Xie, X.-L. 2020. Pyramid Attention Aggregation Network for Semantic Segmentation of Surgical Instruments. In AAAI, volume 34, 11782–11790.
  24. 24.Poole, B.; Ozair, S.; Van Den Oord, A.; Alemi, A.; and Tucker, G. 2019. On Variational Bounds of Mutual Information. In ICML, 5171–5180. PMLR.
  25. 25.Shademan, A.; Decker, R. S.; Opfermann, J. D.; Leonard, S.; Krieger, A.; and Kim, P. C. 2016. Supervised Autonomous Robotic Soft Tissue Surgery. Science Translational Medicine, 8(337): 337ra64–337ra64.
  26. 26.Shvets, A. A.; Rakhlin, A.; Kalinin, A. A.; and Iglovikov, V. I. 2018. Automatic Instrument Segmentation in RobotAssisted Surgery Using Deep Learning. In ICMLA, 624– 628. IEEE.
  27. 27.van den Oord, A.; Li, Y.; and Vinyals, O. 2019. Representation Learning with Contrastive Predictive Coding. arXiv:1807.03748.
  28. 28.Wald, T.; Roy, S.; Koehler, G.; Disch, N.; Rokuss, M. R.; Holzschuh, J.; Zimmerer, D.; and Maier-Hein, K. 2023. SAM. MD: Zero-Shot Medical Image Segmentation Capabilities of the Segment Anything Model. In Medical Imaging with Deep Learning, short paper track.
  29. 29.Wang, A.; Islam, M.; Xu, M.; Zhang, Y.; and Ren, H. 2023a. SAM Meets Robotic Surgery: An Empirical Study in Robustness Perspective. arXiv:2304.14674.
  30. 30.Wang, A.; Islam, M.; Xu, M.; Zhang, Y.; and Ren, H. 2023b. SAM Meets Robotic Surgery: An Empirical Study on Generalization, Robustness and Adaptation. In MICCAI Workshops.
  31. 31.Wang, D.; Zhang, J.; Du, B.; Xu, M.; Liu, L.; Tao, D.; and Zhang, L. 2023c. SAMRS: Scaling-up Remote Sensing Segmentation Dataset with Segment Anything Model. In NeurIPS Datasets and Benchmarks Track.
  32. 32.Wu, J.; Zhang, Y.; Fu, R.; Fang, H.; Liu, Y.; Wang, Z.; Xu, Y.; and Jin, Y. 2023. Medical SAM Adapter: Adapting Segment Anything Model for Medical Image Segmentation. arXiv:2304.12620.
  33. 33.Yan, Z.; Li, J.; Li, X.; Zhou, R.; Zhang, W.; Feng, Y.; Diao, W.; Fu, K.; and Sun, X. 2023. RingMo-SAM: A Foundation Model for Segment Anything in Multimodal RemoteSensing Images. IEEE Transactions on Geoscience and Remote Sensing, 61: 1–16.
  34. 34.Yang, J.; Gao, M.; Li, Z.; Gao, S.; Wang, F.; and Zheng, F. 2023. Track Anything: Segment Anything Meets Videos. arXiv:2304.11968.
  35. 35.Yang, L.; Fan, Y.; and Xu, N. 2019. Video Instance Segmentation. In ICCV, 5188–5197.
  36. 36.Yue, W.; Liao, H.; Xia, Y.; Lam, V.; Luo, J.; and Wang, Z. 2023. Cascade Multi-Level Transformer Network for Surgical Workflow Analysis. IEEE Transactions on Medical Imaging.
  37. 37.Zhang, J.; and Tao, D. 2020. Empowering Things with Intelligence: A Survey of the Progress, Challenges, and Opportunities in Artificial Intelligence of Things. IEEE Internet of Things Journal, 8(10): 7789–7817.
  38. 38.Zhang, K.; and Liu, D. 2023. Customized Segment Anything Model for Medical Image Segmentation. arXiv:2304.13785.
  39. 39.Zhang, R.; Jiang, Z.; Guo, Z.; Yan, S.; Pan, J.; Ma, X.; Dong, H.; Gao, P.; and Li, H. 2023. Personalize Segment Anything Model with One Shot. arXiv:2305.03048.
  40. 40.Zhao, Z.; Jin, Y.; Gao, X.; Dou, Q.; and Heng, P.-A. 2020. Learning Motion Flows for Semi-supervised Instrument Segmentation from Robotic Surgical Video. In MICCAI, 679–689. Springer.
  41. 41.Zhao, Z.; Jin, Y.; and Heng, P.-A. 2022. TraSeTR: Track-toSegment Transformer with Contrastive Query for Instancelevel Instrument Segmentation in Robotic Surgery. In ICRA, 11186–11193. IEEE.

Citation

MLA
Yue, W., et al. “SurgicalSAM: Efficient Class Promptable Surgical Instrument Segmentation”. arXiv, 2023, http://arxiv.org/abs/2308.08746v2.
APA
Yue, W., Zhang, J., Hu, K., Xia, Y., Luo, J., & Wang, Z. (2023). SurgicalSAM: Efficient Class Promptable Surgical Instrument Segmentation. arXiv. http://arxiv.org/abs/2308.08746v2
Chicago
Yue, W., J. Zhang, K. Hu, Y. Xia, J. Luo, and Z. Wang. 2023. “SurgicalSAM: Efficient Class Promptable Surgical Instrument Segmentation”. arXiv. http://arxiv.org/abs/2308.08746v2.
Harvard
Yue, W. et al. (2023) “SurgicalSAM: Efficient Class Promptable Surgical Instrument Segmentation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2308.08746v2.
Vancouver
1. Yue W, Zhang J, Hu K, Xia Y, Luo J, Wang Z (2023) SurgicalSAM: Efficient Class Promptable Surgical Instrument Segmentation. arXiv

BibTeX

@article{yue2023surgicalsam,
  title = {SurgicalSAM: Efficient Class Promptable Surgical Instrument Segmentation},
  author = {Yue, Wenxi and Zhang, Jing and Hu, Kun and Xia, Yong and Luo, Jiebo and Wang, Zhiyong},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2308.08746v2},
  eprint = {2308.08746}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF