SurgicalSAM: Efficient Class Promptable Surgical Instrument Segmentation
Wenxi YueJing ZhangKun HuYong XiaJiebo LuoZhiyong Wang
Proposes an end-to-end framework that efficiently adapts the Segment Anything Model for surgical instrument segmentation by replacing fragile manual bounding-box prompts with learned class prototypes and contrastive learning.
Accurate identification and segmentation of surgical instruments in video feeds is essential for advancing computer-assisted surgery and improving operating room safety. While large artificial intelligence foundation models like the Segment Anything Model offer strong general segmentation capabilities, directly applying them to surgery has proven difficult. Surgical instruments differ substantially from everyday objects, share high visual similarities across categories, and standard foundation models rely heavily on precise manual points or bounding boxes that are impractical during live clinical workflows.
To overcome these limitations, the article evaluates and demonstrates SurgicalSAM, a framework designed to efficiently adapt foundation segmentation models to surgical environments without requiring manual spatial prompts. The objective is to achieve state-of-the-art segmentation accuracy across instrument categories while keeping computational demands low and streamlining deployment into a single, automated step.
The authors developed an end-to-end tuning method that uses category labels as prompts instead of manual bounding boxes. By keeping the massive image-processing backbone frozen and only training a lightweight prompt encoder and mask decoder, the model learns category-specific visual prototypes. A contrastive learning technique ensures these prototypes clearly distinguish between visually similar tools. The system was validated against existing specialist architectures and foundation model baselines across benchmark surgical video datasets, specifically EndoVis2017 and EndoVis2018.
The evaluation revealed several key findings. First, SurgicalSAM matched or outperformed leading specialist models across both datasets while tuning only 4.65 million parameters, compared to nearly 69 million in competing state-of-the-art models. Second, it demonstrated superior generalisation when trained on one dataset and tested on another, achieving an 11.43 percentage-point gain in cross-dataset accuracy over top specialist architectures. Third, training efficiency increased dramatically: training was more than 10 times faster than comparable specialized models, utilizing less than one-sixth of the graphic processor memory. Finally, eliminating the need for precise spatial coordinate prompts made the system significantly more robust against noise compared to zero-shot foundation models.
These results demonstrate that adapting foundation models using lightweight, class-based tuning is a highly viable path for medical artificial intelligence. The approach drastically reduces development and compute costs while eliminating the latency and failure points associated with multi-stage detection systems. Enhanced cross-dataset reliability suggests these models can better adapt to variations across hospital setups and surgical tools.
Engineering and clinical teams should consider shifting from training large custom models from scratch toward adapting general foundation models using class-prototype strategies. Before broad clinical integration, next steps should include testing on larger, multicenter clinical datasets with greater procedural variety to establish performance across diverse operating conditions.
The findings are supported by strong benchmark results, but confidence should be framed around the dataset scale, as the benchmarks encompass limited video quantities and fixed instrument categories. Continued validation in real-time robotic surgery environments will be necessary before operational deployment.
- Paper: Segment Anything, Alexander M. Kirillov et al. (2023). Introduces the Segment Anything Model (SAM) foundation architecture and promptable segmentation formulation that SurgicalSAM directly adapts for surgical instrument segmentation.
- Paper: Segment anything in medical images, Jun Ma et al. (2023). Demonstrates the foundational adaptation of SAM to medical imaging via bounding box prompts and decoder tuning, establishing the domain-specific fine-tuning paradigm that SurgicalSAM builds upon.
- Paper: PANet: Few-Shot Image Semantic Segmentation With Prototype Alignment, Kaixin Wang et al. (2019). Establishes prototype-based metric learning and class alignment representations for segmentation, which directly inform SurgicalSAM's category-level visual prototype formulation.
- Paper: Unleashing the Potential of SAM for Medical Adaptation via Hierarchical Decoding, Zhiheng Cheng et al. (2024). Extends efficient, prompt-free adaptation of SAM in medical contexts through hierarchical decoding and mask-guided attention mechanisms.
- Paper: SAM 2: Segment Anything in Images and Videos, Nikhila Ravi et al. (2025). Generalizes the Segment Anything framework from static frames to continuous video tracking with streaming memory transformers, offering a broader video foundation for surgical applications.
- Paper: SAM 3: Segment Anything with Concepts, Nicolas Carion et al. (2025). Advances promptable foundation segmentation by enabling concept-level noun phrase and exemplar-driven instance tracking across video streams.
