Built independently by an author, for readers. Read the story and support ChapterPal

keyword

promptable segmentation

Promptable segmentation is a computer vision task and framework in which an artificial intelligence model generates precise segmentation masks for target regions based on user-provided guiding cues, known as prompts. Unlike conventional segmentation systems that are restricted to fixed, predefined categories of objects, promptable segmentation allows users to interactively isolate arbitrary objects, parts, or regions using diverse prompt modalities such as coordinate points, bounding boxes, coarse scribbles, or natural language descriptions. This paradigm enables flexible zero-shot generalization across novel domains and unseen visual concepts without requiring task-specific retraining. Additionally, promptable segmentation architectures are designed to handle multi-granularity ambiguity, allowing models to dynamically adapt their output from whole objects to sub-parts across two-dimensional images, video sequences, and three-dimensional spatial representations.

4 items

Segment Any 3D Gaussians

Segment Any 3D Gaussians

Jiazhong Cen, Jiemin Fang, Chen Yang, Lingxi Xie, Xiaopeng Zhang, Wei Shen, Qi Tian

OrganizationsAI InstituteHuaweiShanghai Jiao Tong University

Why you should read this

Proposes an efficient framework that integrates scale-gated affinity features into 3D Gaussian Splatting, enabling real-time, multi-granularity 3D object segmentation from 2D visual prompts in milliseconds by distilling knowledge from the Segment Anything Model.

This paper presents SAGA (Segment Any 3D GAussians), a highly efficient 3D promptable segmentation method based on 3D Gaussian Splatting (3D-GS). Given 2D visual prompts as input, SAGA can segment the corresponding 3D target represented by 3D Gaussians within 4 ms. This is achieved by attaching a scale-gated affinity feature to each 3D Gaussian to endow it a new property towards multi-granularity segmentation. Specifically, a scale-aware contrastive training strategy is proposed for the scale-gated affinity feature learning. It 1) distills the segmentation capability of the Segment Anything Model (SAM) from 2D masks into the affinity features and 2) employs a soft scale-gate mechanism to deal with multi-granularity ambiguity in 3D segmentation through adjusting the magnitude of each feature channel according to a specified 3D physical scale. Evaluations demonstrate that SAGA achieves real-time multi-granularity segmentation with quality comparable to state-of-the-art methods. As one of the first methods addressing promptable segmentation in 3D-GS, the simplicity and effectiveness of SAGA pave the way for future advancements in this field.

Added

2026-10-06

SAM-6D: Segment Anything Model Meets Zero-Shot 6D Object Pose Estimation

SAM-6D: Segment Anything Model Meets Zero-Shot 6D Object Pose Estimation

Jiehong Lin, Lihua Liu, Dekun Lu, Kui Jia

OrganizationsDexForce Technology Co., Ltd.South China University of TechnologyThe Chinese University of Hong Kong

Why you should read this

Presents SAM-6D, a framework that couples the zero-shot capabilities of the Segment Anything Model with a two-stage 3D point-matching network using background tokens to detect and estimate 6D poses of unseen objects in cluttered RGB-D scenes.

Zero-shot 6D object pose estimation involves the detection of novel objects with their 6D poses in cluttered scenes, presenting significant challenges for model generalizability. Fortunately, the recent Segment Anything Model (SAM) has showcased remarkable zero-shot transfer performance, which provides a promising solution to tackle this task. Motivated by this, we introduce SAM-6D, a novel framework designed to realize the task through two steps, including instance segmentation and pose estimation. Given the target objects, SAM-6D employs two dedicated sub-networks, namely Instance Segmentation Model (ISM) and Pose Estimation Model (PEM), to perform these steps on cluttered RGB-D images. ISM takes SAM as an advanced starting point to generate all possible object proposals and selectively preserves valid ones through meticulously crafted object matching scores in terms of semantics, appearance and geometry. By treating pose estimation as a partial-to-partial point matching problem, PEM performs a two-stage point matching process featuring a novel design of background tokens to construct dense 3D-3D correspondence,

Added

2026-09-26

Hierarchical Open-vocabulary Universal Image Segmentation

Hierarchical Open-vocabulary Universal Image Segmentation

Xudong Wang, Shufan Li, Konstantinos Kallidromitis, Yusuke Kato, Kazuki Kozuka, Trevor Darrell

OrganizationsPanasonicUniversity of California Berkeley

Why you should read this

Presents HIPIE, a unified open-vocabulary framework that resolves segmentation ambiguity across multiple granularities by incorporating hierarchical visual representations alongside decoupled text-image fusion mechanisms for stuff and thing categories.

Open-vocabulary image segmentation aims to partition an image into semantic regions according to arbitrary text descriptions. However, complex visual scenes can be naturally decomposed into simpler parts and abstracted at multiple levels of granularity, introducing inherent segmentation ambiguity. Unlike existing methods that typically sidestep this ambiguity and treat it as an external factor, our approach actively incorporates a hierarchical representation encompassing different semantic-levels into the learning process. We also propose a decoupled text-image fusion mechanism and representation learning modules for both “things” and “stuff”.1 Additionally, we systematically examine the differences that exist in the textual and visual features between these types of categories. Our resulting model, named HIPIE, tackles HIerarchical, oPen-vocabulary, and unIvErsal segmentation tasks within a unified framework. Benchmarked on over 40 datasets, e.g., ADE20K, COCO, Pascal-VOC Part, RefCOCO/RefCOCOg, ODinW and SeginW, HIPIE achieves the state-of-the-art results at various levels of image comprehension, including semantic-level (e.g., semantic segmentation), instance-level (e.g., panoptic/referring segmentation and object detection), as well as part-level (e.g., part/subpart segmentation) tasks.

Added

2026-09-26

Segment anything in medical images

Segment anything in medical images

Jun Ma, Yuting He, Feifei Li, Li-Jun Han, Chenyu You, Bo Wang

OrganizationsNew York UniversityUniversity Health NetworkUniversity of TorontoUniversity of Western OntarioVector InstituteYale University

Why you should read this

Presents MedSAM, a universal medical image segmentation foundation model trained on over 1.5 million image-mask pairs across 10 modalities that outperforms specialized models across dozens of clinical benchmarks.

Medical image segmentation is a critical component in clinical practice, facilitating accurate diagnosis, treatment planning, and disease monitoring. However, existing methods, often tailored to specific modalities or disease types, lack generalizability across the diverse spectrum of medical image segmentation tasks. Here we present MedSAM, a foundation model designed for bridging this gap by enabling universal medical image segmentation. The model is developed on a large-scale medical image dataset with 1,570,263 image-mask pairs, covering 10 imaging modalities and over 30 cancer types. We conduct a comprehensive evaluation on 86 internal validation tasks and 60 external validation tasks, demonstrating better accuracy and robustness than modality-wise specialist models. By delivering accurate and efficient segmentation across a wide spectrum of tasks, MedSAM holds significant potential to expedite the evolution of diagnostic tools and the personalization of treatment plans.

Added

2026-09-24