Pointly-Supervised Instance Segmentation
Bowen ChengOmkar ParkhiAlexander Kirillov
Proposes a point-based annotation scheme combining bounding boxes with just ten labeled random points per object, allowing standard instance segmentation models to match up to 98% of fully supervised performance while cutting annotation time by five times.
High-performance computer vision systems for instance segmentation—the task of detecting objects and outlining their exact pixel boundaries—typically rely on full mask annotations. Manually drawing these pixel-level polygon masks is a severe operational bottleneck, taking roughly 79 seconds per object in standard datasets. While weakly supervised approaches using only bounding boxes reduce labeling time, they exhibit a substantial performance drop of 15% or more compared to fully supervised baselines. Consequently, organizations face high data acquisition costs and lengthy timelines when deploying segmentation models to new domains or object classes.
The article demonstrates a point-based annotation scheme that pairs object bounding boxes with binary labels for uniformly sampled points inside each box. The objective is to evaluate whether this low-cost supervision can train standard, off-the-shelf segmentation models to match full mask supervision performance across diverse datasets while drastically reducing labeling time and pipeline complexity.
To assess this strategy, the authors evaluated standard architectures (such as Mask R-CNN, CondInst, and PointRend) across multiple prominent benchmark datasets, including COCO, PASCAL VOC, Cityscapes, and LVIS. Instead of requiring human annotators to select click locations, the scheme automatically samples random coordinates within each box and prompts annotators to perform a simple binary choice: object or background. This setup enabled rigorous simulation across existing ground truth datasets, supplemented by real-world human timing trials using a custom labeling interface. The authors also introduced Implicit PointRend, an architectural variant tailored for point-level supervision that dynamically generates instance-specific parameters to predict masks from coordinates and image features.
The findings show that training Mask R-CNN with only 10 annotated random points per instance (termed P10) achieves 94% to 98% of its fully supervised performance across all tested datasets. Human annotator trials revealed that classifying a single point takes only 0.9 seconds, meaning a complete object annotation (a 7-second bounding box plus 10 points) requires 16 seconds—making the process approximately 5 times faster than standard polygon mask collection. Furthermore, the annotation scheme proved highly robust; the model retained performance even with a 5% label error rate or across different random point selections. In downstream evaluations, pre-training models on point-annotated data achieved 100% of the transfer performance of models pre-trained on full masks, and the proposed Implicit PointRend model trained on 10 points matched the absolute accuracy of fully supervised Mask R-CNN.
These results establish that complete manual mask delineation is largely unnecessary for achieving production-grade instance segmentation. For engineering and data operations, adopting point-supervised labeling can cut annotation labor budgets and lead times by up to 80% with minimal risk to model performance. The workflow eliminates the need for specialized interactive models or multi-stage pseudo-mask generation, enabling immediate integration into standard training pipelines through simple prediction interpolation.
Organizations initiating new vision labeling campaigns should transition from manual polygon tracing to a hybrid box-plus-point workflow, targeting approximately 10 random points per instance. Teams working with high-capacity backbones or implicit models should also incorporate point subsampling data augmentation to prevent overfitting. Where top-tier performance is required, self-training pipelines can be deployed to close the remaining gap to full supervision.
The findings are supported by consistent results across diverse domains ranging from autonomous driving scenes to large-scale vocabularies exceeding 1,000 categories. However, practitioners should note that labeling noise increases slightly around subtle object boundaries, and models with extremely coarse intermediate feature resolutions require architectural adaptations (such as Implicit PointRend) to fully exploit point-level supervisory signals.
- Paper: Mask R-CNN, Kaiming He et al. (2017). Mask R-CNN introduces the standard two-stage instance segmentation architecture and RoIAlign mechanism that the source directly uses as its primary baseline and benchmark for point-based supervision.
- Paper: Per-Pixel Classification is Not All You Need for Semantic Segmentation, Bowen Cheng et al. (2021). MaskFormer establishes the unified mask classification paradigm that underpins modern query-based segmentation pipelines evaluated under point-level supervisory signals.
- Paper: YOLACT: Real-Time Instance Segmentation, Daniel Bolya et al. (2019). YOLACT demonstrates single-stage dynamic instance mask generation via linear combination of prototypes, providing architectural foundations for implicit and dynamic segmentation heads.
- Paper: Simple Copy-Paste is a Strong Data Augmentation Method for Instance Segmentation, Golnaz Ghiasi et al. (2020). Simple Copy-Paste provides foundational data augmentation strategies and low-data regime analysis for instance segmentation models.
- Paper: Simultaneous Detection and Segmentation, Bharath Hariharan et al. (2014). This paper establishes the simultaneous detection and instance segmentation formulation that defines the problem domain simplified by pointly-supervised training.
- Paper: Segment Anything, Alexander M. Kirillov et al. (2023). Segment Anything scales promptable segmentation using sparse inputs such as single points and bounding boxes to a foundational zero-shot segmentation paradigm.
- Paper: Mask Transfiner for High-Quality Instance Segmentation, Lei Ke et al. (2022). Mask Transfiner extends sparse point-based computation by selectively processing error-prone boundary regions using a multi-scale quadtree and transformer refinement.
- Paper: Masked-attention Mask Transformer for Universal Image Segmentation, Bowen Cheng et al. (2022). Mask2Former extends universal mask classification architectures with masked attention and point-sampled loss formulations across all standard segmentation tasks.
- Paper: Token Contrast for Weakly-Supervised Semantic Segmentation, Lixiang Ru et al. (2023). Token Contrast investigates weak supervision alternatives using vision transformers to address regional over-smoothing without dense pixel masks.
- Paper: Weakly Supervised Semantic Segmentation via Adversarial Learning of Classifier and Reconstructor, Hyeokjun Kweon et al. (2023). This work explores an alternative weakly supervised segmentation strategy using adversarial reconstruction to delineate accurate boundaries from minimal supervisory tags.
