Pointly-Supervised Instance Segmentation

Bowen ChengOmkar ParkhiAlexander Kirillov

article2022CVPR153 citations

Proposes a point-based annotation scheme combining bounding boxes with just ten labeled random points per object, allowing standard instance segmentation models to match up to 98% of fully supervised performance while cutting annotation time by five times.

Listen

High-performance computer vision systems for instance segmentation—the task of detecting objects and outlining their exact pixel boundaries—typically rely on full mask annotations. Manually drawing these pixel-level polygon masks is a severe operational bottleneck, taking roughly 79 seconds per object in standard datasets. While weakly supervised approaches using only bounding boxes reduce labeling time, they exhibit a substantial performance drop of 15% or more compared to fully supervised baselines. Consequently, organizations face high data acquisition costs and lengthy timelines when deploying segmentation models to new domains or object classes.

The article demonstrates a point-based annotation scheme that pairs object bounding boxes with binary labels for uniformly sampled points inside each box. The objective is to evaluate whether this low-cost supervision can train standard, off-the-shelf segmentation models to match full mask supervision performance across diverse datasets while drastically reducing labeling time and pipeline complexity.

To assess this strategy, the authors evaluated standard architectures (such as Mask R-CNN, CondInst, and PointRend) across multiple prominent benchmark datasets, including COCO, PASCAL VOC, Cityscapes, and LVIS. Instead of requiring human annotators to select click locations, the scheme automatically samples random coordinates within each box and prompts annotators to perform a simple binary choice: object or background. This setup enabled rigorous simulation across existing ground truth datasets, supplemented by real-world human timing trials using a custom labeling interface. The authors also introduced Implicit PointRend, an architectural variant tailored for point-level supervision that dynamically generates instance-specific parameters to predict masks from coordinates and image features.

The findings show that training Mask R-CNN with only 10 annotated random points per instance (termed P10) achieves 94% to 98% of its fully supervised performance across all tested datasets. Human annotator trials revealed that classifying a single point takes only 0.9 seconds, meaning a complete object annotation (a 7-second bounding box plus 10 points) requires 16 seconds—making the process approximately 5 times faster than standard polygon mask collection. Furthermore, the annotation scheme proved highly robust; the model retained performance even with a 5% label error rate or across different random point selections. In downstream evaluations, pre-training models on point-annotated data achieved 100% of the transfer performance of models pre-trained on full masks, and the proposed Implicit PointRend model trained on 10 points matched the absolute accuracy of fully supervised Mask R-CNN.

These results establish that complete manual mask delineation is largely unnecessary for achieving production-grade instance segmentation. For engineering and data operations, adopting point-supervised labeling can cut annotation labor budgets and lead times by up to 80% with minimal risk to model performance. The workflow eliminates the need for specialized interactive models or multi-stage pseudo-mask generation, enabling immediate integration into standard training pipelines through simple prediction interpolation.

Organizations initiating new vision labeling campaigns should transition from manual polygon tracing to a hybrid box-plus-point workflow, targeting approximately 10 random points per instance. Teams working with high-capacity backbones or implicit models should also incorporate point subsampling data augmentation to prevent overfitting. Where top-tier performance is required, self-training pipelines can be deployed to close the remaining gap to full supervision.

The findings are supported by consistent results across diverse domains ranging from autonomous driving scenes to large-scale vocabularies exceeding 1,000 categories. However, practitioners should note that labeling noise increases slightly around subtle object boundaries, and models with extremely coarse intermediate feature resolutions require architectural adaptations (such as Implicit PointRend) to fully exploit point-level supervisory signals.

arXiv: 2104.06404
  • Paper: Mask R-CNN, Kaiming He et al. (2017). Mask R-CNN introduces the standard two-stage instance segmentation architecture and RoIAlign mechanism that the source directly uses as its primary baseline and benchmark for point-based supervision.
  • Paper: Per-Pixel Classification is Not All You Need for Semantic Segmentation, Bowen Cheng et al. (2021). MaskFormer establishes the unified mask classification paradigm that underpins modern query-based segmentation pipelines evaluated under point-level supervisory signals.
  • Paper: YOLACT: Real-Time Instance Segmentation, Daniel Bolya et al. (2019). YOLACT demonstrates single-stage dynamic instance mask generation via linear combination of prototypes, providing architectural foundations for implicit and dynamic segmentation heads.
  • Paper: Simple Copy-Paste is a Strong Data Augmentation Method for Instance Segmentation, Golnaz Ghiasi et al. (2020). Simple Copy-Paste provides foundational data augmentation strategies and low-data regime analysis for instance segmentation models.
  • Paper: Simultaneous Detection and Segmentation, Bharath Hariharan et al. (2014). This paper establishes the simultaneous detection and instance segmentation formulation that defines the problem domain simplified by pointly-supervised training.
Cover for Pointly-Supervised Instance Segmentation

Abstract

We propose an embarrassingly simple point annotation scheme to collect weak supervision for instance segmentation. In addition to bounding boxes, we collect binary labels for a set of points uniformly sampled inside each bounding box. We show that the existing instance segmentation models developed for full mask supervision can be seamlessly trained with point-based supervision collected via our scheme. Remarkably, Mask R-CNN trained on COCO, PASCAL VOC, Cityscapes, and LVIS with only 10 annotated random points per object achieves 94%–98% of its fully-supervised performance, setting a strong baseline for weakly-supervised instance segmentation. The new point annotation scheme is approximately 5 times faster than annotating full object masks, making high-quality instance segmentation more accessible in practice.

Inspired by the point-based annotation form, we propose a modification to PointRend instance segmentation module. For each object, the new architecture, called Implicit PointRend, generates parameters for a function that makes the final point-level mask prediction. Implicit PointRend is more straightforward and uses a single point-level mask loss. Our experiments show that the new module is more suitable for the point-based supervision.

Table of Contents

  • 1. Introduction
  • 2. Related Works
  • 3. Pointly-Supervised Instance Segmentation
  • 3.1. Annotation format and collection
  • 3.2. Training with points
  • 3.3. Experiments with point supervision
  • 3.3.1 Ablation of the annotation design
  • 3.3.2 Main results
  • 3.3.3 Annotation time and performance trade-off.
  • 4. Implicit PointRend Model
  • 4.1. Point-wise representation and point head
  • 4.2. Point selection for inference and training
  • 4.3. Implementation details
  • 4.4. Experimental evaluation
  • 5. Conclusion
  • References

Knowls

  1. Knowl 1 — Point-Supervised Instance Segmentation Annotation Scheme

    model/method

    The PNP_N annotation scheme collects weak supervision for instance segmentation by combining bounding boxes with NN uniformly sampled points per instance. The collection proceeds in two stages:

    1. Bounding Box Collection: Object bounding boxes are obtained using standard methods, such as extreme clicking (which requires approximately 77 seconds per object).
    2. Point Labeling: Exactly NN random spatial locations are sampled uniformly within each bounding box. Annotators perform a sequential binary classification task, labeling each point as object (foreground) or background, given the box and category label.

    A human annotator takes approximately 0.90.9 seconds per point. For N=10N = 10 (P10P_{10}), the total annotation time per instance is: Time(P10)=7+10×0.9=16 seconds\text{Time}(P_{10}) = 7 + 10 \times 0.9 = 16\text{ seconds} This is roughly 5 times faster than polygon-based mask annotation in COCO, which takes an average of 79.279.2 seconds per instance.

    Because points are sampled uniformly at random rather than through annotator clicks, the scheme avoids spatial correlation biases of human clicking and can be simulated deterministically on datasets with existing full mask ground truth by querying pixel labels at sampled coordinates.

  2. Knowl 2 — Training Instance Segmentation Models via Point Bilinear Interpolation

    model/method

    Standard instance segmentation models predicting mask logits on a discrete regular grid (such as a 28×2828 \times 28 Region of Interest (RoI) grid in Mask R-CNN or an image-level grid in CondInst) can be directly trained with point-based supervision without architectural modifications.

    Given an object with ground-truth point coordinates (xi,yi)(x_i, y_i) and binary labels yi∈{0,1}y_i \in \{0, 1\} for i∈{1,…,N}i \in \{1, \dots, N\}, continuous mask predictions y^i\hat{y}_i at (xi,yi)(x_i, y_i) are computed via bilinear interpolation from the network's grid logits. The point-level mask loss is formulated as a binary cross-entropy loss over the annotated points: Lmask=−1N∑i=1N[yilog⁡y^i+(1−yi)log⁡(1−y^i)]\mathcal{L}_{\text{mask}} = -\frac{1}{N} \sum_{i=1}^N \Big[ y_i \log \hat{y}_i + (1 - y_i) \log(1 - \hat{y}_i) \Big] Gradients are backpropagated through the bilinear interpolation weights to the grid representation. For region-based detectors, ground-truth points falling outside predicted bounding boxes are ignored during training.

  3. Knowl 3 — Implicit PointRend Architecture

    model/method

    Implicit PointRend is an instance segmentation module that replaces the coarse 7×77 \times 7 mask head of standard PointRend with an instance-conditioned implicit representation, utilizing a single point-level loss.

    Architecture:

    • Parameter Head: Takes Region of Interest (RoI) features and dynamically predicts the parameters (weights and biases) of a point head MLP for each detected object instance.
    • Point Head: A Multi-Layer Perceptron (MLP) containing 3 hidden layers with 256 channels, ReLU activations, and a final sigmoid activation. The point head takes as input the concatenation of:
      1. Fine-grained image features extracted from the P2P_2 level of Feature Pyramid Networks (FPN) at the point's image coordinate via bilinear interpolation (256256 dimensions).
      2. Positional encoding of the point's coordinates (x,y)(x, y) relative to the bounding box center, encoded using random Fourier features.

    Training and Inference:

    • Loss Function: Trained with a single point-wise binary cross-entropy loss applied to the point head outputs, along with an L2L_2 regularization penalty (weight 10−510^{-5}) on the predicted dynamic MLP parameters to prevent unbounded growth.
    • Sampling: Uses simple uniform point sampling during mask training instead of uncertainty-based importance sampling.
    • Inference: Employs adaptive subdivision rendering, starting with a 28×2828 \times 28 uniform grid and iteratively upsampling by a factor of 2 up to 224×224224 \times 224 by selecting the N=282N = 28^2 most uncertain points at each step.
  4. Knowl 4 — Point Subsampling Data Augmentation

    model/method

    When training instance segmentation models with limited point supervision (e.g., P10P_{10}), higher-capacity backbones (such as ResNeXt-101) and implicit dynamic architectures can suffer from reduced training data variability and overfit.

    To mitigate this, point subsampling data augmentation randomly samples a subset equal to half of the available ground-truth points per instance at each training iteration (e.g., selecting 5 random points per object out of the 10 available in P10P_{10}). The mask loss is then computed exclusively over this subsampled point set. This simple iteration-level subsampling increases data variability and improves the performance of high-capacity models.

  5. Knowl 5 — Performance of Mask R-CNN with 10-Point Supervision Across Datasets

    data/table

    Mask R-CNN with a ResNet-50-FPN backbone trained with 10 randomly sampled points per instance (P10P_{10}) achieves 94%–97% of the performance of the same model trained with full polygon mask supervision (MM) across four distinct benchmarks without modifying the architecture or default hyperparameters:

    Dataset Metric Full Mask (MM) 10 Points (P10P_{10}) Retained Performance
    COCO val2017 AP 37.2 36.1 97%
    PASCAL VOC val AP50\text{AP}_{50} 66.3 64.2 97%
    Cityscapes val AP 32.7 30.7 94%
    LVISv1.0 val AP 22.8 21.5 94%

    These results demonstrate that point supervision generalizes across different dataset scales (from 20 categories in PASCAL VOC to over 1,000 categories in LVIS) and scene domains (common objects to street scenes).

  6. Knowl 6 — Implicit PointRend vs. PointRend and Mask R-CNN Under Point Supervision

    data/table

    Implicit PointRend overcomes the limitations of standard PointRend under point supervision on COCO val2017 using a ResNet-50-FPN backbone:

    Method Mask Supervision AP (MM) 10-Point Supervision AP (P10P_{10})
    PointRend 38.3 35.7
    Implicit PointRend 38.5 36.9

    While PointRend drops by 2.62.6 AP under P10P_{10} supervision due to low-resolution coarse head conflicts, Implicit PointRend gains +1.2+1.2 AP over PointRend under P10P_{10}.

    Furthermore, Implicit PointRend trained with 10 points matches or exceeds fully supervised Mask R-CNN across various backbone capacities on COCO:

    Method Supervision R50 AP R101 AP X101 AP
    Mask R-CNN MM 37.2 38.6 39.5
    Mask R-CNN P10P_{10} 36.0 37.8 38.5
    Implicit PointRend P10P_{10} 36.9 38.5 39.7
  7. Knowl 7 — Robustness of Point Supervision to Number of Points, Sampling Seed, and Label Noise

    empirical result

    Ablation experiments on COCO val2017 with Mask R-CNN (ResNet-50-FPN) evaluate point supervision sensitivity:

    1. Number of Points (NN): Mask AP increases sharply with point count up to 10 points (P1≈31.8P_1 \approx 31.8, P10=36.1P_{10} = 36.1) and shows diminishing returns beyond (P20=36.4P_{20} = 36.4, an improvement of only +0.3+0.3 AP despite doubling annotation time).
    2. Sampling Seed Variability: Training Mask R-CNN across 5 independent random uniform point samplings of P10P_{10} on COCO train2017 results in an AP variation of only 0.10.1 AP on val2017, showing minimal sensitivity to exact point locations.
    3. Label Noise Sensitivity:
      • Corrupting 5%5\% of randomly selected point labels in P10P_{10} reduces AP by 0.20.2 (from 36.136.1 to 35.935.9 AP).
      • Corrupting 5%5\% of point labels located specifically closest to object boundaries reduces AP by 0.40.4 (to 35.735.7 AP).

    This resilience indicates that rigorous manual verification steps can be relaxed during point data collection.

  8. Knowl 8 — Transfer Learning and Self-Training with Point Supervision

    empirical result

    Point-based supervision can be effectively combined with transfer learning and self-training pipelines:

    • Self-Training on COCO: Generating pseudo-masks from a P10P_{10}-supervised Mask R-CNN (36.1 AP) and retraining Mask R-CNN on these pseudo-masks achieves 36.736.7 AP (+0.6+0.6 AP gain), closing the gap to 98%98\% of a fully supervised Mask R-CNN trained with an equivalent 6×6\times schedule.
    • Downstream Fine-Tuning on PASCAL VOC:
      • Pre-training on COCO with full masks (COCO-MM) and fine-tuning on PASCAL VOC with P10P_{10} achieves 74.0% AP5074.0\% \text{ AP}_{50}, closing the gap with full mask fine-tuning (74.5% AP5074.5\% \text{ AP}_{50}).
      • Pre-training on COCO with 10 points (COCO-P10P_{10}) followed by full mask fine-tuning on PASCAL VOC achieves 74.5% AP5074.5\% \text{ AP}_{50}, performing identically to pre-training with full COCO masks (74.5% AP5074.5\% \text{ AP}_{50}).
  9. Knowl 9 — Annotation Time Efficiency of Point Supervision

    empirical result

    When evaluated under fixed annotation time budgets (worker days) on COCO val2017, Mask R-CNN trained with P10P_{10} point supervision significantly outperforms models trained with full mask supervision (Mask R-CNN) and bounding-box-only supervision (BoxInst):

    • To achieve a given Mask AP, pointly-supervised Mask R-CNN requires 1.7×1.7\times less total annotation time than full mask supervision.
    • At equal total annotation time budgets, Mask R-CNN trained with P10P_{10} achieves up to a +2.9+2.9 AP gain over full mask supervision trained on the subset of data annotatable within that same time budget.
  10. Knowl 10 — Point Supervision Across Alternative Instance Segmentation Architectures

    data/table

    Point supervision (P10P_{10}) evaluated across different instance segmentation architectures with a ResNet-50-FPN backbone on COCO val2017 shows varying sensitivity to point-level losses:

    Model Full Mask AP (MM) 10-Point AP (P10P_{10}) Relative Performance
    Mask R-CNN 37.2 36.1 97%
    CondInst 37.5 35.7^* 95%
    PointRend 38.3 35.7^* 93%

    Note: CondInst and PointRend were trained with point subsampling data augmentation, which improved their performance.

    While all three models achieve comparable absolute performance under P10P_{10} (35.735.7–36.136.1 AP), standard PointRend experiences the largest drop from full mask supervision (93%93\% vs. 97%97\% for Mask R-CNN), because its 7×77 \times 7 coarse mask head receives ambiguous supervision signals when conflicting points fall in the same coarse bin.

Coverage note — None was omitted; all key contributions—including the annotation scheme, training methodology, point augmentation, Implicit PointRend architecture, empirical results across multiple datasets and backbones, sensitivity analyses, self-training/transfer learning, and annotation time trade-offs—are fully covered.

References

  1. 1.Jiwoon Ahn, Sunghyun Cho, and Suha Kwak. Weakly supervised learning of instance segmentation with inter-pixel relations. In CVPR, 2019.
  2. 2.Aditya Arun, CV Jawahar, and M Pawan Kumar. Weakly supervised instance segmentation by learning annotation consistent instances. In ECCV, 2020.
  3. 3.Min Bai and Raquel Urtasun. Deep watershed transform for instance segmentation. In CVPR, 2017.
  4. 4.Amy Bearman, Olga Russakovsky, Vittorio Ferrari, and Li Fei-Fei. What’s the point: Semantic segmentation with point supervision. In ECCV, 2016.
  5. 5.Rodrigo Benenson, Stefan Popov, and Vittorio Ferrari. Large-scale interactive object segmentation with human annotators. In CVPR, 2019.
  6. 6.Daniel Bolya, Chong Zhou, Fanyi Xiao, and Yong Jae Lee. YOLACT++: Better real-time instance segmentation. PAMI, 2020.
  7. 7.Liang-Chieh Chen, Raphael Gontijo Lopes, Bowen Cheng, Maxwell D Collins, Ekin D Cubuk, Barret Zoph, Hartwig Adam, and Jonathon Shlens. Naive-student: Leveraging semi-supervised learning in video sequences for urban scene segmentation. In ECCV, 2020.
  8. 8.Bowen Cheng, Maxwell D Collins, Yukun Zhu, Ting Liu, Thomas S Huang, Hartwig Adam, and Liang-Chieh Chen. Panoptic-DeepLab: A simple, strong, and fast baseline for bottom-up panoptic segmentation. In CVPR, 2020.
  9. 9.Hisham Cholakkal, Guolei Sun, Fahad Shahbaz Khan, and Ling Shao. Object counting and instance segmentation with image-level supervision. In CVPR, 2019.
  10. 10.Herbert H Clark. Coordinating with each other in a material world. Discourse studies, 2005.
  11. 11.Marius Cordts, Mohamed Omran, Sebastian Ramos, Timo Rehfeld, Markus Enzweiler, Rodrigo Benenson, Uwe Franke, Stefan Roth, and Bernt Schiele. The Cityscapes dataset for semantic urban scene understanding. In CVPR, 2016.
  12. 12.Mark Everingham, SM Ali Eslami, Luc Van Gool, Christopher KI Williams, John Winn, and Andrew Zisserman. The PASCAL visual object classes challenge: A retrospective. IJCV, 2015.
  13. 13.Ruochen Fan, Qibin Hou, Ming-Ming Cheng, Gang Yu, Ralph R Martin, and Shi-Min Hu. Associating inter-image salient instances for weakly supervised semantic segmentation. In ECCV, 2018.
  14. 14.Chaz Firestone and Brian J Scholl. “please tap the shape, anywhere you like” shape skeletons in human vision revealed by an exceedingly simple measure. Psychological science, 2014.
  15. 15.Weifeng Ge, Sheng Guo, Weilin Huang, and Matthew R Scott. Label-PEnet: Sequential label propagation and enhancement networks for weakly supervised instance segmentation. In ICCV, 2019.
  16. 16.Kyle Genova, Forrester Cole, Avneesh Sud, Aaron Sarna, and Thomas Funkhouser. Local deep implicit functions for 3D shape. In CVPR, 2020.
  17. 17.Agrim Gupta, Piotr Dollar, and Ross Girshick. LVIS: A dataset for large vocabulary instance segmentation. In ICCV, 2019.
  18. 18.Michael Gygli and Vittorio Ferrari. Efficient object annotation via speaking and pointing. IJCV, 2019.
  19. 19.Kaiming He, Georgia Gkioxari, Piotr Dollar, and Ross Girshick. Mask R-CNN. In ICCV, 2017.
  20. 20.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In CVPR, 2016.
  21. 21.Cheng-Chun Hsu, Kuang-Jui Hsu, Chung-Chi Tsai, Yen-Yu Lin, and Yung-Yu Chuang. Weakly supervised instance segmentation using the bounding box tightness prior. In NIPS, 2019.
  22. 22.Chiyu Jiang, Avneesh Sud, Ameesh Makadia, Jingwei Huang, Matthias Nießner, and Thomas Funkhouser. Local implicit grid representations for 3D scenes. In CVPR, 2020.
  23. 23.Anna Khoreva, Rodrigo Benenson, Jan Hosang, Matthias Hein, and Bernt Schiele. Simple does it: Weakly supervised instance and semantic segmentation. In CVPR, 2017.
  24. 24.Alexander Kirillov, Evgeny Levinkov, Bjoern Andres, Bogdan Savchynskyy, and Carsten Rother. InstanceCut: from edges to instances with multicut. In CVPR, 2017.
  25. 25.Alexander Kirillov, Yuxin Wu, Kaiming He, and Ross Girshick. PointRend: Image segmentation as rendering. In CVPR, 2020.
  26. 26.Alina Kuznetsova, Hassan Rom, Neil Alldrin, Jasper Uijlings, Ivan Krasin, Jordi Pont-Tuset, Shahab Kamali, Stefan Popov, Matteo Malloci, Alexander Kolesnikov, Tom Duerig, and Vittorio Ferrari. The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale. IJCV, 2020.
  27. 27.Issam H Laradji, Negar Rostamzadeh, Pedro O Pinheiro, David Vazquez, and Mark Schmidt. Where are the blobs: Counting by localization with point supervision. In ECCV, 2018.
  28. 28.Issam H Laradji, Negar Rostamzadeh, Pedro O Pinheiro, David Vazquez, and Mark Schmidt. Proposal-based instance segmentation with point supervision. In ICIP, 2020.
  29. 29.Zhuwen Li, Qifeng Chen, and Vladlen Koltun. Interactive image segmentation with latent diversity. In CVPR, 2018.
  30. 30.JunHao Liew, Yunchao Wei, Wei Xiong, Sim-Heng Ong, and Jiashi Feng. Regional interactive image segmentation networks. In ICCV, 2017.
  31. 31.Di Lin, Jifeng Dai, Jiaya Jia, Kaiming He, and Jian Sun. ScribbleSup: Scribble-supervised convolutional networks for semantic segmentation. In CVPR, 2016.
  32. 32.Tsung-Yi Lin, Piotr Dollar, Ross Girshick, Kaiming He, Bharath Hariharan, and Serge Belongie. Feature pyramid networks for object detection. In CVPR, 2017.
  33. 33.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollar, and C Lawrence Zitnick. Microsoft COCO: Common objects in context. In ECCV, 2014.
  34. 34.Yun Liu, Yu-Huan Wu, Pei-Song Wen, Yu-Jun Shi, Yu Qiu, and Ming-Ming Cheng. Leveraging instance-, image- and dataset-level information for weakly supervised instance segmentation. PAMI, 2020.
  35. 35.Kevis-Kokitsi Maninis, Sergi Caelles, Jordi Pont-Tuset, and Luc Van Gool. Deep extreme cut: From extreme points to object segmentation. In CVPR, 2018.
  36. 36.Lars Mescheder, Michael Oechsle, Michael Niemeyer, Sebastian Nowozin, and Andreas Geiger. Occupancy networks: Learning 3D reconstruction in function space. In CVPR, 2019.
  37. 37.Pascal Mettes and Cees GM Snoek. Pointly-supervised action localization. IJCV, 2019.
  38. 38.Pascal Mettes, Jan C Van Gemert, and Cees GM Snoek. Spot on: Action localization from pointly-supervised proposals. In ECCV, 2016.
  39. 39.Fausto Milletari, Nassir Navab, and Seyed-Ahmad Ahmadi. V-Net: Fully convolutional neural networks for volumetric medical image segmentation. In International conference on 3D vision (3DV), 2016.
  40. 40.Vinod Nair and Geoffrey E Hinton. Rectified linear units improve restricted boltzmann machines. In ICML, 2010.
  41. 41.Dim P Papadopoulos, Jasper RR Uijlings, Frank Keller, and Vittorio Ferrari. Extreme clicking for efficient object annotation. In ICCV, 2017.
  42. 42.Dim P Papadopoulos, Jasper RR Uijlings, Frank Keller, and Vittorio Ferrari. Training object class detectors with click supervision. In CVPR, 2017.
  43. 43.Jordi Pont-Tuset, Pablo Arbelaez, Jonathan T Barron, Ferran Marques, and Jitendra Malik. Multiscale combinatorial grouping for image segmentation and object proposal generation. PAMI, 2016.
  44. 44.Rui Qian, Yunchao Wei, Honghui Shi, Jiachen Li, Jiaying Liu, and Thomas Huang. Weakly supervised scene parsing with point-based distance metric learning. In AAAI, 2019.
  45. 45.Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. Faster R-CNN: Towards real-time object detection with region proposal networks. In NIPS, 2015.
  46. 46.Zhongzheng Ren, Zhiding Yu, Xiaodong Yang, Ming-Yu Liu, Alexander G. Schwing, and Jan Kautz. UFO2^2: A unified framework towards omni-supervised object detection. In ECCV, 2020.
  47. 47.Ellen Riloff and Janyce Wiebe. Learning extraction patterns for subjective expressions. In EMNLP, 2003.
  48. 48.Carsten Rother, Vladimir Kolmogorov, and Andrew Blake. “GrabCut” interactive foreground extraction using iterated graph cuts. ACM transactions on graphics (TOG), 2004.
  49. 49.H Scudder. Probability of error of some adaptive pattern-recognition machines. IEEE Transactions on Information Theory, 1965.
  50. 50.Matthew Tancik, Pratul P Srinivasan, Ben Mildenhall, Sara Fridovich-Keil, Nithin Raghavan, Utkarsh Singhal, Ravi Ramamoorthi, Jonathan T Barron, and Ren Ng. Fourier features let networks learn high frequency functions in low dimensional domains. In NIPS, 2020.
  51. 51.Zhi Tian, Chunhua Shen, and Hao Chen. Conditional convolutions for instance segmentation. In ECCV, 2020.
  52. 52.Zhi Tian, Chunhua Shen, Xinlong Wang, and Hao Chen. BoxInst: High-performance instance segmentation with box annotations. In CVPR, 2021.
  53. 53.Jasper RR Uijlings, Koen EA Van De Sande, Theo Gevers, and Arnold WM Smeulders. Selective search for object recognition. IJCV, 2013.
  54. 54.Turner Whitted. An improved illumination model for shaded display. In ACM Siggraph 2005 Courses, 2005.
  55. 55.Yuxin Wu, Alexander Kirillov, Francisco Massa, Wan-Yen Lo, and Ross Girshick. Detectron2. https://github.com/facebookresearch/detectron2, 2019.
  56. 56.Qizhe Xie, Minh-Thang Luong, Eduard Hovy, and Quoc V Le. Self-training with noisy student improves imagenet classification. In CVPR, 2020.
  57. 57.Saining Xie, Ross Girshick, Piotr Dollár, Zhuowen Tu, and Kaiming He. Aggregated residual transformations for deep neural networks. In CVPR, 2017.
  58. 58.Ning Xu, Brian Price, Scott Cohen, Jimei Yang, and Thomas S Huang. Deep interactive object selection. In CVPR, 2016.
  59. 59.David Yarowsky. Unsupervised word sense disambiguation rivaling supervised methods. In ACL, 1995.
  60. 60.Yanzhao Zhou, Yi Zhu, Qixiang Ye, Qiang Qiu, and Jianbin Jiao. Weakly supervised instance segmentation using class peak response. In CVPR, 2018.
  61. 61.Yi Zhu, Yanzhao Zhou, Huijuan Xu, Qixiang Ye, David Doermann, and Jianbin Jiao. Learning instance activation maps for weakly supervised instance segmentation. In CVPR, 2019.
  62. 62.Barret Zoph, Golnaz Ghiasi, Tsung-Yi Lin, Yin Cui, Hanxiao Liu, Ekin D Cubuk, and Quoc V Le. Rethinking pre-training and self-training. In NIPS, 2020.

Citation

MLA
Cheng, B., et al. “Pointly-Supervised Instance Segmentation”. arXiv, 2021, http://arxiv.org/abs/2104.06404v2.
APA
Cheng, B., Parkhi, O., & Kirillov, A. (2021). Pointly-Supervised Instance Segmentation. arXiv. http://arxiv.org/abs/2104.06404v2
Chicago
Cheng, B., O. Parkhi, and A. Kirillov. 2021. “Pointly-Supervised Instance Segmentation”. arXiv. http://arxiv.org/abs/2104.06404v2.
Harvard
Cheng, B., Parkhi, O. and Kirillov, A. (2021) “Pointly-Supervised Instance Segmentation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2104.06404v2.
Vancouver
1. Cheng B, Parkhi O, Kirillov A (2021) Pointly-Supervised Instance Segmentation. arXiv

BibTeX

@article{cheng2021pointly,
  title = {Pointly-Supervised Instance Segmentation},
  author = {Cheng, Bowen and Parkhi, Omkar and Kirillov, Alexander},
  year = {2021},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2104.06404v2},
  eprint = {2104.06404}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE