Instance Segmentation with Mask-supervised Polygonal Boundary Transformers

Justin LazarowWeijian XuZhuowen Tu

article2022CVPR66 citations

Introduces BoundaryFormer, an end-to-end framework that directly predicts object boundary polygons supervised solely by pixel masks via differentiable rasterization, matching or outperforming Mask R-CNN on standard benchmarks without requiring polygon annotations.

Listen

In computer vision, instance segmentation identifies and delineates individual objects within an image. While modern applications such as 3D reconstruction and tracking benefit heavily from compact, structured boundary lines known as polygons, most existing systems predict dense, pixel-by-pixel masks. Previous attempts to predict polygons directly have struggled to match the accuracy of mask-based models, largely because standard training datasets and evaluation benchmarks are built around masks, and converting between the two formats introduces errors and engineering complexity.

The article demonstrates that an end-to-end boundary-predicting system can equal or exceed the performance of leading pixel-mask systems without requiring specialized polygon annotations. It introduces BoundaryFormer, an architecture designed to predict object contours directly while training entirely on standard pixel-mask labels.

The researchers evaluated BoundaryFormer using the benchmark MS-COCO dataset, containing approximately 118,000 natural training images across 80 object classes, and the Cityscapes dataset of complex urban driving scenes. The model starts with an initial coarse boundary within a detected bounding box and iteratively refines the vertex positions using attention mechanisms across four layers. To make the model trainable from standard masks, the authors developed a custom differentiable rasterizer implemented on graphical processing units. This tool continuously converts predicted polygons into soft masks during training, allowing the model to be guided by standard mask loss functions without intermediate 3D mesh triangulation.

The results show that BoundaryFormer achieves parity with and often outperforms established mask-based and boundary-based models. On the standard COCO benchmark, BoundaryFormer achieved an average precision of 36.4 with a standard backbone, surpassing the baseline Mask R-CNN score of 36.1 and outperforming the leading contour-based method, which scored 34.6. When scaled with larger network backbones and longer training, its accuracy reached 39.4. Additionally, when transferring a model pre-trained on COCO to the Cityscapes dataset, BoundaryFormer scored 38.3 average precision, noticeably outperforming Mask R-CNN at 36.5. An ablation test also showed that attempting to train point-based models on polygons derived artificially from masks reduced accuracy dramatically from 34.5 to 23.1, confirming the necessity of direct mask-supervised training.

These findings indicate that technical teams can deploy vector boundary representations as drop-in replacements for dense mask heads in standard detection pipelines without sacrificing accuracy. Vectorized boundaries produce continuous, differentiable outputs that eliminate boundary artifacts and internal holes, offering operational advantages for downstream geometric and spatial analysis. Moreover, the superior cross-dataset transferability suggests that polygon-based representations generalize better across differing visual domains, reducing retraining risk and adaptation costs.

Organizations developing spatial intelligence, autonomous navigation, or piece-wise 3D mapping systems should consider piloting boundary-regression modules in place of conventional mask heads. Future development should focus on extending the framework to handle heavily fragmented or occluded objects that require multiple separate polygon boundaries, as well as optimizing the architecture to further reduce memory and compute requirements. While confidence in the benchmark results is high across diverse standard datasets, stakeholders should note that the system still relies on dense pixel masks during training and currently approximates heavily occluded, multi-part objects with single continuous contours.

  • Paper: Mask R-CNN, Kaiming He et al. (2017). Mask R-CNN established the canonical two-stage, proposal-based mask prediction framework against which BoundaryFormer benchmarks and integrates its polygon boundary regression.
  • Paper: Deformable DETR: Deformable Transformers for End-to-End Object Detection, Xizhou Zhu et al. (2021). Deformable DETR provides the foundational deformable attention mechanisms that BoundaryFormer leverages across multiple layers for iterative vertex refinement.
  • Paper: Per-Pixel Classification is Not All You Need for Semantic Segmentation, Bowen Cheng et al. (2021). MaskFormer establishes the mask-classification transformer paradigm, motivating BoundaryFormer's goal of training explicit geometric boundaries under mask supervision.
  • Paper: Path Aggregation Network for Instance Segmentation, Shu Liu et al. (2018). PANet advances multi-scale feature propagation for instance segmentation, forming a core architectural reference for boundary and mask feature extraction.
  • Paper: Hybrid Task Cascade for Instance Segmentation, Kai Chen et al. (2019). Hybrid Task Cascade demonstrates progressive multi-stage refinement for instance segmentation, informing iterative boundary decoding strategies.
  • Paper: Cell Detection with Star-convex Polygons, Uwe Schmidt et al. (2018). StarDist pioneered direct parametric polygonal boundary prediction for instance segmentation as an explicit alternative to dense pixel masks.
  • Paper: Simultaneous Detection and Segmentation, Bharath Hariharan et al. (2014). This seminal paper formulated simultaneous detection and segmentation, defining the dual instance-level localization and delineation problem tackled by BoundaryFormer.
Cover for Instance Segmentation with Mask-supervised Polygonal Boundary Transformers

Abstract

In this paper, we present an end-to-end instance segmentation method that regresses a polygonal boundary for each object instance. This sparse, vectorized boundary representation for objects, while attractive in many downstream computer vision tasks, quickly runs into issues of parity that need to be addressed: parity in supervision and parity in performance when compared to existing pixel-based methods. This is due in part to object instances being annotated with ground-truth in the form of polygonal boundaries or segmentation masks, yet being evaluated in a conventional manner using only segmentation masks. Our method, BoundaryFormer, is a Transformer based architecture that directly predicts polygons yet uses instance mask segmentations as the ground-truth supervision for computing the loss. We achieve this by developing an end-to-end differentiable model that solely relies on supervision within the mask space through differentiable rasterization. BoundaryFormer matches or surpasses the Mask R-CNN method in terms of instance segmentation quality on both COCO and Cityscapes while exhibiting significantly better transferability across datasets.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 2.1. Point-based Losses
  • 2.2. Mask-based Losses
  • 3. Method
  • 3.1. Setting
  • 3.2. Instance Segmentation with Mask-supervised Boundary Regression Transformers
  • 3.3. Mask-based Supervision
  • 3.3.1. A polygon-specific rasterizer
  • 3.3.2. Alignment
  • 3.4. Coarse to fine upsampling
  • 3.5. Loss
  • 4. Experiments
  • 4.1. Training Details
  • 4.2. COCO
  • 4.3. Cityscapes
  • 4.4. Ablation studies
  • 5. Qualitative Analysis
  • 6. Conclusion
  • References

Knowls

  1. Knowl 1 — BoundaryFormer’s polygonal instance representation

    model/method

    BoundaryFormer replaces a conventional dense mask-prediction head with a polygon-regression head. For each detected object instance ii, the model predicts an ordered set of KK two-dimensional vertices Vi={(xi,0,yi,0),…,(xi,K−1,yi,K−1)}V_i=\{(x_{i,0},y_{i,0}),\ldots,(x_{i,K-1},y_{i,K-1})\} that defines the object boundary. The polygon is rasterized into a soft mask during training and into a hard mask at inference, so the final prediction remains compatible with standard mask-based instance-segmentation evaluation.

    The model uses only pixel-wise instance masks as ground-truth supervision and does not require annotated polygons. Consequently, it can serve as a drop-in replacement for a mask head in both two-stage detectors such as R-CNN and single-stage detectors such as FCOS, with either region-of-interest features or full-image features.

  2. Knowl 2 — Transformer boundary-refinement architecture

    model/method

    Given an image I∈RH×W×3I\in\mathbb{R}^{H\times W\times 3}, an FPN detector produces multiscale feature maps F={P2,P3,P4,P5}F=\{P_2,P_3,P_4,P_5\} and object boxes Bi=(li,ti,wi,hi)B_i=(l_i,t_i,w_i,h_i), where li,til_i,t_i are the left and top coordinates and wi,hiw_i,h_i are width and height. BoundaryFormer initializes an ellipse inscribed in each predicted box as the initial polygon Vi(0)V_i^{(0)}.

    The polygon is refined through LL Transformer layers according to

    Vi(j+1)=gj ⁣(F,Vi(j)),j=0,…,L−1,V_i^{(j+1)}=g_j\!\left(F,V_i^{(j)}\right),\qquad j=0,\ldots,L-1,

    where gjg_j predicts a two-dimensional offset for every polygon vertex using an MLP. Each vertex has a learned point embedding augmented with a Transformer-style point positional encoding. Within each object, point embeddings use ordinary self-attention, while point embeddings attend to the multiscale FPN features using deformable attention. The resulting architecture treats the boundary as a structured set of point queries while retaining access to image evidence at multiple feature resolutions.

  3. Knowl 3 — Direct differentiable polygon-to-mask rasterization

    equation

    For a predicted polygon VV and a target rasterization grid of size X′×Y′X'\times Y', BoundaryFormer processes every pixel index (x,y)(x,y) with 0≤x<X′0\le x<X' and 0≤y<Y′0\le y<Y'. A point-in-polygon test produces a sign C(V,x,y)C(V,x,y) equal to 11 when the pixel lies inside VV and −1-1 otherwise. The pixel is projected to the closest segment of the polygon boundary, and the resulting Euclidean distance in pixel units is denoted by D(V,x,y)≥0D(V,x,y)\ge 0. The differentiable rasterized value is

    I(x,y)=σ ⁣(C(V,x,y)D(V,x,y)τ),I(x,y)=\sigma\!\left(\frac{C(V,x,y)D(V,x,y)}{\tau}\right),

    where I(x,y)∈(0,1)I(x,y)\in(0,1) is the soft-mask value, σ(z)=1/(1+e−z)\sigma(z)=1/(1+e^{-z}) is the logistic sigmoid, and τ>0\tau>0 is the rasterization-smoothness parameter. The formulation supplies gradients from the polygon boundary directly and avoids triangulating the polygon into a mesh. The rasterizer and its backward pass were implemented in CUDA, allowing it to be used throughout training.

  4. Knowl 4 — Mask-only objective with supervision at every refinement stage

    equation

    Let Lbbox\mathcal{L}_{\mathrm{bbox}} be the detector’s bounding-box loss, let RR be the number of foreground-matched detected boxes in an image, let MiM_i be the ground-truth binary mask for matched instance ii at resolution X′×Y′X'\times Y', and let Ii(ℓ)I_i^{(\ell)} be the differentiably rasterized polygon predicted at refinement layer ℓ\ell. BoundaryFormer minimizes

    L=Lbbox+∑ℓ=1L∑i=1RDice⁡ ⁣(Ii(ℓ),Mi).\mathcal{L}=\mathcal{L}_{\mathrm{bbox}}+\sum_{\ell=1}^{L}\sum_{i=1}^{R}\operatorname{Dice}\!\left(I_i^{(\ell)},M_i\right).

    Here LL is the number of polygon-refinement layers and Dice⁡(⋅,⋅)\operatorname{Dice}(\cdot,\cdot) is the Dice mask loss. Every intermediate polygon and the final polygon are therefore trained against the same mask target. The objective requires no polygonal ground truth and optimizes the representation in the same mask space used for evaluation.

  5. Knowl 5 — Coarse-to-fine polygon upsampling

    algorithm

    BoundaryFormer reduces the cost of processing many polygon vertices by starting with a small number of control points and doubling the polygon resolution after each refinement stage. If Vi(ℓ)V_i^{(\ell)} contains K(ℓ)K^{(\ell)} ordered vertices, the refinement first updates those existing vertices and then inserts one new vertex between every consecutive pair. The inserted vertex is initialized at the geometric midpoint, while its point embedding is a separate learned embedding rather than the average of the neighboring embeddings.

    Input: Initial polygon V_i with K^(0) vertices, feature maps F, and L refinement layers
    Output: Final polygon V_i^(L)
    for layer ell from 0 to L - 1 do
        Refine every existing vertex using Transformer attention to polygon points and F
        Predict a 2D offset for each existing vertex and update V_i
        for every consecutive vertex pair (x_j, y_j), (x_(j+1), y_(j+1)) do
            Insert midpoint ((x_j + x_(j+1))/2, (y_j + y_(j+1))/2)
            Initialize the inserted point with its learned point embedding
        end for
    end for
    return V_i^(L)

    The paper uses a base control-point count that is usually 88, four Transformer layers, and final polygon sizes of 6464 points on COCO and 128128 points on Cityscapes. Processing fewer points in early layers reduces the nominal work from approximately O(O K L)O(O\,K\,L) for OO objects, KK points, and LL layers; the authors report about a 1.5×1.5\times training and memory improvement over using the final point count at every layer.

  6. Knowl 6 — Training and rasterization alignment choices

    experimental setup

    The standard comparison uses a ResNet-50-FPN backbone and keeps the detector settings matched to Mask R-CNN except for replacing its mask head with BoundaryFormer. Models are trained end-to-end with Adam using settings based on Swin Transformer training; larger models use weight decay 0.200.20. Deformable-attention settings follow the Deformable DETR decoder.

    During training, polygons are rasterized inside their predicted boxes at a fixed 64×6464\times64 resolution, and the ground-truth masks are clipped to the same boxes. The smoothness parameter is set to τ=0.1\tau=0.1. At inference, differentiability is unnecessary, so masks are rasterized with the COCO API’s ordinary hard rasterizer rather than the training rasterizer.

    Coordinate alignment is essential: BoundaryFormer subtracts half a pixel from all polygon coordinates before rasterization to match the pixel convention of the COCO API. Without this correction, mask AP falls from 36.136.1 to 35.335.3. In the RoI-less full-image setting, adding a global-pooling-like operation to the feature processing raises mask AP from 34.234.2 to 36.136.1; the authors hypothesize that the difference is caused by aliasing in FPN deconvolutional features.

  7. Knowl 7 — COCO instance-segmentation performance

    data/table

    On MS-COCO validation, BoundaryFormer reaches mask quality comparable to or better than mask-based and contour-based baselines while retaining polygonal outputs. APAP and AP50AP_{50} are mask average precision at the standard and 0.500.50 IoU thresholds, respectively; APbboxAP_{\mathrm{bbox}} is bounding-box AP. An asterisk denotes a model retrained with Adam.

    Method Backbone Detector AP AP50_{50} APbbox_{\mathrm{bbox}}
    Mask R-CNN R50-FPN R-CNN 35.2 56.3 38.6
    Mask R-CNN∗^* R50-FPN R-CNN 35.8 56.8 38.8
    BMask R-CNN R50-FPN R-CNN 36.6 56.7 39.4
    BMask R-CNN∗^* R50-FPN R-CNN 36.4 56.3 37.8
    DANCE R50-FPN FCOS 34.5 55.3 40.2
    BoundaryFormer R50-FPN FCOS 35.8 55.7 40.2
    BoundaryFormer R50-FPN R-CNN 36.1 56.7 38.8

    On COCO test-dev, BoundaryFormer also scales with model size and training duration: its AP is 36.436.4 with an R50-FPN R-CNN and a 1×1\times schedule, 37.737.7 with an R101-FPN and a 1×1\times schedule, and 39.439.4 with an R101-FPN and a 3×3\times schedule. The corresponding Mask R-CNN results are 36.136.1 for R50-FPN with a 1×1\times schedule and 39.239.2 for R101-FPN with a 3×3\times schedule. Thus, the polygon output does not impose a measurable COCO mask-quality penalty relative to the strong mask baseline.

  8. Knowl 8 — Cityscapes transfer performance

    data/table

    BoundaryFormer maintains competitive performance on Cityscapes despite the dataset’s frequent fragmented objects, for which a single polygon is intrinsically restrictive. APAP and AP50AP_{50} denote mask average precision at the standard and 0.500.50 IoU thresholds. The supervision column records the source of training supervision, and the end-to-end column indicates whether the entire model is differentiable and trained jointly. A dash indicates that the cited result did not report the metric.

    Method Initialization Polygon initialization End-to-end Supervision AP AP50_{50}
    Mask R-CNN∗^* ImageNet N/A no masks 34.2 60.7
    Mask R-CNN∗^* COCO N/A no masks 36.5 62.0
    DeepSnake ImageNet extreme points yes polygons 28.2 –
    UPSNet COCO predicted masks no masks 37.8 –
    PolyTransform COCO predicted masks no both 40.2 –
    BoundaryFormer ImageNet ellipse yes masks 34.7 60.8
    BoundaryFormer COCO ellipse yes masks 38.3 62.9

    The BoundaryFormer results are averages over three Cityscapes runs. With ImageNet initialization, BoundaryFormer is essentially tied with the ImageNet-initialized Mask R-CNN. With COCO initialization, it exceeds the COCO-initialized Mask R-CNN by 1.81.8 AP, reaching 38.338.3 AP, while requiring only mask supervision and remaining end-to-end differentiable.

  9. Knowl 9 — Ablation evidence for mask supervision and polygon resolution

    empirical result

    The experiments support both direct mask supervision and the selected rasterization settings. When DANCE, a point-supervised contour model, is trained on the original annotated polygons, it obtains 34.534.5 mask AP on COCO. Replacing those polygons with contours generated from masks using a standard border-following procedure reduces its mask AP to 23.123.1. This demonstrates that converting masks into pseudo-polygons can introduce substantial supervision errors.

    BoundaryFormer’s COCO validation ablations are:

    Rasterization resolution 14×\times14 20×\times20 40×\times40 64×\times64 80×\times80
    Mask AP 34.5 35.5 35.9 36.1 36.1

    For rasterization smoothness, values around τ=0.1\tau=0.1 are robust, whereas τ=1.0\tau=1.0 reduces mask AP to 35.635.6 because the soft rasterization is not sharp enough. The reported point-count/layer ablations are:

    L/K1/KLL/K_1/K_L 2/32/64 3/16/64 4/4/32 4/8/64 4/64/64 4/16/128
    Mask AP 35.0 35.7 35.2 36.1 36.1 36.2

    Two refinement layers incur only a one-point AP drop relative to the best configuration, fewer points produce a similar moderate drop, and coarse-to-fine point growth performs as well as using the dense final point count at every layer.

  10. Knowl 10 — Limitations of a sparse single-polygon output

    limitation

    BoundaryFormer predicts one polygon for each detected object, so it cannot always represent fragmented or multiply connected objects faithfully. This is especially problematic on Cityscapes, where occlusion often causes one annotated instance to consist of several disconnected regions. The model nevertheless learns a reasonable single-polygon approximation from mask supervision, but it does not explicitly predict multiple polygons or introduce special fragmentation handling.

    A second limitation is that the sparse boundary representation still requires dense object masks for training. The paper identifies reducing this dense supervision to a sparser, more boundary-centric signal as an open direction, along with incorporating further mask-based advances and designing more computationally efficient architectures.

Coverage note — The qualitative visual examples and societal-impact discussion were not made separate knowls because they illustrate the quantified behavior or general risks already captured above rather than adding independent technical results.

References

  1. 1.Pnpoly - point inclusion in polygon test. https://wrf.ecse.rpi.edu/Research/Short_Notes/pnpoly.html. Accessed: 2021-10-30. 5
  2. 2.Serge Belongie, Jitendra Malik, and Jan Puzicha. Shape matching and object recognition using shape contexts. IEEE transactions on pattern analysis and machine intelligence, 24(4):509–522, 2002. 1
  3. 3.Maxim Berman, Amal Rannen Triki, and Matthew B Blaschko. The lovasz-softmax loss: A tractable surrogate for the optimization of the intersection-over-union measure in neural networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 4413–4421, 2018. 5
  4. 4.Andrew Blake and Michael Isard. Active contours: the application of techniques from graphics, vision, control theory and statistics to visual tracking of shapes in motion. Springer Science & Business Media, 2012. 1
  5. 5.Anne-Laure Chauve, Patrick Labatut, and Jean-Philippe Pons. Robust piecewise-planar 3d reconstruction and completion from large-scale unstructured point data. In CVPR, pages 1261–1268, 2010. 1
  6. 6.Tianheng Cheng, Xinggang Wang, Lichao Huang, and Wenyu Liu. Boundary-preserving mask r-cnn. In ECCV, pages 660–676, 2020. 3, 6, 7
  7. 7.Marius Cordts, Mohamed Omran, Sebastian Ramos, Timo Rehfeld, Markus Enzweiler, Rodrigo Benenson, Uwe Franke, Stefan Roth, and Bernt Schiele. The cityscapes dataset for semantic urban scene understanding. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 3213–3223, 2016. 2, 5, 6
  8. 8.Agrim Gupta, Piotr Dollar, and Ross Girshick. Lvis: A dataset for large vocabulary instance segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5356–5364, 2019. 7
  9. 9.Shir Gur, Tal Shaharabany, and Lior Wolf. End to end trainable active contours via differentiable rendering. arXiv preprint arXiv:1912.00367, 2019. 2, 3, 4
  10. 10.Kaiming He, Georgia Gkioxari, Piotr Dollar, and Ross Girshick. Mask r-cnn. In Proceedings of the IEEE international conference on computer vision, pages 2961–2969, 2017. 1, 2, 3, 5, 6, 7
  11. 11.Hiroharu Kato, Yoshitaka Ushiku, and Tatsuya Harada. Neural 3d mesh renderer. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018. 3
  12. 12.Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014. 6, 7
  13. 13.Ke Li, Shijie Wang, Xiang Zhang, Yifan Xu, Weijian Xu, and Zhuowen Tu. Pose recognition with cascade transformers. In CVPR, pages 1944–1953, 2021. 2
  14. 14.Justin Liang, Namdar Homayounfar, Wei-Chiu Ma, Yuwen Xiong, Rui Hu, and Raquel Urtasun. Polytransform: Deep polygon transformer for instance segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9131–9140, 2020. 2, 3, 4, 6, 7
  15. 15.Tsung-Yi Lin, Piotr Dollar, Ross Girshick, Kaiming He, Bharath Hariharan, and Serge Belongie. Feature pyramid networks for object detection. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2117–2125, 2017. 5
  16. 16.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollar, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In European conference on computer vision, pages 740–755. Springer, 2014. 2, 3, 4, 5, 6
  17. 17.Huan Ling, Jun Gao, Amlan Kar, Wenzheng Chen, and Sanja Fidler. Fast interactive object annotation with curve-gcn. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5257–5266, 2019. 3, 4
  18. 18.Chen Liu, Jimei Yang, Duygu Ceylan, Ersin Yumer, and Yasutaka Furukawa. Planenet: Piece-wise planar reconstruction from a single rgb image. In CVPR, 2018. 2
  19. 19.Shichen Liu, Tianye Li, Weikai Chen, and Hao Li. Soft rasterizer: A differentiable renderer for image-based 3d reasoning. The IEEE International Conference on Computer Vision (ICCV), Oct 2019. 5
  20. 20.Zichen Liu, Jun Hao Liew, Xiangyu Chen, and Jiashi Feng. Dance: A deep attentive contour model for efficient instance segmentation. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), pages 345–354, January 2021. 3, 4, 6, 7
  21. 21.Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. arXiv preprint arXiv:2103.14030, 2021. 6
  22. 22.Matthew M Loper and Michael J Black. Opendr: An approximate differentiable renderer. In European Conference on Computer Vision, pages 154–169. Springer, 2014. 3
  23. 23.David Martin, Charless Fowlkes, Doron Tal, and Jitendra Malik. A database of human segmented natural images and its application to evaluating segmentation algorithms and measuring ecological statistics. In ICCV, volume 2, pages 416–423, 2001. 1
  24. 24.Fausto Milletari, Nassir Navab, and Seyed-Ahmad Ahmadi. V-net: Fully convolutional neural networks for volumetric medical image segmentation. In 2016 fourth international conference on 3D vision (3DV), pages 565–571. IEEE, 2016. 5
  25. 25.Sida Peng, Wen Jiang, Huaijin Pi, Xiuli Li, Hujun Bao, and Xiaowei Zhou. Deep snake for real-time instance segmentation. In CVPR, 2020. 2, 3, 7
  26. 26.Dzung L Pham, Chenyang Xu, and Jerry L Prince. Current methods in medical image segmentation. Annual review of biomedical engineering, 2(1):315–337, 2000. 1
  27. 27.Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. Faster r-cnn: Towards real-time object detection with region proposal networks. Advances in neural information processing systems, 28:91–99, 2015. 5
  28. 28.Jamie Shotton, John Winn, Carsten Rother, and Antonio Criminisi. Textonboost: Joint appearance, shape and context modeling for multi-class object recognition and segmentation. In ECCV, 2006. 1
  29. 29.Zhi Tian, Chunhua Shen, Hao Chen, and Tong He. Fcos: Fully convolutional one-stage object detection. In Proceedings of the IEEE/CVF international conference on computer vision, pages 9627–9636, 2019. 2, 3, 5
  30. 30.Zhuowen Tu. Auto-context and its application to high-level vision tasks. In CVPR, 2008. 1
  31. 31.Zhuowen Tu, Xiangrong Chen, Alan L Yuille, and Song-Chun Zhu. Image parsing: Unifying segmentation, detection, and recognition. International Journal of computer vision, 63(2):113–140, 2005. 1
  32. 32.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Advances in neural information processing systems, pages 5998–6008, 2017. 3, 6
  33. 33.Yuxin Wu, Alexander Kirillov, Francisco Massa, Wan-Yen Lo, and Ross Girshick. Detectron2. https://github.com/facebookresearch/detectron2, 2019. 6
  34. 34.Enze Xie, Peize Sun, Xiaoge Song, Wenhai Wang, Xuebo Liu, Ding Liang, Chunhua Shen, and Ping Luo. Polarmask: Single shot instance segmentation with polar representation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 12193–12202, 2020. 2
  35. 35.Yifan Xu, Weijian Xu, David Cheung, and Zhuowen Tu. Line segment detection using transformers without edges. In ICCV, pages 4257–4266, 2021. 2
  36. 36.Xizhou Zhu, Weijie Su, Lewei Lu, Bin Li, Xiaogang Wang, and Jifeng Dai. Deformable detr: Deformable transformers for end-to-end object detection. arXiv preprint arXiv:2010.04159, 2020. 3, 6

Citation

MLA
Lazarow, J., et al. “Instance Segmentation with Mask-supervised Polygonal Boundary Transformers”. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022, pp. 4372–81, https://doi.org/10.1109/CVPR52688.2022.00434.
APA
Lazarow, J., Xu, W., & Tu, Z. (2022). Instance Segmentation with Mask-supervised Polygonal Boundary Transformers. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 4372–4381. https://doi.org/10.1109/CVPR52688.2022.00434
Chicago
Lazarow, J., W. Xu, and Z. Tu. 2022. “Instance Segmentation with Mask-supervised Polygonal Boundary Transformers”. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 4372–81. https://doi.org/10.1109/CVPR52688.2022.00434.
Harvard
Lazarow, J., Xu, W. and Tu, Z. (2022) “Instance Segmentation with Mask-supervised Polygonal Boundary Transformers”, 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp. 4372–4381. Available at: https://doi.org/10.1109/CVPR52688.2022.00434.
Vancouver
1. Lazarow J, Xu W, Tu Z (2022) Instance Segmentation with Mask-supervised Polygonal Boundary Transformers. In: 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp 4372–4381

BibTeX

@inproceedings{Lazarow_2022, title={Instance Segmentation with Mask-supervised Polygonal Boundary Transformers}, url={http://dx.doi.org/10.1109/CVPR52688.2022.00434}, DOI={10.1109/cvpr52688.2022.00434}, booktitle={2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)}, publisher={IEEE}, author={Lazarow, Justin and Xu, Weijian and Tu, Zhuowen}, year={2022}, month=June, pages={4372–4381} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE