Instance Segmentation with Mask-supervised Polygonal Boundary Transformers
Justin LazarowWeijian XuZhuowen Tu
Introduces BoundaryFormer, an end-to-end framework that directly predicts object boundary polygons supervised solely by pixel masks via differentiable rasterization, matching or outperforming Mask R-CNN on standard benchmarks without requiring polygon annotations.
In computer vision, instance segmentation identifies and delineates individual objects within an image. While modern applications such as 3D reconstruction and tracking benefit heavily from compact, structured boundary lines known as polygons, most existing systems predict dense, pixel-by-pixel masks. Previous attempts to predict polygons directly have struggled to match the accuracy of mask-based models, largely because standard training datasets and evaluation benchmarks are built around masks, and converting between the two formats introduces errors and engineering complexity.
The article demonstrates that an end-to-end boundary-predicting system can equal or exceed the performance of leading pixel-mask systems without requiring specialized polygon annotations. It introduces BoundaryFormer, an architecture designed to predict object contours directly while training entirely on standard pixel-mask labels.
The researchers evaluated BoundaryFormer using the benchmark MS-COCO dataset, containing approximately 118,000 natural training images across 80 object classes, and the Cityscapes dataset of complex urban driving scenes. The model starts with an initial coarse boundary within a detected bounding box and iteratively refines the vertex positions using attention mechanisms across four layers. To make the model trainable from standard masks, the authors developed a custom differentiable rasterizer implemented on graphical processing units. This tool continuously converts predicted polygons into soft masks during training, allowing the model to be guided by standard mask loss functions without intermediate 3D mesh triangulation.
The results show that BoundaryFormer achieves parity with and often outperforms established mask-based and boundary-based models. On the standard COCO benchmark, BoundaryFormer achieved an average precision of 36.4 with a standard backbone, surpassing the baseline Mask R-CNN score of 36.1 and outperforming the leading contour-based method, which scored 34.6. When scaled with larger network backbones and longer training, its accuracy reached 39.4. Additionally, when transferring a model pre-trained on COCO to the Cityscapes dataset, BoundaryFormer scored 38.3 average precision, noticeably outperforming Mask R-CNN at 36.5. An ablation test also showed that attempting to train point-based models on polygons derived artificially from masks reduced accuracy dramatically from 34.5 to 23.1, confirming the necessity of direct mask-supervised training.
These findings indicate that technical teams can deploy vector boundary representations as drop-in replacements for dense mask heads in standard detection pipelines without sacrificing accuracy. Vectorized boundaries produce continuous, differentiable outputs that eliminate boundary artifacts and internal holes, offering operational advantages for downstream geometric and spatial analysis. Moreover, the superior cross-dataset transferability suggests that polygon-based representations generalize better across differing visual domains, reducing retraining risk and adaptation costs.
Organizations developing spatial intelligence, autonomous navigation, or piece-wise 3D mapping systems should consider piloting boundary-regression modules in place of conventional mask heads. Future development should focus on extending the framework to handle heavily fragmented or occluded objects that require multiple separate polygon boundaries, as well as optimizing the architecture to further reduce memory and compute requirements. While confidence in the benchmark results is high across diverse standard datasets, stakeholders should note that the system still relies on dense pixel masks during training and currently approximates heavily occluded, multi-part objects with single continuous contours.
- Paper: Mask R-CNN, Kaiming He et al. (2017). Mask R-CNN established the canonical two-stage, proposal-based mask prediction framework against which BoundaryFormer benchmarks and integrates its polygon boundary regression.
- Paper: Deformable DETR: Deformable Transformers for End-to-End Object Detection, Xizhou Zhu et al. (2021). Deformable DETR provides the foundational deformable attention mechanisms that BoundaryFormer leverages across multiple layers for iterative vertex refinement.
- Paper: Per-Pixel Classification is Not All You Need for Semantic Segmentation, Bowen Cheng et al. (2021). MaskFormer establishes the mask-classification transformer paradigm, motivating BoundaryFormer's goal of training explicit geometric boundaries under mask supervision.
- Paper: Path Aggregation Network for Instance Segmentation, Shu Liu et al. (2018). PANet advances multi-scale feature propagation for instance segmentation, forming a core architectural reference for boundary and mask feature extraction.
- Paper: Hybrid Task Cascade for Instance Segmentation, Kai Chen et al. (2019). Hybrid Task Cascade demonstrates progressive multi-stage refinement for instance segmentation, informing iterative boundary decoding strategies.
- Paper: Cell Detection with Star-convex Polygons, Uwe Schmidt et al. (2018). StarDist pioneered direct parametric polygonal boundary prediction for instance segmentation as an explicit alternative to dense pixel masks.
- Paper: Simultaneous Detection and Segmentation, Bharath Hariharan et al. (2014). This seminal paper formulated simultaneous detection and segmentation, defining the dual instance-level localization and delineation problem tackled by BoundaryFormer.
- Paper: Masked-attention Mask Transformer for Universal Image Segmentation, Bowen Cheng et al. (2022). Mask2Former unifies transformer-based segmentation with masked attention, extending mask-level supervision principles across universal segmentation tasks.
- Paper: Mask Transfiner for High-Quality Instance Segmentation, Lei Ke et al. (2022). Mask Transfiner addresses high-frequency edge refinement in instance segmentation by adaptively processing sparse boundary regions via quadtree transformers.
- Paper: Pointly-Supervised Instance Segmentation, Bowen Cheng et al. (2022). Pointly-Supervised Instance Segmentation explores sparse point supervision for training segmentation heads, offering an alternative efficient supervision paradigm to differentiable rasterization.
- Paper: Superpoint Transformer for 3D Scene Instance Segmentation, Jiahao Sun et al. (2023). SPFormer generalizes transformer-based query decoders and geometric grouping to 3D point cloud instance segmentation.
- Paper: TubeFormer-DeepLab: Video Mask Transformer, Dahun Kim et al. (2022). TubeFormer-DeepLab extends mask transformer frameworks from static images to spatio-temporal video instance and panoptic segmentation.
