Mask Transfiner for High-Quality Instance Segmentation

Lei KeMartin DanelljanXia LiYu-Wing TaiChi-Keung TangFisher Yu

article2022CVPR169 citations

Proposes a transformer-based framework that represents image regions as quadtrees to efficiently detect and correct errors in boundary pixels, significantly boosting instance mask accuracy across two-stage and query-based detectors on COCO, BDD100K, and Cityscapes benchmarks.

Listen

Modern computer vision systems rely heavily on instance segmentation to detect objects and delineate their exact boundaries. While recent detection algorithms have advanced rapidly, the quality of predicted segmentation masks remains coarse, frequently oversmoothing object edges and missing fine details. Resolving these errors traditionally requires processing entire high-resolution feature maps on uniform pixel grids, which incurs prohibitively high computational and memory costs for real-world deployments.

The article introduces and evaluates Mask Transfiner, an end-to-end framework designed to produce high-precision instance segmentation masks efficiently. The authors' primary objective is to demonstrate that focusing computation exclusively on sparse, error-prone boundary regions across multiple feature scales can bridge the accuracy gap between object detection and detailed mask generation without excessive computational overhead.

The approach identifies "incoherent regions"—areas where spatial detail is lost during downsampling—and represents them in a hierarchical quadtree data structure. A lightweight convolutional detector predicts these incoherent areas across multi-scale feature pyramids. Instead of uniform grid convolutions, a transformer-based refinement network jointly processes the sparse quadtree nodes in parallel, incorporating fine-grained features, semantic priors, spatial positions, and surrounding context. The framework was evaluated across standard computer vision benchmarks including the COCO, Cityscapes, and BDD100K datasets, testing both two-stage and query-based base detectors.

The analysis yielded several key findings. First, prediction errors are overwhelmingly concentrated in incoherent regions: while these areas account for only 14% of an object's bounding box area, they contain 43% of all wrongly predicted pixels. Second, Mask Transfiner sets state-of-the-art results across major benchmarks, improving mask accuracy by up to 3.0 points on COCO and BDD100K, and raising boundary accuracy by 6.6 points on Cityscapes over baseline models. Third, the quadtree design achieves these gains with high computational efficiency, cutting memory usage by approximately two-thirds compared to non-local attention baselines and producing high-resolution outputs using half the floating-point operations of standard transformers.

These findings indicate that instance segmentation systems can achieve sharp boundary fidelity without incurring unsustainable computing and memory costs. For applied domains such as autonomous driving and urban scene understanding, high-frequency boundary accuracy directly reduces risks associated with misidentifying small object details like side mirrors or thin obstacles. Furthermore, the framework's compatibility with both two-stage and query-based architectures makes it broadly applicable to existing computer vision pipelines.

Organizations developing high-precision visual perception systems should consider adopting sparse, hierarchical quadtree refinement to replace uniform high-resolution processing. When configuring the system, teams should select a three-level quadtree depth with an output resolution around 112x112, where accuracy gains and operational speed (roughly 7 frames per second) achieve an optimal balance before returns diminish at higher resolutions.

The primary limitation noted in the article is the reliance on fully supervised training requiring detailed ground-truth mask annotations. While confidence in the benchmark improvements is high across diverse backbones and datasets, future initiatives should focus on extending the framework to weakly supervised or semi-supervised settings to reduce data annotation costs in production environments.

  • Paper: Mask R-CNN, Kaiming He et al. (2017). Mask R-CNN establishes the foundational two-stage instance segmentation paradigm whose coarse mask predictions Mask Transfiner specifically sets out to refine.
  • Paper: Cascade R-CNN: High Quality Object Detection and Instance Segmentation, Zhaowei Cai et al. (2019). Cascade R-CNN introduces progressive multi-stage refinement for detection and segmentation, directly motivating the high-quality mask refinement objectives of Mask Transfiner.
  • Paper: Path Aggregation Network for Instance Segmentation, Shu Liu et al. (2018). PANet enhances feature hierarchy and spatial localization in instance segmentation networks, serving as an architectural baseline improved by Mask Transfiner's sparse refinement.
  • Paper: Hybrid Task Cascade for Instance Segmentation, Kai Chen et al. (2019). Hybrid Task Cascade introduces intertwined multi-stage mask feature refinement, providing key context on multi-stage instance segmentation pipelines refined by Transfiner.
  • Paper: Feature Pyramid Networks for Object Detection, Tsung-Yi Lin et al. (2017). Feature Pyramid Networks supply the foundational multi-scale feature representation underlying the two-stage instance segmentation architectures that Mask Transfiner builds on.
  • Paper: End-to-End Object Detection with Transformers, Nicolas Carion et al. (2020). DETR introduces end-to-end query-based transformer modeling for vision, forming the foundation of the query-based instance segmentation frameworks enhanced by Transfiner.
  • Paper: Per-Pixel Classification is Not All You Need for Semantic Segmentation, Bowen Cheng et al. (2021). MaskFormer establishes transformer-based mask classification for segmentation, representing the query-based segmentation paradigm whose boundary quality Transfiner aims to rectify.
Cover for Mask Transfiner for High-Quality Instance Segmentation

Abstract

Two-stage and query-based instance segmentation methods have achieved remarkable results. However, their segmented masks are still very coarse. In this paper, we present Mask Transfiner for high-quality and efficient instance segmentation. Instead of operating on regular dense tensors, our Mask Transfiner decomposes and represents the image regions as a quadtree. Our transformer-based approach only processes detected error-prone tree nodes and self-corrects their errors in parallel. While these sparse pixels only constitute a small proportion of the total number, they are critical to the final mask quality. This allows Mask Transfiner to predict highly accurate instance masks, at a low computational cost. Extensive experiments demonstrate that Mask Transfiner outperforms current instance segmentation methods on three popular benchmarks, significantly improving both two-stage and query-based frameworks by a large margin of +3.0 mask AP on COCO and BDD100K, and +6.6 boundary AP on Cityscapes. Our code and trained models are available at https://github.com/SysCV/transfiner.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Mask Transfiner
  • 3.1. Incoherent Regions
  • 3.2. Quadtree for Mask Refinement
  • 3.3. Mask Transfiner Architecture
  • 4. Experiments
  • 4.1. Experimental Setup
  • 4.2. Ablation Experiments
  • 4.3. Comparison with State-of-the-art
  • 5. Conclusion
  • References

Knowls

  1. Knowl 1 — Definition and Mathematical Formulation of Incoherent Mask Regions

    definition

    In instance segmentation, spatial downsampling operations and coarse feature representations cause substantial loss of high-frequency spatial details and object boundary accuracy. An incoherent region is formally defined as the set of pixels where binary mask information is lost when downsampled and subsequently reconstructed via upsampling.

    Let MlM_l denote the binary ground-truth instance mask of an object at scale level l∈{0,…,L}l \in \{0, \dots, L\}, where l=0l = 0 represents the finest scale and l=Ll = L represents the coarsest scale, with resolutions between adjacent levels differing by a factor of 2. Let S↓\mathcal{S}_\downarrow and S↑\mathcal{S}_\uparrow denote 2×2\times nearest-neighbor downsampling and upsampling operations, respectively. The binary incoherent region map DlD_l at scale level ll is defined as:

    Dl=O↓(Ml−1⊕S↑(S↓(Ml−1)))D_l = \mathcal{O}_\downarrow \left(M_{l-1} \oplus \mathcal{S}_\uparrow (\mathcal{S}_\downarrow (M_{l-1}))\right)

    where ⊕\oplus is the logical exclusive OR (XOR) operation, and O↓\mathcal{O}_\downarrow is a 2×2\times downsampling operation that performs a logical OR over each 2×22 \times 2 pixel neighborhood.

    A pixel (x,y)(x, y) is classified as incoherent (Dl(x,y)=1D_l(x, y) = 1) if the original fine mask Ml−1M_{l-1} differs from its downsampled-and-upsampled reconstruction in at least one pixel within the corresponding 2×22 \times 2 finer window. Incoherent regions are sparsely distributed along object boundaries and high-frequency patterns.

  2. Knowl 2 — Hierarchical Point Quadtree Construction and Incoherent Region Detector

    model/method

    Mask Transfiner decomposes ambiguous image regions into a multi-scale point quadtree to selectively process sparse, error-prone pixels on high-resolution features without evaluating full uniform dense grids.

    1. RoI Feature Pyramid Extraction: Given an input image and detected instance bounding box proposals of width WW and height HH, Region of Interest (RoI) features are extracted across three levels {Pi,Pi−1,Pi−2}\{P_i, P_{i-1}, P_{i-2}\} of a Feature Pyramid Network (FPN) with increasing square spatial dimensions {28×28,56×56,112×112}\{28\times 28, 56\times 56, 112\times 112\}. The base pyramid level index ii is calculated as: i=⌊i0+log⁡2(WH224)⌋i = \left\lfloor i_0 + \log_2\left(\frac{\sqrt{WH}}{224}\right)\right\rfloor where i0=4i_0 = 4.

    2. Cascaded Incoherent Region Detection: A lightweight fully convolutional network (FCN) detects incoherent regions across scales:

    • At the coarsest RoI level (28×2828\times 28), the concatenated FPN features and initial coarse mask predictions are processed by four 3×33\times 3 convolutional layers followed by a binary classifier to predict root incoherent mask regions.
    • For finer levels (56×5656\times 56 and 112×112112\times 112), the predicted lower-resolution incoherent masks are upsampled and fused with neighboring finer FPN features using a single 1×11\times 1 convolution to guide finer-level incoherence predictions.
    1. Point Quadtree Decomposition: Each quadtree node corresponds to an individual feature point. Incoherent points detected at the coarsest level serve as root nodes. Each root node subdivides into four quadrant points at the adjacent finer resolution level. Only quadrant points that are independently classified as incoherent undergo further subdivision, strictly constraining node expansion to error-prone regions.
  3. Knowl 3 — Mask Transfiner Transformer Architecture: Node Encoder, Sequence Encoder, and Pixel Decoder

    model/method

    Because incoherent quadtree nodes are sparse and non-contiguously distributed across feature levels, Mask Transfiner replaces standard grid convolutions with a transformer refinement architecture consisting of three components:

    1. Node Encoder: For each incoherent quadtree node, the node encoder constructs a CC-dimensional feature representation fusing four distinct cues:
    • Fine-grained features: Sampled from the corresponding spatial coordinate and level of the RoI FPN feature pyramid.
    • Coarse mask prediction: The initial coarse probability value predicted by the base detector mask head.
    • Surrounding local context: Flattened features extracted from a 3×33\times 3 neighborhood around the node, compressed by a fully connected (FC) layer.
    • Relative positional encoding: Encodes the spatial (x,y)(x, y) coordinate within the bounding box RoI. The fine features, coarse mask cue, and local context features are concatenated and projected by an FC layer to dimension CC, and the positional embedding is added elementwise.
    1. Sequence Encoder: All encoded incoherent nodes across all quadtree levels are flattened into an unordered query sequence of shape C×NC \times N, where N≪H×WN \ll H \times W is the total number of incoherent nodes across all levels. To supply dense positive and negative references, all feature points from a coarse 14×1414\times 14 RoI grid are appended. The sequence is processed by 3 transformer encoder layers (each containing 4-head multi-head self-attention and a feed-forward network FFN), performing non-local spatial and cross-scale context reasoning in parallel.

    2. Pixel Decoder: A two-layer multi-layer perceptron (MLP) operates on each attended node representation to output the final refined binary segmentation probability.

  4. Knowl 4 — Coarse-to-Fine Quadtree Mask Propagation Algorithm

    algorithm

    After the transformer refines the segmentation probabilities of all sparse incoherent quadtree nodes, Mask Transfiner reconstructs the full-resolution output mask via hierarchical coarse-to-fine propagation.

    Input: Initial low-resolution coarse mask McoarseM_{\text{coarse}} at root resolution, Quadtree levels l∈{1,…,L}l \in \{1, \dots, L\} where l=1l=1 is coarsest root and l=Ll=L is finest, Refined label probabilities y^v∈[0,1]\hat{y}_v \in [0, 1] for all incoherent nodes v∈Vincv \in \mathcal{V}_{\text{inc}}
    Output: High-resolution instance mask MfinalM_{\text{final}} at scale LL
    Initialize mask at root level: M(1)←McoarseM^{(1)} \leftarrow M_{\text{coarse}}
    for each root incoherent node u∈Vinc(1)u \in \mathcal{V}_{\text{inc}}^{(1)} at level 1:
        M(1)(u)←y^uM^{(1)}(u) \leftarrow \hat{y}_u
    for l=2l = 2 to LL:
        Upsample previous level mask: M(l)←NearestNeighborInterpolate2×(M(l−1))M^{(l)} \leftarrow \text{NearestNeighborInterpolate}_{2\times}(M^{(l-1)})
        for each incoherent node v∈Vinc(l)v \in \mathcal{V}_{\text{inc}}^{(l)} at level ll:
            M(l)(v)←y^vM^{(l)}(v) \leftarrow \hat{y}_v
    Mfinal←M(L)M_{\text{final}} \leftarrow M^{(L)}
    return MfinalM_{\text{final}}

    By initializing each finer level via nearest-neighbor upsampling of the previously corrected mask, coherent child nodes (which were not flagged as incoherent and thus not recomputed by the transformer) inherit the refined values of their parents, enlarging the effective refinement area across intermediate tree levels at negligible computational cost.

  5. Knowl 5 — Multi-Task Training Loss for Mask Transfiner

    equation

    Mask Transfiner is trained in an end-to-end multi-task manner using the following composite loss function:

    L=λ1LDetect+λ2LCoarse+λ3LRefine+λ4LInc\mathcal{L} = \lambda_1 \mathcal{L}_{\text{Detect}} + \lambda_2 \mathcal{L}_{\text{Coarse}} + \lambda_3 \mathcal{L}_{\text{Refine}} + \lambda_4 \mathcal{L}_{\text{Inc}}

    where:

    • LDetect\mathcal{L}_{\text{Detect}} is the base object detector loss (including bounding box classification and regression losses for Faster R-CNN or query-based DETR architectures).
    • LCoarse\mathcal{L}_{\text{Coarse}} is the binary cross-entropy loss on the base network's initial coarse mask head prediction.
    • LRefine\mathcal{L}_{\text{Refine}} is the L1L_1 loss between predicted probabilities and binary ground-truth labels for all incoherent quadtree nodes.
    • LInc\mathcal{L}_{\text{Inc}} is the binary cross-entropy loss for the multi-level incoherent region detector.
    • λ1,λ2,λ3,λ4\lambda_1, \lambda_2, \lambda_3, \lambda_4 are loss balancing hyperparameters set to λ1=1.0\lambda_1 = 1.0, λ2=1.0\lambda_2 = 1.0, λ3=1.0\lambda_3 = 1.0, and λ4=0.5\lambda_4 = 0.5.
  6. Knowl 6 — Error Concentration and Oracle Performance of Incoherent Regions on COCO

    empirical result

    An empirical analysis on the COCO validation set demonstrates that mask prediction errors are overwhelmingly concentrated in spatially sparse incoherent regions.

    Bounding Box Area (%) Error Recall (RecallErr\text{Recall}_{\text{Err}}) Coarse Accuracy APCoarse\text{AP}_{\text{Coarse}} APGT\text{AP}_{\text{GT}} (Oracle)
    14% 43% 56% 35.5 51.0

    Key findings include:

    • Incoherent regions occupy only 14% of the total bounding box area, yet they encompass 43% of all wrongly predicted pixels per object (RecallErr\text{Recall}_{\text{Err}}).
    • Inside incoherent regions, the accuracy of standard coarse mask heads is only 56%.
    • In an oracle experiment fixing the bounding box detections and replacing coarse predictions inside incoherent regions with ground-truth labels (leaving all other pixels unchanged), average precision (AP\text{AP}) increases by +15.5 points from 35.5 to 51.0, demonstrating that selective refinement of incoherent regions is sufficient to achieve high mask fidelity.
  7. Knowl 7 — Instance Segmentation Performance on COCO Benchmark

    data/table

    Mask Transfiner was evaluated on COCO test-dev and COCO 2017 validation sets across two-stage (Faster R-CNN) and query-based detection frameworks. Metrics include standard mask AP\text{AP}, boundary quality APB\text{AP}^B (Boundary IoU), and high-quality LVIS annotation metric AP⋆\text{AP}^\star.

    Method Backbone Type AP\text{AP} APval⋆\text{AP}^\star_{\text{val}} APvalB\text{AP}^B_{\text{val}} APBox\text{AP}_{\text{Box}} APS\text{AP}_S APL\text{AP}_L
    Mask R-CNN R50-FPN Two-stage 37.5 38.2 21.2 41.3 21.1 48.3
    PointRend R50-FPN Two-stage 38.1 39.7 23.5 41.5 18.8 49.4
    BMask R-CNN R50-FPN Two-stage 37.8 39.8 23.5 41.6 19.7 49.6
    BPR R50-FPN Two-stage 38.4 40.2 24.3 41.3 20.2 49.7
    Mask Transfiner R50-FPN Two-stage 39.4 42.3 26.0 41.8 22.3 50.2
    Mask Transfiner†^{\dagger} R50-FPN Two-stage 40.5 43.1 26.8 43.2 22.8 52.5
    Mask R-CNN R101-FPN Two-stage 38.8 39.3 23.1 43.1 21.8 50.5
    PointRend R101-FPN Two-stage 39.6 41.4 25.3 43.3 19.8 53.7
    HTC R101-FPN Two-stage 39.7 42.5 25.4 45.9 21.0 53.5
    RefineMask R101-FPN Two-stage 39.4 42.3 26.8 43.8 21.6 53.1
    BCNet R101-FPN Two-stage 39.8 41.9 26.1 43.5 22.7 51.1
    Mask Transfiner R101-FPN Two-stage 40.7 43.6 27.3 43.9 23.1 53.8
    Mask Transfiner†^{\dagger} R101-FPN Two-stage 42.2 45.0 28.6 45.8 24.1 55.4
    ISTR R50-FPN Query 38.6 39.5 23.0 46.8 22.1 50.6
    QueryInst R50-FPN Query 39.9 42.1 25.1 44.5 22.9 51.9
    SOLQ R50-FPN Query 39.7 39.8 23.3 47.8 21.5 53.1
    Mask Transfiner R50-FPN Query 41.6 45.4 28.2 46.5 24.2 55.2

    (†\dagger denotes training with Deformable Convolutional Networks / DCN).

    Key comparisons demonstrate:

    • With ResNet-50-FPN in a two-stage setting, Mask Transfiner achieves 39.4 AP (+1.9 over Mask R-CNN) and 26.0 APB\text{AP}^B (+4.8 over Mask R-CNN).
    • With ResNet-50-FPN in a query-based setting, Mask Transfiner achieves 41.6 AP and 28.2 APB\text{AP}^B, outperforming QueryInst (39.9 AP) and SOLQ (39.7 AP) by +1.7 and +1.9 AP.
  8. Knowl 8 — High-Resolution Instance Segmentation on Cityscapes and BDD100K Benchmarks

    data/table

    Mask Transfiner was evaluated on urban scene benchmarks containing high-resolution images (2048×10242048 \times 1024 on Cityscapes) and complex object boundaries.

    Cityscapes Validation Set (ResNet-50-FPN, Two-Stage):

    Method APB\text{AP}^B AP50B\text{AP}^B_{50} AP\text{AP} AP50\text{AP}_{50}
    Mask R-CNN Baseline 11.4 37.4 33.8 61.5
    PointRend 16.7 47.2 35.9 61.8
    BMask R-CNN 15.7 46.2 36.2 62.6
    Panoptic-DeepLab 16.5 47.7 35.3 57.9
    RefineMask 17.4 49.2 37.6 63.3
    Mask Transfiner (Ours) 18.0 49.8 37.9 64.1

    On Cityscapes, Mask Transfiner achieves 37.9 mask AP\text{AP} and 18.0 boundary APB\text{AP}^B, outperforming the Mask R-CNN baseline by +6.6 APB\text{AP}^B and exceeding RefineMask and PointRend.

    BDD100K Validation Set:

    Method Backbone APmask\text{AP}_{\text{mask}} APbox\text{AP}_{\text{box}}
    Mask R-CNN Baseline R101-FPN 20.5 26.1
    Cascade Mask R-CNN R101-FPN 19.8 24.7
    Mask R-CNN + DCNv2 R101-FPN 20.9 26.0
    HRNet HRNet-w32 22.5 28.2
    Mask Transfiner (Ours) R101-FPN 23.6 26.2

    On BDD100K, Mask Transfiner achieves 23.6 APmask\text{AP}_{\text{mask}}, surpassing the Mask R-CNN baseline by +3.1 AP at comparable bounding box accuracy.

  9. Knowl 9 — Ablation Analysis of Node Encoding, Transformer Efficiency, and Quadtree Dynamics

    empirical result

    Ablation experiments on the COCO validation set using ResNet-50-FPN evaluate individual architectural choices in Mask Transfiner:

    1. Node Encoding Cues:
    • Fine-grained FPN features only: 33.8 AP, 20.1 APB\text{AP}^B, 37.0 AP⋆\text{AP}^\star.
    • Adding coarse mask prediction cues: 34.2 AP, 20.4 APB\text{AP}^B, 37.3 AP⋆\text{AP}^\star (+0.4 AP).
    • Adding relative positional encodings: 36.8 AP, 23.9 extAPB ext{AP}^B, 40.1 AP⋆\text{AP}^\star (+2.6 AP, +3.5 APB\text{AP}^B), demonstrating that positional signals are critical because transformers are permutation-invariant while segmentation is spatial.
    • Adding 3×33\times 3 local context features: 37.3 AP, 24.2 APB\text{AP}^B, 40.5 AP⋆\text{AP}^\star (+0.5 AP).
    1. Refinement Module Architecture and Resource Efficiency:
    • MLP refinement on incoherent regions: 36.4 AP, 23.7 APB\text{AP}^B.
    • Non-Local Attention (NLA) on 112×112112\times 112 grid: 36.3 AP, 24.6 GFLOPs, 8347 MB memory, 4.6 FPS.
    • Standard Transformer on full 56×5656\times 56 grid: 36.5 AP, 68.3 GFLOPs, 17359 MB memory, 2.1 FPS (standard transformer runs out of memory at 112×112112\times 112).
    • Mask Transfiner (112×112112\times 112 quadtree): 37.3 AP, 16.8 GFLOPs, 2316 MB memory, 7.1 FPS. Mask Transfiner achieves higher accuracy with over 3.6×3.6\times less memory than NLA and under one-fourth the compute of a standard grid transformer.
    1. Multi-Level Joint Refinement and Quadtree Mask Propagation:
    • Jointly feeding all 3 quadtree levels into a single transformer sequence outperforms separately refining each level with multiple sequences by +0.6 AP⋆\text{AP}^\star.
    • Hierarchical mask propagation from coarse to fine intermediate levels improves AP from 36.5 to 37.0 over only writing refined labels to the finest leaf nodes.
  10. Knowl 10 — Limitation: Dependence on Fully Supervised Pixel-Level Instance Annotations

    limitation

    Mask Transfiner requires fully supervised training with high-quality instance mask annotations. Ground-truth incoherent region detection targets, coarse mask head supervision, and point refinement losses all depend directly on dense pixel-level ground truth during training.

Coverage note — None. All substantial contributions—including the formulation of incoherent regions, point quadtree construction, transformer refinement architecture, quadtree propagation algorithm, loss formulation, empirical oracle analyses, benchmark comparisons, and ablation studies—have been represented.

References

  1. 1.Daniel Bolya, Chong Zhou, Fanyi Xiao, and Yong Jae Lee. Yolact: real-time instance segmentation. In ICCV, 2019. 2
  2. 2.Zhaowei Cai and Nuno Vasconcelos. Cascade r-cnn: Delving into high quality object detection. In CVPR, 2018. 2
  3. 3.Zhaowei Cai and Nuno Vasconcelos. Cascade r-cnn: High quality object detection and instance segmentation. 2019. 8
  4. 4.Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. End-to-end object detection with transformers. In ECCV, 2020. 2, 5, 6, 7, 8
  5. 5.Hao Chen, Kunyang Sun, Zhi Tian, Chunhua Shen, Yongming Huang, and Youliang Yan. BlendMask: Top-down meets bottom-up for instance segmentation. In CVPR, 2020. 2
  6. 6.Kai Chen, Jiangmiao Pang, Jiaqi Wang, Yu Xiong, Xiaoxiao Li, Shuyang Sun, Wansen Feng, Ziwei Liu, Jianping Shi, Wanli Ouyang, et al. Hybrid task cascade for instance segmentation. In CVPR, 2019. 2, 8
  7. 7.Liang-Chieh Chen, Alexander Hermans, George Papandreou, Florian Schroff, Peng Wang, and Hartwig Adam. Masklab: Instance segmentation by refining object detection with semantic and direction features. In CVPR, 2018. 2
  8. 8.Xinlei Chen, Ross Girshick, Kaiming He, and Piotr Doll'ar. Tensormask: A foundation for dense object segmentation. In ICCV, 2019. 2
  9. 9.Bowen Cheng, Maxwell D Collins, Yukun Zhu, Ting Liu, Thomas S Huang, Hartwig Adam, and Liang-Chieh Chen. Panoptic-deeplab: A simple, strong, and fast baseline for bottom-up panoptic segmentation. In CVPR, 2020. 8
  10. 10.Bowen Cheng, Ross Girshick, Piotr Doll'ar, Alexander C Berg, and Alexander Kirillov. Boundary iou: Improving object-centric image segmentation evaluation. In CVPR, 2021. 6, 8
  11. 11.Ho Kei Cheng, Jihoon Chung, Yu-Wing Tai, and Chi-Keung Tang. Cascadepsp: toward class-agnostic and very high-resolution segmentation via global and local refinement. In CVPR, 2020. 2
  12. 12.Tianheng Cheng, Xinggang Wang, Lichao Huang, and Wenyu Liu. Boundary-preserving mask r-cnn. In ECCV, 2020. 1, 2, 8
  13. 13.Marius Cordts, Mohamed Omran, Sebastian Ramos, Timo Rehfeld, Markus Enzweiler, Rodrigo Benenson, Uwe Franke, Stefan Roth, and Bernt Schiele. The cityscapes dataset for semantic urban scene understanding. In CVPR, 2016. 6
  14. 14.Henghui Ding, Xudong Jiang, Ai Qun Liu, Nadia Magnenat Thalmann, and Gang Wang. Boundary-aware feature propagation for scene segmentation. In ICCV, 2019. 2
  15. 15.Bin Dong, Fangao Zeng, Tiancai Wang, Xiangyu Zhang, and Yichen Wei. Solq: Segmenting objects by learning queries. In NeurIPS, 2021. 1, 2, 3, 4, 8
  16. 16.Qi Fan, Lei Ke, Wenjie Pei, Chi-Keung Tang, and Yu-Wing Tai. Commonality-parsing network across shape and appearance for partially supervised instance segmentation. In ECCV, 2020. 2
  17. 17.Yuxin Fang, Shusheng Yang, Xinggang Wang, Yu Li, Chen Fang, Ying Shan, Bin Feng, and Wenyu Liu. Instances as queries. In ICCV, 2021. 1, 2, 8
  18. 18.Raphael A Finkel and Jon Louis Bentley. Quad trees a data structure for retrieval on composite keys. Acta informatica, 4(1):1–9, 1974. 1
  19. 19.Ruohao Guo, Dantong Niu, Liao Qu, and Zhenbo Li. Sotr: Segmenting objects with transformers. In ICCV, 2021. 2
  20. 20.Agrim Gupta, Piotr Dollar, and Ross Girshick. Lvis: A dataset for large vocabulary instance segmentation. In CVPR, 2019. 6, 8
  21. 21.Kaiming He, Georgia Gkioxari, Piotr Doll'ar, and Ross Girshick. Mask r-cnn. In ICCV, 2017. 1, 2, 3, 4, 5, 7, 8
  22. 22.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In CVPR, 2016. 6, 7, 8
  23. 23.Jie Hu, Liujuan Cao, Yao Lu, ShengChuan Zhang, Ke Li, Feiyue Huang, Ling Shao, and Rongrong Ji. Istr: End-to-end instance segmentation via transformers. arXiv preprint arXiv:2105.00637, 2021. 1, 2, 3, 8
  24. 24.Zhaojin Huang, Lichao Huang, Yongchao Gong, Chang Huang, and Xinggang Wang. Mask scoring r-cnn. In CVPR, 2019. 1, 2, 8
  25. 25.Lei Ke, Xia Li, Martin Danelljan, Yu-Wing Tai, Chi-Keung Tang, and Fisher Yu. Prototypical cross-attention networks for multiple object tracking and segmentation. In Advances in Neural Information Processing Systems, 2021. 2
  26. 26.Lei Ke, Yu-Wing Tai, and Chi-Keung Tang. Deep occlusion-aware instance segmentation with overlapping bilayers. In CVPR, 2021. 2, 6, 8
  27. 27.Lei Ke, Yu-Wing Tai, and Chi-Keung Tang. Occlusion-aware video object inpainting. In ICCV, 2021. 2
  28. 28.Alexander Kirillov, Yuxin Wu, Kaiming He, and Ross Girshick. Pointrend: Image segmentation as rendering. In CVPR, 2020. 1, 2, 4, 6, 7, 8
  29. 29.Weicheng Kuo, Anelia Angelova, Jitendra Malik, and Tsung-Yi Lin. Shapemask: Learning to segment novel objects by refining shape priors. In ICCV, 2019. 2
  30. 30.Youngwan Lee and Jongyoul Park. Centermask: Real-time anchor-free instance segmentation. In CVPR, 2020. 2, 6
  31. 31.Yi Li, Haozhi Qi, Jifeng Dai, Xiangyang Ji, and Yichen Wei. Fully convolutional instance-aware semantic segmentation. In CVPR, 2017. 2
  32. 32.Tsung-Yi Lin, Piotr Doll'ar, Ross Girshick, Kaiming He, Bharath Hariharan, and Serge Belongie. Feature pyramid networks for object detection. In CVPR, 2017. 4
  33. 33.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll'ar, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In ECCV, 2014. 1, 2, 6
  34. 34.Shu Liu, Lu Qi, Haifang Qin, Jianping Shi, and Jiaya Jia. Path aggregation network for instance segmentation. In CVPR, 2018. 1
  35. 35.Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. Faster r-cnn: Towards real-time object detection with region proposal networks. In NeurIPS, 2015. 2, 5, 6
  36. 36.Chufeng Tang, Hang Chen, Xiao Li, Jianmin Li, Zhaoxiang Zhang, and Xiaolin Hu. Look closer to segment better: Boundary patch refinement for instance segmentation. 2021. 2, 8
  37. 37.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In NeurIPS, 2017. 7
  38. 38.Jingdong Wang, Ke Sun, Tianheng Cheng, Borui Jiang, Chaorui Deng, Yang Zhao, Dong Liu, Yadong Mu, Mingkui Tan, Xinggang Wang, et al. Deep high-resolution representation learning for visual recognition. TPAMI, 2020. 1, 2, 8
  39. 39.Xiaolong Wang, Ross Girshick, Abhinav Gupta, and Kaiming He. Non-local neural networks. In CVPR, 2018. 7
  40. 40.Xinlong Wang, Tao Kong, Chunhua Shen, Yuning Jiang, and Lei Li. Solo: Segmenting objects by locations. arXiv preprint arXiv:1912.04488, 2019. 2
  41. 41.Xinlong Wang, Rufeng Zhang, Tao Kong, Lei Li, and Chunhua Shen. Solov2: Dynamic and fast instance segmentation. In NeurIPS, 2020. 2
  42. 42.Yuqing Wang, Zhaoliang Xu, Xinlong Wang, Chunhua Shen, Baoshan Cheng, Hao Shen, and Huaxia Xia. End-to-end video instance segmentation with transformers. In CVPR, 2021. 2
  43. 43.Yuxin Wu, Alexander Kirillov, Francisco Massa, Wan-Yen Lo, and Ross Girshick. Detectron2. https://github.com/facebookresearch/detectron2, 2019. 6
  44. 44.Enze Xie, Peize Sun, Xiaoge Song, Wenhai Wang, Xuebo Liu, Ding Liang, Chunhua Shen, and Ping Luo. Polarmask: Single shot instance segmentation with polar representation. In CVPR, 2020. 2
  45. 45.Fisher Yu, Haofeng Chen, Xin Wang, Wenqi Xian, Yingying Chen, Fangchen Liu, Vashisht Madhavan, and Trevor Darrell. Bdd100k: A diverse driving dataset for heterogeneous multitask learning. In CVPR, 2020. 6
  46. 46.Yuhui Yuan, Jingyi Xie, Xilin Chen, and Jingdong Wang. Segfix: Model-agnostic boundary refinement for segmentation. In ECCV, 2020. 2
  47. 47.Gang Zhang, Xin Lu, Jingru Tan, Jianmin Li, Zhaoxiang Zhang, Quanquan Li, and Xiaolin Hu. Refinemask: Towards high-quality instance segmentation with fine-grained features. In CVPR, 2021. 2, 8
  48. 48.Wenwei Zhang, Jiangmiao Pang, Kai Chen, and Chen Change Loy. K-net: Towards unified image segmentation. In NeurIPS, 2021. 2
  49. 49.Xizhou Zhu, Han Hu, Stephen Lin, and Jifeng Dai. Deformable convnets v2: More deformable, better results. In CVPR, 2019. 8

Citation

MLA
Ke, L., et al. “Mask Transfiner for High-Quality Instance Segmentation”. arXiv, 2021, http://arxiv.org/abs/2111.13673v1.
APA
Ke, L., Danelljan, M., Li, X., Tai, Y.-W., Tang, C.-K., & Yu, F. (2021). Mask Transfiner for High-Quality Instance Segmentation. arXiv. http://arxiv.org/abs/2111.13673v1
Chicago
Ke, L., M. Danelljan, X. Li, Y.-W. Tai, C.-K. Tang, and F. Yu. 2021. “Mask Transfiner for High-Quality Instance Segmentation”. arXiv. http://arxiv.org/abs/2111.13673v1.
Harvard
Ke, L. et al. (2021) “Mask Transfiner for High-Quality Instance Segmentation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2111.13673v1.
Vancouver
1. Ke L, Danelljan M, Li X, Tai Y-W, Tang C-K, Yu F (2021) Mask Transfiner for High-Quality Instance Segmentation. arXiv

BibTeX

@article{ke2021mask,
  title = {Mask Transfiner for High-Quality Instance Segmentation},
  author = {Ke, Lei and Danelljan, Martin and Li, Xia and Tai, Yu-Wing and Tang, Chi-Keung and Yu, Fisher},
  year = {2021},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2111.13673v1},
  eprint = {2111.13673}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE