Panoptic Feature Pyramid Networks

Alexander KirillovRoss B. GirshickKaiming HePiotr Dollár

article2019CVPR1,542 citations

Presents Panoptic FPN, a simple yet highly effective framework that unifies instance and semantic segmentation into a single architecture by adding a lightweight semantic segmentation branch to Mask R-CNN over a shared Feature Pyramid Network backbone.

Listen

Modern computer vision systems rely heavily on two core visual recognition tasks: instance segmentation, which detects and segments individual countable objects such as cars and pedestrians, and semantic segmentation, which assigns pixel-level category labels to uncountable background regions such as sky and roads. Although real-world applications require understanding both elements simultaneously—a unified objective known as panoptic segmentation—leading methods traditionally execute these tasks using separate, disconnected neural networks. Operating separate pipelines doubles computational expense and memory overhead, complicating deployment in resource-constrained environments.

The article demonstrates that a single, unified deep learning architecture can simultaneously perform both instance and semantic segmentation at state-of-the-art accuracy levels without duplicating computational workloads. To evaluate this, the authors introduce Panoptic Feature Pyramid Networks (Panoptic FPN), a framework that integrates dense pixel prediction directly into an established object detection architecture.

The approach builds upon the standard Mask R-CNN object detector with a Feature Pyramid Network (FPN) backbone. While retaining the original region-based branch for identifying individual objects, the authors attach a parallel, lightweight dense-prediction branch designed to extract background semantic labels directly from multi-scale feature maps. The unified network is evaluated through rigorous multi-task training experiments across two prominent computer vision benchmarks: the large-scale COCO dataset and the urban street-scene Cityscapes dataset.

The key findings show that a single Panoptic FPN matches the accuracy of two independent, specialized networks while cutting total computational demand by approximately 50%. When evaluated under an identical computational budget, a single deeper Panoptic FPN outperforms two separate networks across all primary metrics. Additionally, the lightweight semantic segmentation branch independently matches top dilation-based systems while using roughly half the memory activations and computation. On the COCO panoptic leaderboard, the single-network model outperformed competing unified methods by approximately 9 points in panoptic quality, and exceeded existing alternatives on Cityscapes by 4.3 points.

These results demonstrate that organizations can deploy unified visual recognition systems that lower compute costs, reduce hardware memory footprint, and simplify maintenance pipelines without compromising visual precision. Furthermore, the findings challenge the long-standing assumption that high-accuracy semantic segmentation requires computationally intensive dilated convolutions or complex symmetric decoders.

Teams designing computer vision systems should adopt unified multi-task architectures like Panoptic FPN as their baseline rather than maintaining separate pipelines. When training joint models, practitioners must balance task loss weights and merge training losses per batch rather than alternating tasks. Future development should explore more advanced multi-task feature sharing strategies and investigate integrating complementary architectural enhancements to further boost joint performance.

  • Paper: Panoptic Segmentation, Alexander Kirillov et al. (2018). This foundational paper introduces and defines the panoptic segmentation task and its evaluation metric, providing the core problem formulation that Panoptic FPN aims to solve.
  • Paper: Mask R-CNN, Kaiming He et al. (2017). Panoptic FPN directly builds upon Mask R-CNN as its instance segmentation engine, making an understanding of its ROI-based architecture and RoIAlign essential.
  • Paper: Feature Pyramid Networks for Object Detection, Tsung-Yi Lin et al. (2017). This paper establishes Feature Pyramid Networks, the multi-scale feature backbone that Panoptic FPN augments with a dense semantic segmentation head.
  • Paper: Path Aggregation Network for Instance Segmentation, Shu Liu et al. (2018). It enhances multi-scale feature aggregation in FPN-based architectures, influencing the design of dense feature fusion heads for dense prediction tasks.
  • Paper: Fully convolutional networks for semantic segmentation, Jonathan Long et al. (2015). It provides the foundational framework for dense per-pixel semantic segmentation via fully convolutional networks adapted by Panoptic FPN's semantic branch.
  • Paper: Unified Perceptual Parsing for Scene Understanding, Tete Xiao et al. (2018). This work explores multi-task visual understanding using Feature Pyramid Networks, serving as a direct precursor to shared-backbone parsing architectures.
Cover for Panoptic Feature Pyramid Networks

Abstract

The recently introduced panoptic segmentation task has renewed our community's interest in unifying the tasks of instance segmentation (for thing classes) and semantic segmentation (for stuff classes). However, current state-of-the-art methods for this joint task use separate and dissimilar networks for instance and semantic segmentation, without performing any shared computation. In this work, we aim to unify these methods at the architectural level, designing a single network for both tasks. Our approach is to endow Mask R-CNN, a popular instance segmentation method, with a semantic segmentation branch using a shared Feature Pyramid Network (FPN) backbone. Surprisingly, this simple baseline not only remains effective for instance segmentation, but also yields a lightweight, top-performing method for semantic segmentation. In this work, we perform a detailed study of this minimally extended version of Mask R-CNN with FPN, which we refer to as Panoptic FPN, and show it is a robust and accurate baseline for both tasks. Given its effectiveness and conceptual simplicity, we hope our method can serve as a strong baseline and aid future research in panoptic segmentation.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Panoptic Feature Pyramid Network
  • 3.1 Model Architecture
  • 3.2 Inference and Training
  • 3.3 Analysis
  • 4 Experiments
  • 4.1 Experimental Setup
  • 4.2 FPN for Semantic Segmentation
  • 4.3 Multi-Task Training
  • 4.4 Panoptic FPN
  • 5 Conclusion
  • References

Knowls

  1. Knowl 1 — Panoptic Feature Pyramid Network Architecture

    model/method

    Panoptic Feature Pyramid Network (Panoptic FPN) is a unified, single-network architecture for panoptic segmentation that simultaneously outputs instance segmentations for thing classes and dense semantic segmentations for stuff classes.

    The network comprises three main components:

    1. Shared Feature Pyramid Network (FPN) Backbone: A standard convolutional feedforward backbone (such as ResNet or ResNeXt) extracts multi-scale bottom-up features. A top-down pathway with lateral connections progressively upsamples higher-level semantic features and adds transformed lower-level features, generating a multi-scale feature pyramid spanning scales from 1/321/32 to 1/41/4 resolution (P5P_5 to P2P_2), each with a uniform channel dimension (typically 256 channels).
    2. Region-Based Instance Segmentation Branch: Identical to Mask R-CNN, this branch operates on Region of Interest (RoI) pooled features extracted from the various FPN levels. It applies shared convolutional and fully connected heads for bounding-box classification and regression, together with a Fully Convolutional Network (FCN) branch that predicts a binary segmentation mask for each candidate region.
    3. Dense Semantic Segmentation Branch: Attached directly to the same multi-scale FPN feature maps in parallel to the instance branch, this lightweight branch merges features from all pyramid levels (1/32,1/16,1/8,1/41/32, 1/16, 1/8, 1/4) into a single unified feature representation at 1/41/4 resolution, followed by a pixel-wise classification layer predicting class logits at the original image resolution.

    Both branches share the entire FPN feature backbone without modifying the backbone's design, enabling end-to-end multi-task training and inference.

  2. Knowl 2 — Semantic Segmentation Branch Architecture in Panoptic FPN

    model/method

    The semantic segmentation branch in Panoptic FPN generates dense pixel-wise semantic segmentations by merging multi-scale features from all pyramid levels of a Feature Pyramid Network (FPN) into a single feature representation at 1/41/4 image scale.

    For each FPN scale (1/32,1/16,1/8,1/41/32, 1/16, 1/8, 1/4, each having 256256 input channels):

    1. Upsampling Stages: Features at each scale are mapped through a sequence of identical upsampling blocks until they reach 1/41/4 scale:
      • The 1/321/32 scale level undergoes 33 successive stages (reaching 1/161/16, 1/81/8, then 1/41/4).
      • The 1/161/16 scale level undergoes 22 successive stages (reaching 1/81/8, then 1/41/4).
      • The 1/81/8 scale level undergoes 11 stage (reaching 1/41/4).
      • The 1/41/4 scale level undergoes a single 3×33 \times 3 convolution block without spatial upsampling. Each upsampling stage consists of a 3×33 \times 3 convolution reducing/maintaining channel dimension to 128128, Group Normalization (GN), ReLU non-linearity, and 2×2\times bilinear upsampling.
    2. Feature Aggregation: The resulting four feature maps, all now at 1/41/4 spatial resolution and with 128128 channels, are combined via element-wise summation.
    3. Final Classification: The summed feature map undergoes a final 1×11 \times 1 convolution to map from 128128 channels to CC class logits, followed by 4×4\times bilinear upsampling and a softmax function to generate per-pixel probability distributions at the original image resolution.

    For datasets distinguishing thing and stuff categories (such as COCO), the branch predicts all stuff classes plus an extra 'other' class assigned to all pixels belonging to object instances, preventing stuff class predictions on foreground thing regions.

  3. Knowl 3 — Joint Multi-Task Loss Formulation for Panoptic FPN

    equation

    The overall multi-task training objective for Panoptic FPN combines instance segmentation and semantic segmentation losses via task-specific balancing weights:

    L=λi(Lc+Lb+Lm)+λsLsL = \lambda_i (L_c + L_b + L_m) + \lambda_s L_s

    where:

    • LcL_c is the classification cross-entropy loss over RoI proposals in the instance branch, normalized by the number of sampled RoIs.
    • LbL_b is the bounding-box regression loss over positive RoI proposals, normalized by the number of sampled RoIs.
    • LmL_m is the binary mask cross-entropy loss for candidate instance regions, normalized by the number of foreground RoIs.
    • LsL_s is the per-pixel cross-entropy loss between predicted semantic class distributions and ground-truth pixel labels, normalized by the total number of labeled pixels in the image.
    • λi∈[0,∞)\lambda_i \in [0, \infty) is the scalar weight controlling the contribution of the instance segmentation branch.
    • λs∈[0,∞)\lambda_s \in [0, \infty) is the scalar weight controlling the contribution of the semantic segmentation branch.

    Because the loss magnitudes and normalization schemes of the region-based instance branch and dense semantic branch differ significantly, adjusting λi\lambda_i and λs\lambda_s (typically chosen from {0.5,0.75,1.0}\{0.5, 0.75, 1.0\}) prevents one task from dominating or degrading the performance of the other during joint gradient optimization.

  4. Knowl 4 — Panoptic Post-Processing and Output Merging Algorithm

    algorithm

    Panoptic segmentation requires assigning every pixel a single semantic class label (or void) and an instance ID (non-void for things, void for stuff). Panoptic FPN produces overlapping candidate instances from its Mask R-CNN branch and dense semantic predictions from its semantic branch. These outputs are resolved into a non-overlapping panoptic segmentation using a priority-based post-processing procedure.

    Input: Set of predicted instances I={(Mk,ck,sk)}k=1KI = \{(M_k, c_k, s_k)\}_{k=1}^K where MkM_k is a binary mask, ckc_k is a thing class label, and sk∈[0,1]s_k \in [0, 1] is a confidence score; dense semantic segmentation map S∈{1,…,C}(H×W)S \in \{1, \dots, C\}^{(H \times W)}; stuff area threshold τarea\tau_{\text{area}}.
    Output: Panoptic segmentation map P∈(ClassID×InstanceID)(H×W)P \in (\text{ClassID} \times \text{InstanceID})^{(H \times W)}.
    Initialize P(p)←(void,void)P(p) \leftarrow (\text{void}, \text{void}) for all pixels p∈H×Wp \in H \times W
    Sort instance predictions II in descending order of confidence score sks_k
    for each instance (Mk,ck,sk)∈I(M_k, c_k, s_k) \in I do
        for each pixel pp where Mk(p)=1M_k(p) = 1 do
            if P(p)=(void,void)P(p) = (\text{void}, \text{void}) then
                P(p)←(ck,k)P(p) \leftarrow (c_k, k)
            end if
        end for
    end for
    for each contiguous connected component segment UU of class cstuffc_{\text{stuff}} in SS do
        if cstuff="other"c_{\text{stuff}} = \text{"other"} or Area(U)<τarea\text{Area}(U) < \tau_{\text{area}} then
            continue
        end if
        for each pixel p∈Up \in U do
            if P(p)=(void,void)P(p) = (\text{void}, \text{void}) then
                P(p)←(cstuff,void)P(p) \leftarrow (c_{\text{stuff}}, \text{void})
            end if
        end for
    end for
    return PP

    This procedure resolves instance-instance overlaps by score priority, gives instance masks strict precedence over dense semantic stuff predictions, and filters out spurious stuff segments labeled 'other' or under the minimum area threshold τarea\tau_{\text{area}}.

  5. Knowl 5 — Computational and Memory Efficiency of FPN for Semantic Segmentation

    empirical result

    When evaluated on a 2-megapixel image with a ResNet-101 backbone, the computational cost (multiply-adds) and memory footprint (activation volume) of different feature resolution strategies compare as follows:

    1. Dilated Convolutions (Dilation-8 vs Dilation-16): Standard dilation-8 backbones, which maintain a 1/81/8 feature resolution by replacing stride-2 convolutions with dilated filters, require approximately 3×3\times more compute (∼6.0×1012\sim 6.0 \times 10^{12} mult-adds vs ∼1.9×1012\sim 1.9 \times 10^{12}) and significantly higher activation memory (∼6.0×109\sim 6.0 \times 10^{9} vs ∼1.9×109\sim 1.9 \times 10^{9} activations) compared to dilation-16.
    2. Symmetric Decoders (U-Net Style): Mirror-image decoders with lateral connections that upsample features to 1/41/4 resolution consume roughly 2×2\times the compute and memory of FPN.
    3. Feature Pyramid Network (FPN): FPN generates a 1/41/4 scale feature output with a computational overhead (∼0.5×1012\sim 0.5 \times 10^{12} mult-adds and ∼0.8×109\sim 0.8 \times 10^{9} activations) that is roughly equivalent to a low-resolution dilation-16 network (which only produces a 1/161/16 output resolution), while delivering a 4×4\times higher spatial resolution.

    Thus, FPN serves as a lightweight, asymmetric encoder-decoder that avoids the heavy memory and computational overhead of dilated convolutions while generating higher-resolution feature representations.

  6. Knowl 6 — Joint Panoptic FPN vs Separate Task-Specific Networks

    data/table

    Joint training with a single Panoptic FPN matches the accuracy of two independent, task-dedicated FPN models (one for instance segmentation, one for semantic segmentation) at roughly half the total computation. When given equal computational budgets, a single Panoptic FPN with a larger backbone significantly outperforms two separate networks.

    Evaluations were performed on the COCO and Cityscapes validation datasets using Panoptic Quality (PQPQ), Thing Panoptic Quality (PQThPQ^{\text{Th}}), Stuff Panoptic Quality (PQStPQ^{\text{St}}), Mask Average Precision (APAP), and Semantic Mean Intersection-over-Union (mIoUmIoU):

    Dataset Backbone AP PQTh\text{PQ}^{\text{Th}} mIoU PQSt\text{PQ}^{\text{St}} PQ
    COCO ResNet-50-FPN ×2\times 2 (Separate) 33.9 46.6 40.2 27.9 39.2
    COCO ResNet-50-FPN (Joint) 33.3 45.9 41.0 28.7 39.0
    COCO ResNet-101-FPN (Joint, Equal Budget) 35.2 47.5 42.1 29.5 40.3
    Cityscapes ResNet-50-FPN ×2\times 2 (Separate) 32.2 51.3 74.5 62.4 57.7
    Cityscapes ResNet-50-FPN (Joint) 32.0 51.6 75.0 62.2 57.7
    Cityscapes ResNet-101-FPN (Joint, Equal Budget) 33.0 52.0 75.7 62.5 58.1

    On COCO, a single ResNet-50-FPN achieves 39.039.0 PQ compared to 39.239.2 PQ for two separate ResNet-50 models (a difference of only −0.2-0.2 PQ with nearly 50% fewer FLOPs). Under an equal compute budget, a single ResNet-101 Panoptic FPN outperforms two ResNet-50 networks by +1.1+1.1 PQ on COCO and +0.4+0.4 PQ on Cityscapes.

  7. Knowl 7 — Semantic FPN Single-Task Segmentation Performance

    data/table

    When trained purely for semantic segmentation (termed Semantic FPN), attaching the lightweight dense prediction branch to an FPN backbone achieves accuracy competitive with state-of-the-art dilation-based architectures on Cityscapes validation and COCO-Stuff benchmarks, while requiring fewer FLOPs and less activation memory.

    On Cityscapes validation (trained only on fine annotations; FLOPs in multiply-adds ×1012\times 10^{12}, memory in activations ×109\times 10^9 on 2 MP2\,\text{MP} images):

    Method Backbone mIoU FLOPs Memory
    DeepLabV3 ResNet-101-D8 77.8 1.9 1.9
    PSANet101 ResNet-101-D8 77.9 2.0 2.0
    Mapillary WideResNet-38-D8 79.4 4.3 1.7
    DeepLabV3+ Xception-71-D16 79.6 0.5 1.9
    Semantic FPN ResNet-101-FPN 77.7 0.5 0.8
    Semantic FPN ResNeXt-101-FPN 79.1 0.8 1.4

    On the COCO-Stuff 2017 Challenge:

    Entry / Team Backbone mIoU fIoU
    Vllab Stacked Hourglass 12.4 38.8
    DeepLab VGG16 VGG-16 20.2 47.5
    Oxford ResNeXt-101 24.1 50.6
    G-RMI Inception ResNet v2 26.6 51.9
    Semantic FPN ResNeXt-152-FPN 28.8 55.7

    Semantic FPN with ResNeXt-152 won the 2017 COCO-Stuff Segmentation Challenge without ensembling, outperforming all competitors by at least 2.22.2 points in mIoU and 3.83.8 points in fIoU.

  8. Knowl 8 — Effects of Multi-Task Loss Weighting on Individual Task Performance

    empirical result

    In multi-task training of Panoptic FPN (ResNet-50-FPN), setting the relative loss weights λs\lambda_s (semantic loss weight) and λi\lambda_i (instance loss weight) allows one task to serve as beneficial auxiliary supervision for the other, outperforming single-task baselines:

    1. Instance Segmentation with Auxiliary Semantic Loss (λi=1\lambda_i = 1):
      • On COCO, setting λs=0.1\lambda_s = 0.1 improves mask AP from 33.933.9 (single-task baseline λs=0.0\lambda_s = 0.0) to 34.034.0, and thing panoptic quality PQThPQ^{\text{Th}} from 46.646.6 to 46.846.8. Higher semantic loss weights (e.g., λs=1.0\lambda_s = 1.0) degrade mask AP to 32.132.1.
      • On Cityscapes, setting λs=1.0\lambda_s = 1.0 improves mask AP from 32.232.2 (λs=0.0\lambda_s = 0.0) to 33.233.2 (+1.0 AP) and PQThPQ^{\text{Th}} from 51.351.3 to 52.452.4 (+1.1).
    2. Semantic Segmentation with Auxiliary Instance Loss (λs=1\lambda_s = 1):
      • On COCO, setting λi=1.0\lambda_i = 1.0 improves mIoU from 40.240.2 (single-task baseline λi=0.0\lambda_i = 0.0) to 41.541.5 (+1.3 mIoU), frequency-weighted IoU (fIoU) from 67.267.2 to 68.268.2, and stuff panoptic quality PQStPQ^{\text{St}} from 27.927.9 to 29.029.0.
      • On Cityscapes, setting λi=0.25\lambda_i = 0.25 improves mIoU from 74.574.5 (λi=0.0\lambda_i = 0.0) to 75.575.5 (+1.0 mIoU) and instance-level IoU (iIoU) from 55.855.8 to 58.358.3 (+2.5 iIoU).

    Unweighted loss summation (λi=λs=1.0\lambda_i = \lambda_s = 1.0) is suboptimal when optimizing purely for one task, but properly tuned loss weights allow joint training to act as a regularizer that enhances single-task representation learning.

  9. Knowl 9 — Panoptic FPN Benchmark Results on COCO and Cityscapes

    data/table

    Panoptic FPN using a single ResNet-101 backbone achieves state-of-the-art panoptic segmentation performance on the COCO test-dev benchmark among single-network entries, and outperforms competing methods on Cityscapes validation without ensembling or extra coarse training annotations.

    Comparison on COCO test-dev (single-network entries without ensembling):

    Method PQ PQTh\text{PQ}^{\text{Th}} PQSt\text{PQ}^{\text{St}}
    Artemis 16.9 16.8 17.0
    LeChen 26.2 31.0 18.9
    MPS-TU Eindhoven 27.2 29.6 23.4
    MMAP-seg 32.1 38.9 22.0
    Panoptic FPN (ResNet-101) 40.9 48.3 29.7

    Panoptic FPN outperforms all competing single-model systems on COCO test-dev by an 8.8-point PQ margin (40.940.9 vs 32.132.1).

    Comparison on Cityscapes validation:

    Method Extra Coarse PQ PQTh\text{PQ}^{\text{Th}} PQSt\text{PQ}^{\text{St}} mIoU AP
    DIN Yes 53.8 42.5 62.1 80.1 28.6
    Panoptic FPN (ResNet-101) No 58.1 52.0 62.5 75.7 33.0

    On Cityscapes validation, Panoptic FPN outperforms the pixel-grouping method DIN by 4.34.3 points in PQ (58.158.1 vs 53.853.8) and by 9.59.5 points in PQThPQ^{\text{Th}} (52.052.0 vs 42.542.5), despite not using the 20,000 coarsely annotated training images used by DIN.

  10. Knowl 10 — Ablation Analysis of Semantic Branch and Multi-Task Optimization Strategies

    empirical result

    Ablation studies on Panoptic FPN (ResNet-50 backbone) establish design guidelines for the semantic branch architecture and joint training:

    1. Semantic Feature Channel Width:
      • On Cityscapes val / COCO val, channel widths of 64, 128, and 256 achieve semantic mIoU of 74.1/39.674.1 / 39.6, 74.5/40.274.5 / 40.2, and 74.6/40.174.6 / 40.1 respectively. Setting the channel dimension to 128 balances representation power and parameter efficiency.
    2. Feature Aggregation Method:
      • Element-wise summation of multi-scale 1/41/4 feature maps yields 74.574.5 mIoU on Cityscapes and 40.240.2 on COCO, compared to 74.474.4 and 39.939.9 for channel concatenation. Summation is both slightly higher in accuracy and computationally more efficient.
    3. Loss Gradient Scheduling (Combined vs Alternating):
      • Computing combined instance and semantic losses simultaneously in each minibatch achieves 39.039.0 PQ on COCO (33.333.3 AP, 41.041.0 mIoU) and 57.757.7 PQ on Cityscapes (32.032.0 AP, 75.075.0 mIoU).
      • Alternating instance and semantic loss updates across successive iterations (trained for 2×2\times total iterations) degrades performance to 37.537.5 PQ on COCO (31.731.7 AP, 40.240.2 mIoU) and 57.457.4 PQ on Cityscapes (32.032.0 AP, 74.374.3 mIoU).
    4. Shared vs Grouped FPN Channels:
      • Splitting the 256 FPN channels into two dedicated groups of 128 (one for instance segmentation and one for semantic segmentation) results in 38.838.8 PQ on COCO and 57.557.5 PQ on Cityscapes, slightly trailing the fully shared 256-channel configuration (39.039.0 PQ and 57.757.7 PQ, respectively).

Coverage note — None was omitted; all primary architectural innovations, loss formulations, post-processing merging algorithms, efficiency evaluations, multi-task interaction analyses, and core benchmark tables have been captured.

References

  1. 1.A. Arnab and P. H. Torr. Pixelwise instance segmentation with a dynamically instantiated network. In CVPR, 2017. 1, 3, 8
  2. 2.V. Badrinarayanan, A. Kendall, and R. Cipolla. Segnet: A deep convolutional encoder-decoder architecture for image segmentation. PAMI, 2017. 3
  3. 3.S. Bell, C. Lawrence Zitnick, K. Bala, and R. Girshick. Inside-outside net: Detecting objects in context with skip pooling and recurrent neural networks. In CVPR, 2016. 3
  4. 4.P. Bilinski and V. Prisacariu. COCO-Stuff 2017 Challenge: Oxford Active Vision Lab team. 2017. 6
  5. 5.S. R. Bulò, L. Porzi, and P. Kontschieder. In-place activated batchnorm for memory-optimized training of DNNs. In CVPR, 2018. 3, 5, 6
  6. 6.H. Caesar, J. Uijlings, and V. Ferrari. COCO-Stuff: Thing and stuff classes in context. In CVPR, 2018. 2, 5
  7. 7.Z. Cai and N. Vasconcelos. Cascade R-CNN: Delving into high quality object detection. In CVPR, 2018. 3
  8. 8.J. Cao, Y. Pang, and X. Li. Triply supervised decoder networks for joint detection and segmentation. arXiv preprint arXiv:1809.09299, 2018. 3
  9. 9.L.-C. Chen, A. Hermans, G. Papandreou, F. Schroff, P. Wang, and H. Adam. MaskLab: Instance segmentation by refining object detection with semantic and direction features. In CVPR, 2018. 1, 3
  10. 10.L.-C. Chen, G. Papandreou, I. Kokkinos, K. Murphy, and A. L. Yuille. DeepLab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected CRFs. PAMI, 2018. 1, 3, 6
  11. 11.L.-C. Chen, G. Papandreou, F. Schroff, and H. Adam. Rethinking atrous convolution for semantic image segmentation. arXiv:1706.05587, 2017. 6
  12. 12.L.-C. Chen, Y. Zhu, G. Papandreou, F. Schroff, and H. Adam. Encoder-decoder with atrous separable convolution for semantic image segmentation. In ECCV, 2018. 2, 3, 6
  13. 13.J.-T. Chien and H.-T. Chen. COCO-Stuff 2017 Challenge: Vllab team. 2017. 6
  14. 14.M. Cordts, M. Omran, S. Ramos, T. Rehfeld, M. Enzweiler, R. Benenson, U. Franke, S. Roth, and B. Schiele. The cityscapes dataset for semantic urban scene understanding. In CVPR, 2016. 1, 2, 5
  15. 15.J. Dai, H. Qi, Y. Xiong, Y. Li, G. Zhang, H. Hu, and Y. Wei. Deformable convolutional networks. In ICCV, 2017. 3
  16. 16.D. de Geus, P. Meletis, and G. Dubbelman. Panoptic segmentation with a joint semantic and instance segmentation network. arXiv:1809.02110, 2018. 8
  17. 17.N. Dvornik, K. Shmelkov, J. Mairal, and C. Schmid. BlitzNet: A real-time deep network for scene understanding. In ICCV, 2017. 3
  18. 18.M. Everingham, S. A. Eslami, L. Van Gool, C. K. Williams, J. Winn, and A. Zisserman. The PASCAL visual object classes challenge: A retrospective. IJCV, 2015. 1, 5
  19. 19.A. Fathi and K. Murphy. COCO-Stuff 2017 Challenge: G-RMI team. 2017. 6
  20. 20.G. Ghiasi and C. C. Fowlkes. Laplacian pyramid reconstruction and refinement for semantic segmentation. In ECCV, 2016. 3
  21. 21.R. Girshick. Fast R-CNN. In ICCV, 2015. 3
  22. 22.R. Girshick, J. Donahue, T. Darrell, and J. Malik. Rich feature hierarchies for accurate object detection and semantic segmentation. In CVPR, 2014. 3
  23. 23.R. Girshick, I. Radosavovic, G. Gkioxari, P. Dollár, and K. He. Detectron. https://github.com/facebookresearch/detectron, 2018. 2, 5
  24. 24.K. He, G. Gkioxari, P. Dollár, and R. Girshick. Mask R-CNN. In ICCV, 2017. 1, 2, 3, 4, 5
  25. 25.K. He, X. Zhang, S. Ren, and J. Sun. Deep residual learning for image recognition. In CVPR, 2016. 3, 4
  26. 26.S. Honari, J. Yosinski, P. Vincent, and C. Pal. Recombinator networks: Learning coarse-to-fine feature aggregation. In CVPR, 2016. 3
  27. 27.J. Hu, L. Shen, and G. Sun. Squeeze-and-excitation networks. In CVPR, 2018. 6
  28. 28.S. Ioffe and C. Szegedy. Batch normalization: Accelerating deep network training by reducing internal covariate shift. In ICML, 2015. 4
  29. 29.A. Kendall, Y. Gal, and R. Cipolla. Multi-task learning using uncertainty to weigh losses for scene geometry and semantics. In CVPR, 2018. 3, 6
  30. 30.A. Kirillov, K. He, R. Girshick, C. Rother, and P. Dollár. Panoptic segmentation. In CVPR, 2019. 1, 2, 3, 4, 5, 8
  31. 31.A. Kirillov, E. Levinkov, B. Andres, B. Savchynskyy, and C. Rother. InstanceCut: from edges to instances with multicut. In CVPR, 2017. 3
  32. 32.I. Kokkinos. UberNet: Training a universal convolutional neural network for low-, mid-, and high-level vision using diverse datasets and limited memory. In CVPR, 2017. 3, 6
  33. 33.J. Li, A. Raventos, A. Bhargava, T. Tagawa, and A. Gaidon. Learning to fuse things and stuff. arXiv:1812.01192, 2018. 2
  34. 34.Q. Li, A. Arnab, and P. H. Torr. Weakly-and semi-supervised panoptic segmentation. In ECCV, 2018. 8
  35. 35.Y. Li, H. Qi, J. Dai, X. Ji, and Y. Wei. Fully convolutional instance-aware semantic segmentation. In CVPR, 2017. 3
  36. 36.T.-Y. Lin, P. Dollár, R. Girshick, K. He, B. Hariharan, and S. Belongie. Feature pyramid networks for object detection. In CVPR, 2017. 1, 2, 3
  37. 37.T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick. Microsoft COCO: Common objects in context. In ECCV, 2014. 1, 2, 3, 5
  38. 38.S. Liu, J. Jia, S. Fidler, and R. Urtasun. SGN: Sequential grouping networks for instance segmentation. In CVPR, 2017. 3
  39. 39.S. Liu, L. Qi, H. Qin, J. Shi, and J. Jia. Path aggregation network for instance segmentation. In CVPR, 2018. 3
  40. 40.W. Liu, D. Anguelov, D. Erhan, C. Szegedy, S. Reed, C.-Y. Fu, and A. C. Berg. SSD: Single shot multibox detector. In ECCV, 2016. 5, 6
  41. 41.J. Long, E. Shelhamer, and T. Darrell. Fully convolutional networks for semantic segmentation. In CVPR, 2015. 1, 3
  42. 42.I. Misra, A. Shrivastava, A. Gupta, and M. Hebert. Cross-stitch networks for multi-task learning. In CVPR, 2016. 3
  43. 43.G. Neuhold, T. Ollmann, S. Rota Bulò, and P. Kontschieder. The mapillary vistas dataset for semantic understanding of street scenes. In CVPR, 2017. 1, 2, 3
  44. 44.A. Newell, K. Yang, and J. Deng. Stacked hourglass networks for human pose estimation. In ECCV, 2016. 3
  45. 45.C. Peng, T. Xiao, Z. Li, Y. Jiang, X. Zhang, K. Jia, G. Yu, and J. Sun. Megdet: A large mini-batch object detector. In CVPR, 2018. 3
  46. 46.V.-Q. Pham, S. Ito, and T. Kozakaya. BiSeg: Simultaneous instance segmentation and semantic segmentation with fully convolutional networks. In BMVC, 2017. 1, 3
  47. 47.P. O. Pinheiro, T.-Y. Lin, R. Collobert, and P. Dollár. Learning to refine object segments. In ECCV, 2016. 3
  48. 48.S. Ren, K. He, R. Girshick, and J. Sun. Faster R-CNN: Towards real-time object detection with region proposal networks. In NIPS, 2015. 3
  49. 49.O. Ronneberger, P. Fischer, and T. Brox. U-Net: Convolutional networks for biomedical image segmentation. In MICCAI, 2015. 3, 5
  50. 50.O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, A. C. Berg, and L. Fei-Fei. ImageNet Large Scale Visual Recognition Challenge. IJCV, 2015. 4
  51. 51.J. Tighe, M. Niethammer, and S. Lazebnik. Scene parsing with object instances and occlusion ordering. In CVPR, 2014. 2
  52. 52.Z. Tu, X. Chen, A. L. Yuille, and S.-C. Zhu. Image parsing: Unifying segmentation, detection, and recognition. IJCV, 2005. 2
  53. 53.X. Wang, R. Girshick, A. Gupta, and K. He. Non-local neural networks. In CVPR, 2018. 6
  54. 54.Y. Wu and K. He. Group normalization. In ECCV, 2018. 4
  55. 55.S. Xie, R. Girshick, P. Dollár, Z. Tu, and K. He. Aggregated residual transformations for deep neural networks. In CVPR, 2017. 2, 4
  56. 56.J. Yao, S. Fidler, and R. Urtasun. Describing the scene as a whole: Joint object detection, scene classification and semantic segmentation. In CVPR, 2012. 2
  57. 57.F. Yu and V. Koltun. Multi-scale context aggregation by dilated convolutions. In ICLR, 2016. 1, 3
  58. 58.H. Zhao, J. Shi, X. Qi, X. Wang, and J. Jia. Pyramid scene parsing network. In CVPR, 2017. 3
  59. 59.H. Zhao, Y. Zhang, S. Liu, J. Shi, C. C. Loy, D. Lin, and J. Jia. PSANet: Point-wise spatial attention network for scene parsing. In ECCV, 2018. 3, 6
  60. 60.B. Zhou, H. Zhao, X. Puig, S. Fidler, A. Barriuso, and A. Torralba. Scene parsing through ADE20K dataset. In CVPR, 2017. 1, 3

Citation

MLA
Kirillov, A., et al. “Panoptic Feature Pyramid Networks”. arXiv, 2019, http://arxiv.org/abs/1901.02446v2.
APA
Kirillov, A., Girshick, R., He, K., & Dollár, P. (2019). Panoptic Feature Pyramid Networks. arXiv. http://arxiv.org/abs/1901.02446v2
Chicago
Kirillov, A., R. Girshick, K. He, and P. Dollár. 2019. “Panoptic Feature Pyramid Networks”. arXiv. http://arxiv.org/abs/1901.02446v2.
Harvard
Kirillov, A. et al. (2019) “Panoptic Feature Pyramid Networks”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1901.02446v2.
Vancouver
1. Kirillov A, Girshick R, He K, Dollár P (2019) Panoptic Feature Pyramid Networks. arXiv

BibTeX

@article{kirillov2019panoptic,
  title = {Panoptic Feature Pyramid Networks},
  author = {Kirillov, Alexander and Girshick, Ross and He, Kaiming and Dollár, Piotr},
  year = {2019},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1901.02446v2},
  eprint = {1901.02446}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE