PANet: Few-Shot Image Semantic Segmentation With Prototype Alignment

Kaixin WangJun Hao LiewYingtian ZouDaquan ZhouJiashi Feng

article2019ICCV1,511 citations

Proposes a metric learning framework for few-shot image semantic segmentation that learns representative class prototypes and enforces bidirectional alignment between support and query images to improve generalization to unseen categories.

Listen

Deep learning models for image semantic segmentation typically require vast quantities of densely annotated training images, making them expensive to deploy and poor at generalizing to unseen object classes. Few-shot segmentation addresses this bottleneck by enabling models to segment new classes using only a handful of annotated examples. The article introduces PANet (Prototype Alignment Network), a framework designed to accurately segment new object classes from minimal data by combining prototype-based metric learning with a novel alignment regularization.

The core objective of the article is to demonstrate that separating knowledge extraction from the segmentation process and enforcing bidirectional consistency between support examples and query images improves few-shot segmentation performance over complex parametric architectures.

To evaluate this concept, the authors developed a lightweight architecture using a standard convolutional feature extractor. Instead of complex learned decoder modules, the framework computes class prototypes using masked average pooling and classifies query pixels by matching them to the nearest prototype. During training, a prototype alignment regularization enforces mutual consistency by using the predicted query mask to segment the original support image in reverse. The system was benchmarked across standard splits on benchmark datasets, evaluating both fully dense annotations and weak annotations such as scribbles and bounding boxes.

The findings show that PANet establishes a new state-of-the-art across key benchmarks. On standard benchmarks, the method achieved a mean Intersection-over-Union (mIoU) score of 48.1% in the 1-shot setting and 55.7% in the 5-shot setting, outperforming prior state-of-the-art models by 1.8% and 8.6%, respectively. In multi-class settings, it improved accuracy by more than 20% over existing approaches. The prototype alignment regularization reduced the feature distance between query and support representations from 42.6 to 32.2, accelerating training convergence and boosting overall accuracy. Furthermore, PANet demonstrated strong robustness when provided with weak annotations, achieving 44.8% mIoU with scribbles and 45.1% with bounding boxes in 1-shot tests.

These results demonstrate that high-performance visual segmentation does not require complex parameter-heavy decoders or cost-prohibitive full-pixel labeling. Organizations can deploy segmentation models with lower computational overhead, reduced memory footprint, and significantly shorter data annotation timelines. The model's compatibility with scribbles and bounding boxes drastically cuts operational labeling costs while retaining competitive accuracy.

Engineering teams should consider non-parametric prototype metric learning when designing few-shot vision systems and explore weak annotation workflows to minimize data collection costs. Future technical efforts should focus on integrating post-processing refinements to resolve boundary artifacts and investigating interactive segmentation systems where annotations are updated dynamically.

The main limitations include vulnerability to spatially isolated classification errors—such as patchy segmentations resulting from pixel-independent predictions—and difficulties differentiating semantically similar classes with overlapping feature representations, such as chairs and tables. Nonetheless, the experimental results provide high confidence in the model's overall generalization and efficiency across standard vision benchmarks.

arXiv: 1908.06391
  • Paper: Per-Pixel Classification is Not All You Need for Semantic Segmentation, Bowen Cheng et al. (2021). MaskFormer shifts segmentation from per-pixel classification to mask classification, providing a modern alternative framework to PANet's pixel-to-prototype matching approach.
  • Paper: Masked-attention Mask Transformer for Universal Image Segmentation, Bowen Cheng et al. (2022). Mask2Former further advances universal mask attention and query decoding, presenting the next paradigm shift beyond prototype-based dense segmentation architectures.
  • Paper: Image Segmentation Using Deep Learning: A Survey, Shervin Minaee et al. (2020). This comprehensive survey contextualizes metric learning and prototype-based segmentation strategies alongside broader deep learning advances in image segmentation.
  • Paper: Segment Anything, Alexander M. Kirillov et al. (2023). Segment Anything scales promptable few-shot and zero-shot segmentation to foundation-model scale using prompt encoders and dense mask decoders.
  • Paper: SAM 3: Segment Anything with Concepts, Nicolas Carion et al. (2025). SAM 3 extends open-world and few-shot promptable segmentation to unified concept-level visual reasoning across images and video.
Cover for PANet: Few-Shot Image Semantic Segmentation With Prototype Alignment

Abstract

Despite the great progress made by deep CNNs in image semantic segmentation, they typically require a large number of densely-annotated images for training and are difficult to generalize to unseen object categories. Few-shot segmentation has thus been developed to learn to perform segmentation from only a few annotated examples. In this paper, we tackle the challenging few-shot segmentation problem from a metric learning perspective and present PANet, a novel prototype alignment network to better utilize the information of the support set. Our PANet learns class-specific prototype representations from a few support images within an embedding space and then performs segmentation over the query images through matching each pixel to the learned prototypes. With non-parametric metric learning, PANet offers high-quality prototypes that are representative for each semantic class and meanwhile discriminative for different classes. Moreover, PANet introduces a prototype alignment regularization between support and query. With this, PANet fully exploits knowledge from the support and provides better generalization on few-shot segmentation. Significantly, our model achieves the mIoU score of 48.1% and 55.7% on PASCAL-5i for 1-shot and 5-shot settings respectively, surpassing the state-of-the-art method by 1.8% and 8.6%.

Table of Contents

  • 1 Introduction
  • 2 Related work
  • 3 Method
  • 3.1 Problem setting
  • 3.2 Method overview
  • 3.3 Prototype learning
  • 3.4 Non-parametric metric learning
  • 3.5 Prototype alignment regularization (PAR)
  • 3.6 Generalization to weaker annotations
  • 4 Experiments
  • 4.1 Setup
  • 4.2 Comparison with state-of-the-arts
  • 4.3 Analysis on PAR
  • 4.4 Test with weak annotations
  • 5 Conclusion
  • References

Knowls

  1. Knowl 1 — Prototype Alignment Network (PANet) Architecture

    model/method

    The Prototype Alignment Network (PANet) is a metric learning-based framework for few-shot image semantic segmentation that maps support and query images into a shared embedding space without requiring parametric task-specific decoders.

    The PANet architecture consists of:

    1. Backbone Feature Extractor: A shared convolutional neural network based on VGG-16 truncated after the first 5 convolutional blocks. To retain high spatial resolution for dense pixel prediction, the stride of the maxpool4 layer is set to 1. To enlarge the receptive field without loss of resolution, standard convolutions in the conv5 block are replaced with dilated convolutions with a dilation rate of 2.
    2. Support Prototype Extraction: Deep feature maps of support images are masked via late-fusion masked average pooling to extract dense spatial representations into single embedding vectors (prototypes) for each semantic foreground class and for the background.
    3. Non-Parametric Metric Learning: Query segmentation is executed directly by calculating the cosine distance between the query feature vector at each spatial coordinate and each class prototype, followed by a temperature-scaled softmax.
    4. Prototype Alignment Regularization (PAR): During training, a reverse few-shot segmentation task is performed by treating the query image along with its predicted probability mask as a new support set to segment the original support images.

    PANet introduces no extra learnable parameters beyond the shared convolutional backbone, reducing susceptibility to overfitting in few-shot regimes and maintaining the same computational cost at inference time.

  2. Knowl 2 — Class and Background Prototype Extraction via Masked Average Pooling

    equation

    In the Prototype Alignment Network (PANet), class-specific foreground and background prototypes are extracted from support images using a late-fusion masked average pooling strategy over the backbone feature representations.

    Let S={(Ic,k,Mc,k)}\mathcal{S} = \{(I_{c,k}, M_{c,k})\} denote a support set for an episode, where c∈Cc \in \mathcal{C} indexes the semantic classes (∣C∣=C|\mathcal{C}| = C), k∈{1,…,K}k \in \{1, \dots, K\} indexes the KK support images per class, Ic,kI_{c,k} is an input image, and Mc,kM_{c,k} is its corresponding ground-truth segmentation mask. Let Fc,k∈RH×W×DF_{c,k} \in \mathbb{R}^{H \times W \times D} denote the feature map output by the backbone for image Ic,kI_{c,k}, where (x,y)(x,y) indexes spatial coordinates.

    The prototype pc∈RDp_c \in \mathbb{R}^D for foreground class cc is computed as: pc=1K∑k=1K∑x,yFc,k(x,y)1[Mc,k(x,y)=c]∑x,y1[Mc,k(x,y)=c]p_c = \frac{1}{K} \sum_{k=1}^K \frac{\sum_{x,y} F_{c,k}^{(x,y)} \mathbf{1}[M_{c,k}^{(x,y)} = c]}{\sum_{x,y} \mathbf{1}[M_{c,k}^{(x,y)} = c]}

    The background prototype pbg∈RDp_{bg} \in \mathbb{R}^D is computed across all C×KC \times K support images as: pbg=1CK∑c∈C∑k=1K∑x,yFc,k(x,y)1[Mc,k(x,y)∉C]∑x,y1[Mc,k(x,y)∉C]p_{bg} = \frac{1}{CK} \sum_{c \in \mathcal{C}} \sum_{k=1}^K \frac{\sum_{x,y} F_{c,k}^{(x,y)} \mathbf{1}[M_{c,k}^{(x,y)} \notin \mathcal{C}]}{\sum_{x,y} \mathbf{1}[M_{c,k}^{(x,y)} \notin \mathcal{C}]}

    where 1[⋅]\mathbf{1}[\cdot] is an indicator function evaluating to 1 if the condition is satisfied and 0 otherwise. The full set of prototypes for the episode is P={pc∣c∈C}∪{pbg}\mathcal{P} = \{p_c \mid c \in \mathcal{C}\} \cup \{p_{bg}\}.

  3. Knowl 3 — Non-Parametric Metric Learning for Dense Pixel Prediction

    equation

    Dense pixel-level segmentation on a query image is performed in PANet via non-parametric metric learning in the embedding feature space against the prototype set P={pc∣c∈C}∪{pbg}\mathcal{P} = \{p_c \mid c \in \mathcal{C}\} \cup \{p_{bg}\}.

    Let Fq∈RH×W×DF_q \in \mathbb{R}^{H \times W \times D} denote the feature map of the query image IqI_q, and let Fq(x,y)∈RDF_q^{(x,y)} \in \mathbb{R}^D be the feature vector at spatial position (x,y)(x,y). The predicted probability M~q;j(x,y)\tilde{M}_{q;j}^{(x,y)} that spatial location (x,y)(x,y) belongs to class j∈C∪{bg}j \in \mathcal{C} \cup \{bg\} is computed via a scaled softmax over cosine distances: M~q;j(x,y)=exp⁡(−α d(Fq(x,y),pj))∑p∈Pexp⁡(−α d(Fq(x,y),p))\tilde{M}_{q;j}^{(x,y)} = \frac{\exp(-\alpha \, d(F_q^{(x,y)}, p_j))}{\sum_{p \in \mathcal{P}} \exp(-\alpha \, d(F_q^{(x,y)}, p))}

    where d(u,v)=1−u⋅v∥u∥2∥v∥2d(u, v) = 1 - \frac{u \cdot v}{\|u\|_2 \|v\|_2} is the cosine distance, and α\alpha is a fixed scaling factor set to α=20\alpha = 20.

    The discrete predicted query segmentation mask M^q\hat{M}_q is given by: M^q(x,y)=arg⁡max⁡j∈C∪{bg}M~q;j(x,y)\hat{M}_q^{(x,y)} = \arg\max_{j \in \mathcal{C} \cup \{bg\}} \tilde{M}_{q;j}^{(x,y)}

    The segmentation loss Lseg\mathcal{L}_{seg} over the query image is defined as the cross-entropy loss: Lseg=−1N∑x,y∑pj∈P1[Mq(x,y)=j]log⁡M~q;j(x,y)\mathcal{L}_{seg} = -\frac{1}{N} \sum_{x,y} \sum_{p_j \in \mathcal{P}} \mathbf{1}[M_q^{(x,y)} = j] \log \tilde{M}_{q;j}^{(x,y)} where MqM_q is the ground-truth query mask and N=H×WN = H \times W is the total number of spatial locations.

  4. Knowl 4 — Prototype Alignment Regularization (PAR)

    model/method

    Prototype Alignment Regularization (PAR) is a training regularization technique that enforces mutual alignment between support and query feature prototypes by performing few-shot segmentation in a reverse direction.

    While standard few-shot segmentation passes information from the support set S\mathcal{S} to the query image IqI_q, PAR swaps their roles during training:

    1. Using the query feature map FqF_q and the predicted query probability map M~q\tilde{M}_q, a new set of query-derived prototypes Pˉ={pˉc∣c∈C}∪{pˉbg}\bar{\mathcal{P}} = \{\bar{p}_c \mid c \in \mathcal{C}\} \cup \{\bar{p}_{bg}\} is computed via masked average pooling.
    2. The model uses Pˉ\bar{\mathcal{P}} to segment each support image Ic,k∈SI_{c,k} \in \mathcal{S}, computing the predicted probability map M~c,k\tilde{M}_{c,k} for class j∈C∪{bg}j \in \mathcal{C} \cup \{bg\} at spatial location (x,y)(x,y): M~c,k;j(x,y)=exp⁡(−α d(Fc,k(x,y),pˉj))∑pˉ∈Pˉexp⁡(−α d(Fc,k(x,y),pˉ))\tilde{M}_{c,k;j}^{(x,y)} = \frac{\exp(-\alpha \, d(F_{c,k}^{(x,y)}, \bar{p}_j))}{\sum_{\bar{p} \in \bar{\mathcal{P}}} \exp(-\alpha \, d(F_{c,k}^{(x,y)}, \bar{p}))}
    3. The regularization loss LPAR\mathcal{L}_{PAR} is calculated against the ground-truth support masks Mc,kM_{c,k}: LPAR=−1CKN∑c∈C∑k=1K∑x,y∑j∈C∪{bg}1[Mc,k(x,y)=j]log⁡M~c,k;j(x,y)\mathcal{L}_{PAR} = -\frac{1}{CKN} \sum_{c \in \mathcal{C}} \sum_{k=1}^K \sum_{x,y} \sum_{j \in \mathcal{C} \cup \{bg\}} \mathbf{1}[M_{c,k}^{(x,y)} = j] \log \tilde{M}_{c,k;j}^{(x,y)} where C=∣C∣C = |\mathcal{C}| is the number of semantic classes, KK is the number of support shots, and NN is the spatial resolution.

    The overall training loss is: L=Lseg+λLPAR\mathcal{L} = \mathcal{L}_{seg} + \lambda \mathcal{L}_{PAR} where λ=1\lambda = 1 is the regularization weight. PAR is active only during training and introduces zero additional computational overhead during inference.

  5. Knowl 5 — Training and Testing Algorithm for PANet

    algorithm

    The complete episodic training and evaluation pipeline for the Prototype Alignment Network (PANet) operates by extracting prototypes and calculating forward and reverse matching losses.

    Input: Training set Dtrain\mathcal{D}_{train}, Testing set Dtest\mathcal{D}_{test}, feature extractor weights WW, temperature scaling factor α=20\alpha = 20, loss weight λ=1\lambda = 1.
    Output: Evaluated segmentation predictions on Dtest\mathcal{D}_{test}.
    for each episode (Si,Qi)∈Dtrain(\mathcal{S}_i, \mathcal{Q}_i) \in \mathcal{D}_{train} do
        Extract support features Fc,kF_{c,k} for each (Ic,k,Mc,k)∈Si(I_{c,k}, M_{c,k}) \in \mathcal{S}_i using backbone WW
        Extract prototypes P={pc∣c∈Ci}∪{pbg}\mathcal{P} = \{p_c \mid c \in \mathcal{C}_i\} \cup \{p_{bg}\} via masked average pooling on Si\mathcal{S}_i
        Extract query features FqF_q for (Iq,Mq)∈Qi(I_q, M_q) \in \mathcal{Q}_i using backbone WW
        Compute query class probabilities M~q\tilde{M}_q via cosine metric matching against P\mathcal{P} scaled by α\alpha
        Compute query segmentation loss Lseg\mathcal{L}_{seg} using ground truth MqM_q
        Extract query prototypes Pˉ={pˉc∣c∈Ci}∪{pˉbg}\bar{\mathcal{P}} = \{\bar{p}_c \mid c \in \mathcal{C}_i\} \cup \{\bar{p}_{bg}\} via masked average pooling on FqF_q and M~q\tilde{M}_q
        Compute support class probabilities M~c,k\tilde{M}_{c,k} via cosine metric matching against Pˉ\bar{\mathcal{P}} scaled by α\alpha
        Compute regularization loss LPAR\mathcal{L}_{PAR} using ground truth support masks Mc,kM_{c,k}
        Compute total loss L=Lseg+λLPAR\mathcal{L} = \mathcal{L}_{seg} + \lambda \mathcal{L}_{PAR}
        Update backbone weights WW by SGD
    end
    for each episode (Si,Qi)∈Dtest(\mathcal{S}_i, \mathcal{Q}_i) \in \mathcal{D}_{test} do
        Extract support features Fc,kF_{c,k} for each (Ic,k,Mc,k)∈Si(I_{c,k}, M_{c,k}) \in \mathcal{S}_i using backbone WW
        Extract prototypes P={pc∣c∈Ci}∪{pbg}\mathcal{P} = \{p_c \mid c \in \mathcal{C}_i\} \cup \{p_{bg}\} via masked average pooling on Si\mathcal{S}_i
        Extract query features FqF_q for Iq∈QiI_q \in \mathcal{Q}_i using backbone WW
        Compute query class probabilities M~q\tilde{M}_q via cosine metric matching against P\mathcal{P} scaled by α\alpha
        Predict query mask M^q(x,y)=arg⁡max⁡jM~q;j(x,y)\hat{M}_q^{(x,y)} = \arg\max_j \tilde{M}_{q;j}^{(x,y)}
    end

    The feature backbone is initialized with ImageNet pre-trained VGG-16 weights and trained end-to-end using SGD with momentum 0.9, weight decay 0.0005, batch size 1, and an initial learning rate of 1e-3 (decayed by 0.1 every 10,000 iterations) for 30,000 iterations. Input images are resized to 417×417417 \times 417 and augmented with random horizontal flips.

  6. Knowl 6 — Few-Shot Segmentation Performance on PASCAL-5i

    data/table

    PANet was evaluated on the PASCAL-5i5^i benchmark (created from PASCAL VOC 2012 with SBD augmentation, split into 4 folds of 5 classes each) under 1-way 1-shot and 1-way 5-shot protocols across 5 runs of 1,000 episodes.

    Method 1-shot mean-IoU (%) 5-shot mean-IoU (%) Δ\Delta #Params
    split-1 split-2 split-3 split-4 Mean split-1 split-2 split-3 split-4 Mean Mean
    OSLSM 33.6 55.3 40.9 33.5 40.8 35.9 58.1 42.7 39.1 43.9 3.1 272.6M
    co-FCN 36.7 50.6 44.9 32.4 41.1 37.5 50.0 44.1 33.9 41.4 0.3 34.2M
    SG-One 40.2 58.4 48.4 38.4 46.3 41.9 58.6 48.6 39.4 47.1 0.8 19.0M
    PANet-init 30.8 40.7 38.3 31.4 35.3 41.6 52.7 51.6 40.8 46.7 11.4 14.7M
    PANet 42.3 58.0 51.1 41.2 48.1 51.8 64.6 59.8 46.5 55.7 7.6 14.7M

    Under the binary-IoU metric on PASCAL-5i5^i:

    Method 1-shot binary-IoU (%) 5-shot binary-IoU (%) Δ\Delta (%)
    FG-BG 55.0 - -
    Fine-tuning 55.1 55.6 0.5
    OSLSM 61.3 61.5 0.2
    co-FCN 60.1 60.2 0.1
    PL 61.2 62.3 1.1
    A-MCG 61.2 62.2 1.0
    SG-One 63.9 65.9 2.0
    PANet-init 58.9 65.7 6.8
    PANet 66.5 70.7 4.2

    PANet achieves 48.1% mean-IoU in 1-shot (surpassing SG-One by 1.8%) and 55.7% in 5-shot (surpassing SG-One by 8.6%). The parameter count is reduced to 14.7M. The untrained baseline PANet-init (using ImageNet weights without fine-tuning on PASCAL-5i5^i) reaches 46.7% 5-shot mean-IoU, rivaling trained prior architectures.

  7. Knowl 7 — Multi-Way and MS COCO Few-Shot Segmentation Performance

    data/table

    PANet was evaluated on 2-way few-shot segmentation on PASCAL-5i5^i and on 1-way few-shot segmentation on the MS COCO dataset (divided into 4 splits of 20 categories each):

    2-way Few-Shot Segmentation on PASCAL-5i5^i:

    Method mean-IoU (%) binary-IoU (%)
    1-shot 5-shot 1-shot 5-shot
    PL - - 42.7 43.7
    SG-One - 29.4 - -
    PANet 45.1 53.1 64.2 67.9

    On 2-way 5-shot segmentation, PANet outperforms SG-One by 23.7% in mean-IoU (53.1% vs 29.4%) and outperforms PL by 24.2% in binary-IoU.

    1-way Few-Shot Segmentation on MS COCO:

    Method mean-IoU (%) binary-IoU (%)
    1-shot 5-shot 1-shot 5-shot
    A-MCG - - 52.0 54.7
    PANet 20.9 29.7 59.2 63.5

    On MS COCO, PANet achieves 20.9% 1-shot and 29.7% 5-shot mean-IoU, improving over A-MCG in binary-IoU by 7.2% and 8.8% for 1-shot and 5-shot tasks respectively.

  8. Knowl 8 — Impact of Prototype Alignment Regularization on Feature Alignment and Performance

    empirical result

    Ablation and quantitative analysis of Prototype Alignment Regularization (PAR) reveal three main empirical effects:

    1. Segmentation Metric Improvement: On PASCAL-5i5^i, training with PAR improves mean-IoU from 47.2% to 48.1% (+0.9%) in the 1-shot setting and from 54.9% to 55.7% (+0.8%) in the 5-shot setting.
    2. Closer Prototype Distance: Across 1,000 randomly sampled episodes on PASCAL-5i5^i split-1 under the 1-way 5-shot setting, the average Euclidean distance between query-derived prototypes and support-derived prototypes decreases from 42.6 (without PAR) to 32.2 (with PAR), confirming that bidirectional regularization aligns prototypes closer in the embedding space.
    3. Accelerated Training Convergence: PANet trained with PAR converges faster across training iterations and achieves a lower overall loss curve than training without PAR, particularly in the 5-shot setting.
  9. Knowl 9 — Few-Shot Segmentation with Weak Annotations (Scribbles and Bounding Boxes)

    empirical result

    PANet can perform few-shot segmentation at test time using weak support annotations (scribbles or bounding boxes) instead of dense pixel masks, without changing the network architecture:

    Support Annotation Type 1-shot mean-IoU (%) 5-shot mean-IoU (%)
    Dense 48.1 55.7
    Scribble 44.8 54.6
    Bounding Box 45.1 52.8

    In the 1-shot setting, bounding box annotations achieve 45.1% mean-IoU and scribbles achieve 44.8% mean-IoU (compared to 48.1% for dense masks). In the 5-shot setting, scribble annotations achieve 54.6% mean-IoU (within 1.1% of dense masks), outperforming bounding box annotations (52.8% mean-IoU) by 1.8%, because bounding box pooling aggregates extraneous background pixels across multiple support images.

  10. Knowl 10 — Limitations of Pixel-Independent Metric Prediction and Embedding Ambiguity

    limitation

    The Prototype Alignment Network exhibits two main failure modes:

    1. Unnatural Spatial Discontinuities: Because pixel classification is performed independently at each spatial feature location via metric distance calculation without spatial consistency regularization, structured decoders, or post-processing (such as DenseCRF), predictions can produce fragmented or unnatural patches.
    2. Embedding Ambiguity for Visually Similar Classes: When semantic categories share high visual or contextual similarity (such as chairs and tables), their feature prototypes can overlap in the learned metric embedding space, causing the model to misclassify pixels between related classes.

Coverage note — None was omitted; all primary model components, mathematical formulations, algorithms, experimental benchmarks on PASCAL-5i and MS COCO, weak supervision analyses, ablations, and failure modes have been fully covered.

References

  1. 1.Vijay Badrinarayanan, Alex Kendall, and Roberto Cipolla. Segnet: A deep convolutional encoder-decoder architecture for image segmentation. IEEE transactions on pattern analysis and machine intelligence, 39(12):2481–2495, 2017.
  2. 2.Liang-Chieh Chen, George Papandreou, Iasonas Kokkinos, Kevin Murphy, and Alan L Yuille. Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs. IEEE transactions on pattern analysis and machine intelligence, 40(4):834–848, 2018.
  3. 3.Jifeng Dai, Kaiming He, and Jian Sun. Boxsup: Exploiting bounding boxes to supervise convolutional networks for semantic segmentation. In Proceedings of the IEEE International Conference on Computer Vision, pages 1635–1643, 2015.
  4. 4.Nanqing Dong and Eric P Xing. Few-shot semantic segmentation with prototype learning. In BMVC, volume 3, page 4, 2018.
  5. 5.Mark Everingham, Luc Van Gool, Christopher KI Williams, John Winn, and Andrew Zisserman. The pascal visual object classes (voc) challenge. International journal of computer vision, 88(2):303–338, 2010.
  6. 6.Chelsea Finn, Pieter Abbeel, and Sergey Levine. Model-agnostic meta-learning for fast adaptation of deep networks. In Proceedings of the 34th International Conference on Machine Learning-Volume 70, pages 1126–1135. JMLR. org, 2017.
  7. 7.Bharath Hariharan, Pablo Arbelaez, Lubomir Bourdev, ´ Subhransu Maji, and Jitendra Malik. Semantic contours from inverse detectors. 2011.
  8. 8.Tao Hu, Pengwan, Chiliang Zhang, Gang Yu, Yadong Mu, and Cees G. M. Snoek. Attention-based multi-context guiding for few-shot semantic segmentation. 2018.
  9. 9.Di Lin, Jifeng Dai, Jiaya Jia, Kaiming He, and Jian Sun. Scribblesup: Scribble-supervised convolutional networks for semantic segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 3159–3167, 2016.
  10. 10.Guosheng Lin, Anton Milan, Chunhua Shen, and Ian Reid. Refinenet: Multi-path refinement networks for high-resolution semantic segmentation. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 1925–1934, 2017.
  11. 11.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollar, and C Lawrence ´ Zitnick. Microsoft coco: Common objects in context. In European conference on computer vision, pages 740–755. Springer, 2014.
  12. 12.Yanbin Liu, Juho Lee, Minseop Park, Saehoon Kim, Eunho Yang, Sung Ju Hwang, and Yi Yang. Learning to propagate labels: Transductive propagation network for few-shot learning. 2018.
  13. 13.Jonathan Long, Evan Shelhamer, and Trevor Darrell. Fully convolutional networks for semantic segmentation. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 3431–3440, 2015.
  14. 14.Boris Oreshkin, Pau Rodr´ıguez Lopez, and Alexandre La- ´ coste. Tadam: Task dependent adaptive metric for improved few-shot learning. In Advances in Neural Information Processing Systems, pages 719–729, 2018.
  15. 15.George Papandreou, Liang-Chieh Chen, Kevin P Murphy, and Alan L Yuille. Weakly-and semi-supervised learning of a deep convolutional network for semantic image segmentation. In Proceedings of the IEEE international conference on computer vision, pages 1742–1750, 2015.
  16. 16.Kate Rakelly, Evan Shelhamer, Trevor Darrell, Alyosha Efros, and Sergey Levine. Conditional networks for few-shot semantic segmentation. 2018.
  17. 17.Kate Rakelly, Evan Shelhamer, Trevor Darrell, Alexei A Efros, and Sergey Levine. Few-shot segmentation propagation with guided networks. arXiv preprint arXiv:1806.07373, 2018.
  18. 18.Sachin Ravi and Hugo Larochelle. Optimization as a model for few-shot learning. 2016.
  19. 19.Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei. ImageNet Large Scale Visual Recognition Challenge. International Journal of Computer Vision (IJCV), 115(3):211–252, 2015.
  20. 20.Victor Garcia Satorras and Joan Bruna Estrach. Few-shot learning with graph neural networks. In International Conference on Learning Representations, 2018.
  21. 21.Amirreza Shaban, Shray Bansal, Zhen Liu, Irfan Essa, and Byron Boots. One-shot learning for semantic segmentation. arXiv preprint arXiv:1709.03410, 2017.
  22. 22.Karen Simonyan and Andrew Zisserman. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556, 2014.
  23. 23.Jake Snell, Kevin Swersky, and Richard Zemel. Prototypical networks for few-shot learning. In Advances in Neural Information Processing Systems, pages 4077–4087, 2017.
  24. 24.Flood Sung, Yongxin Yang, Li Zhang, Tao Xiang, Philip HS Torr, and Timothy M Hospedales. Learning to compare: Relation network for few-shot learning. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 1199–1208, 2018.
  25. 25.Oriol Vinyals, Charles Blundell, Timothy Lillicrap, Daan Wierstra, et al. Matching networks for one shot learning. In Advances in neural information processing systems, pages 3630–3638, 2016.
  26. 26.Yunchao Wei, Jiashi Feng, Xiaodan Liang, Ming-Ming Cheng, Yao Zhao, and Shuicheng Yan. Object region mining with adversarial erasing: A simple classification to semantic segmentation approach. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 1568–1576, 2017.
  27. 27.Fisher Yu and Vladlen Koltun. Multi-scale context aggregation by dilated convolutions. arXiv preprint arXiv:1511.07122, 2015.
  28. 28.Xiaolin Zhang, Yunchao Wei, Yi Yang, and Thomas Huang. Sg-one: Similarity guidance network for one-shot semantic segmentation. arXiv preprint arXiv:1810.09091, 2018.
  29. 29.Hengshuang Zhao, Jianping Shi, Xiaojuan Qi, Xiaogang Wang, and Jiaya Jia. Pyramid scene parsing network. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2881–2890, 2017.

Citation

MLA
Wang, K., et al. “PANet: Few-Shot Image Semantic Segmentation with Prototype Alignment”. arXiv, 2019, http://arxiv.org/abs/1908.06391v2.
APA
Wang, K., Liew, J. H., Zou, Y., Zhou, D., & Feng, J. (2019). PANet: Few-Shot Image Semantic Segmentation with Prototype Alignment. arXiv. http://arxiv.org/abs/1908.06391v2
Chicago
Wang, K., J. H. Liew, Y. Zou, D. Zhou, and J. Feng. 2019. “PANet: Few-Shot Image Semantic Segmentation with Prototype Alignment”. arXiv. http://arxiv.org/abs/1908.06391v2.
Harvard
Wang, K. et al. (2019) “PANet: Few-Shot Image Semantic Segmentation with Prototype Alignment”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1908.06391v2.
Vancouver
1. Wang K, Liew JH, Zou Y, Zhou D, Feng J (2019) PANet: Few-Shot Image Semantic Segmentation with Prototype Alignment. arXiv

BibTeX

@article{wang2019panet,
  title = {PANet: Few-Shot Image Semantic Segmentation with Prototype Alignment},
  author = {Wang, Kaixin and Liew, Jun Hao and Zou, Yingtian and Zhou, Daquan and Feng, Jiashi},
  year = {2019},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1908.06391v2},
  eprint = {1908.06391}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE