Rethinking the Correlation in Few-Shot Segmentation: A Buoys View

Yuan WangRui SunTianzhu Zhang

article2023CVPR82 citations

Proposes an adaptive buoys correlation network that suppresses false pixel-level matches in few-shot segmentation by using mined representative reference features to rectify support-query correspondence.

Listen

Deep learning models for image segmentation conventionally require vast volumes of manually annotated data, which is expensive and time-consuming to obtain. Few-shot segmentation addresses this bottleneck by enabling models to recognize and isolate new visual categories using only a few annotated reference examples. However, standard methods suffer from severe matching errors due to complex image backgrounds and substantial differences in appearance or pose between the reference and target images. These errors degrade segmentation reliability in low-data regimes.

The main objective of the article is to demonstrate that introducing a set of representative reference features—termed buoys—can effectively suppress false matches and rectify the direct, pixel-level correlation commonly used in few-shot segmentation models.

To achieve this, the article introduces the Adaptive Buoys Correlation network, designed as a modular plug-in for existing segmentation frameworks. The method operates through two core components: a buoy mining module and an adaptive correlation module. The buoy mining module initializes reference points using matrix decomposition to retain primary visual information, followed by attention mechanisms that capture task-specific context and reconcile appearance discrepancies between image pairs. The adaptive correlation module then evaluates matches dynamically by applying an optimal transport algorithm that downweights irrelevant reference points and assesses the structural similarity of helpful features. The approach was systematically evaluated across two standard benchmarks (PASCAL-5i and COCO-20i) using both 1-shot and 5-shot learning configurations on standard ResNet architectures.

The experimental findings show consistent, broad-based improvements over established baseline models. First, incorporating the buoy network improved segmentation accuracy across all baselines, delivering gains of up to 1.7 percentage points in mean intersection-over-union on PASCAL-5i and up to 2.2 percentage points on the more complex COCO-20i dataset. Second, the performance lift was pronounced in both single-sample (1-shot) and multi-sample (5-shot) settings across different neural network backbones. Third, diagnostic ablation experiments revealed that each design component—including matrix-based initialization, contextual aggregation, and adaptive transport scoring—contributed cumulatively to suppressing false background-to-foreground matches. Fourth, the module achieved these gains with only a minor increase in parameters and computational overhead.

These results demonstrate that explicitly refining feature correlation via structured reference points substantially mitigates errors caused by visual ambiguity and background clutter. For organizations deploying computer vision systems, this approach enhances segmentation accuracy in data-constrained environments without necessitating expensive retraining or complete architectural overhauls, thereby reducing data annotation costs and operational risk.

Based on these findings, engineering teams should evaluate integrating this modular buoy network into existing segmentation pipelines that rely on attention or prior mask guidance. When implementing the system, practitioners should balance the number of reference buoys—setting roughly 24 buoys based on empirical tuning—to avoid either information loss or redundancy.

While the findings are supported by consistent empirical improvements across standard benchmark datasets, the evaluations remain focused on standard natural image datasets. Confidence in the underlying method is high, though teams deploying the framework in specialized domains, such as medical imaging or aerial photography, should conduct domain-specific pilot testing to verify performance under distinct visual conditions.

  • Paper: PANet: Few-Shot Image Semantic Segmentation With Prototype Alignment, Kaixin Wang et al. (2019). PANet establishes the foundational prototype-matching and alignment paradigm for few-shot image segmentation that the source paper directly builds upon and modifies.
  • Paper: Prototypical Networks for Few-shot Learning, Jake Snell et al. (2017). Prototypical Networks introduce metric-based nearest-prototype classification for few-shot learning, providing the fundamental theoretical basis for prototype representations in segmentation.
  • Paper: Matching Networks for One Shot Learning, Oriol Vinyals et al. (2016). Matching Networks pioneer episodic metric learning and learned memory-based matching mechanisms that underpin modern few-shot vision architectures.
  • Paper: Learning to Compare: Relation Network for Few-Shot Learning, Flood Sung et al. (2017). Relation Network introduces end-to-end learnable similarity and relation modules for few-shot comparisons, representing the affinity-learning branch analyzed and improved in the source.
  • Paper: TADAM: Task dependent adaptive metric for improved few-shot learning, Boris N. Oreshkin et al. (2018). TADAM develops task-dependent adaptive metric scaling and feature conditioning for few-shot learning, motivating the adaptive correlation mechanisms in the source paper.
  • Paper: Generalizing from a Few Examples, Yaqing Wang et al. (2019). This survey provides a comprehensive formal taxonomy of few-shot learning formulations, error sources, and prior-knowledge strategies necessary to contextualize few-shot segmentation challenges.
  • Paper: Fully convolutional networks for semantic segmentation, Jonathan Long et al. (2015). Fully Convolutional Networks lay the foundational architecture for dense pixel-level prediction upon which deep semantic segmentation and few-shot segmentation networks operate.
Cover for Rethinking the Correlation in Few-Shot Segmentation: A Buoys View

Abstract

Few-shot segmentation (FSS) aims to segment novel objects in a given query image with only a few annotated support images. However, most previous best-performing methods, whether prototypical learning methods or affinity learning methods, neglect to alleviate false matches caused by their own pixel-level correlation. In this work, we rethink how to mitigate the false matches from the perspective of representative reference features (referred to as buoys), and propose a novel adaptive buoys correlation (ABC) network to rectify direct pairwise pixel-level correlation, including a buoys mining module and an adaptive correlation module. The proposed ABC enjoys several merits. First, to learn the buoys well without any correspondence supervision, we customize the buoys mining module according to the three characteristics of representativeness, task awareness and resilience. Second, the proposed adaptive correlation module is responsible for further endowing buoy-correlation-based pixel matching with an adaptive ability. Extensive experimental results with two different backbones on two challenging benchmarks demonstrate that our ABC, as a general plugin, achieves consistent improvements over several leading methods on both 1-shot and 5-shot settings.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Method
  • 3.1. Problem Formulation
  • 3.2. Overview
  • 3.3. Buoys Mining Module
  • 3.4. Adaptive Correlation Module
  • 4. Experiments
  • 4.1. Dataset and Evaluation Metric
  • 4.2. Implementation Details
  • 4.3. Results and Analysis
  • 4.4. Ablation Study
  • 4.5. Visualizations
  • 5. Conclusion
  • 6. Acknowledgments
  • References

Knowls

  1. Knowl 1 — Adaptive Buoys Correlation Network for Few-Shot Segmentation

    model/method

    The Adaptive Buoys Correlation (ABC) network is a generic plug-in framework designed to rectify false matches in few-shot segmentation (FSS) caused by direct pairwise pixel-level correlation. In standard FSS architectures, pairwise pixel-level correlation matrices W=Sim⁡(FQ,FS)∈Rhw×hw\mathbf{W} = \operatorname{Sim}(\mathbf{F}_Q, \mathbf{F}_S) \in \mathbb{R}^{hw \times hw} (where FQ,FS∈Rhw×C\mathbf{F}_Q, \mathbf{F}_S \in \mathbb{R}^{hw \times C} denote reshaped query and support feature maps with spatial dimensions h×wh \times w and channel dimension CC) often suffer from noise and false matches due to large intra-class variations and cluttered backgrounds.

    ABCNet replaces or rectifies direct pixel-to-pixel matching by measuring correlations with respect to a set of KK representative reference features called "buoys". The architecture consists of two primary components:

    1. Buoys Mining Module (BMM): Constructs a set of reference features that possess representativeness (via Singular Value Decomposition initialization and a representation decay loss), task awareness (via cross-aggregation with task features), and resilience to domain shifts (via self-aggregation across support and query buoys).
    2. Adaptive Correlation Module (ACM): Formulates buoy-mediated pixel matching as an optimal transport problem with a cost matrix derived from buoy structural similarities, adaptively downweighting irrelevant reference features for each pixel pair.
  2. Knowl 2 — SVD-Based Buoy Initialization and Representation Decay Loss

    model/method

    To initialize KK representative reference features (buoys) that retain maximal correlation energy from the initial pairwise pixel-level correlation matrix W∈Rhw×hw\mathbf{W} \in \mathbb{R}^{hw \times hw} computed from query features FQ∈Rhw×C\mathbf{F}_Q \in \mathbb{R}^{hw \times C} and support features FS∈Rhw×C\mathbf{F}_S \in \mathbb{R}^{hw \times C}, ABCNet performs truncated Singular Value Decomposition (SVD):

    W≈UΣVT\mathbf{W} \approx \mathbf{U} \mathbf{\Sigma} \mathbf{V}^T

    where U∈Rhw×K\mathbf{U} \in \mathbb{R}^{hw \times K} contains the top-KK left singular vectors, Σ∈RK×K\mathbf{\Sigma} \in \mathbb{R}^{K \times K} contains the top-KK singular values, and V∈Rhw×K\mathbf{V} \in \mathbb{R}^{hw \times K} contains the top-KK right singular vectors. The initial query buoys Fbq∈RK×C\mathbf{F}_{bq} \in \mathbb{R}^{K \times C} and support buoys Fbs∈RK×C\mathbf{F}_{bs} \in \mathbb{R}^{K \times C} are constructed by linearly projecting the feature maps onto the singular vector bases:

    FbqT=FQTU,FbsT=FSTV\mathbf{F}_{bq}^T = \mathbf{F}_Q^T \mathbf{U}, \quad \mathbf{F}_{bs}^T = \mathbf{F}_S^T \mathbf{V}

    To preserve the representational energy of the correlation matrix throughout training and prevent buoy degradation, a representation decay loss Lrd\mathcal{L}_{rd} is added to the training objective:

    Lrd=∥WR−W∥F\mathcal{L}_{rd} = \Vert \mathbf{W}_R - \mathbf{W} \Vert_F

    where WR∈Rhw×hw\mathbf{W}_R \in \mathbb{R}^{hw \times hw} is the buoys-correlation-based matching matrix output by the Adaptive Correlation Module, and ∥⋅∥F\Vert \cdot \Vert_F denotes the Frobenius norm.

  3. Knowl 3 — Buoy Refinement via Cross-Aggregation and Self-Aggregation

    model/method

    The Buoys Mining Module (BMM) refines the initial buoy features Fb⋆=[b1,⋆,…,bK,⋆]∈RK×C\mathbf{F}_{b\star} = [\mathbf{b}_{1,\star}, \dots, \mathbf{b}_{K,\star}] \in \mathbb{R}^{K \times C} (where ⋆∈{S,Q}\star \in \{S, Q\} denotes support or query) through two successive aggregation stages:

    1. Cross-Aggregation Mechanism: To infuse task-specific context, buoys serve as queries (denoted Qi,⋆\mathbf{Q}_{i,\star}) while image features F⋆=[f1,⋆,…,fhw,⋆]∈Rhw×C\mathbf{F}_\star = [\mathbf{f}_{1,\star}, \dots, \mathbf{f}_{hw,\star}] \in \mathbb{R}^{hw \times C} serve as keys Kj,⋆\mathbf{K}_{j,\star} and values Vj,⋆\mathbf{V}_{j,\star}: Qi,⋆=bi,⋆W⋆Q,Kj,⋆=fj,⋆W⋆K,Vj,⋆=fj,⋆W⋆V\mathbf{Q}_{i,\star} = \mathbf{b}_{i,\star} \mathbf{W}_{\star}^{\mathcal{Q}}, \quad \mathbf{K}_{j,\star} = \mathbf{f}_{j,\star} \mathbf{W}_{\star}^{\mathcal{K}}, \quad \mathbf{V}_{j,\star} = \mathbf{f}_{j,\star} \mathbf{W}_{\star}^{\mathcal{V}} where W⋆Q,W⋆K∈RC×Ck\mathbf{W}_\star^{\mathcal{Q}}, \mathbf{W}_\star^{\mathcal{K}} \in \mathbb{R}^{C \times C_k} and W⋆V∈RC×Cv\mathbf{W}_\star^{\mathcal{V}} \in \mathbb{R}^{C \times C_v} are linear projection matrices. Multi-head attention is computed as: si,j=exp⁡(βi,j)∑m=1hwexp⁡(βi,m),βi,j=Qi,⋆Kj,⋆Tdk,b^i,⋆=∑j=1hwsi,jVj,⋆s_{i,j} = \frac{\exp(\beta_{i,j})}{\sum_{m=1}^{hw} \exp(\beta_{i,m})}, \quad \beta_{i,j} = \frac{\mathbf{Q}_{i,\star} \mathbf{K}_{j,\star}^T}{\sqrt{d_k}}, \quad \hat{\mathbf{b}}_{i,\star} = \sum_{j=1}^{hw} s_{i,j} \mathbf{V}_{j,\star} followed by a feed-forward network (FFN), yielding task-aware buoy representations F^bs,F^bq∈RK×C\hat{\mathbf{F}}_{bs}, \hat{\mathbf{F}}_{bq} \in \mathbb{R}^{K \times C}.

    2. Self-Aggregation Mechanism: To bridge the domain gap between support and query features and capture co-occurring foreground objects, the support and query buoys are concatenated: F^B=Concatenate⁡(F^bs,F^bq)∈R2K×C\hat{\mathbf{F}}_B = \operatorname{Concatenate}(\hat{\mathbf{F}}_{bs}, \hat{\mathbf{F}}_{bq}) \in \mathbb{R}^{2K \times C} Two multi-head self-attention layers with scaled dot-product attention followed by an FFN are applied to F^B\hat{\mathbf{F}}_B: FˉB=Softmax⁡(QBKBTdk1)VB\bar{\mathbf{F}}_B = \operatorname{Softmax}\left(\frac{\mathbf{Q}_B \mathbf{K}_B^T}{\sqrt{d_{k1}}}\right)\mathbf{V}_B producing the final contextualized reference buoys FˉB\bar{\mathbf{F}}_B.

  4. Knowl 4 — Optimal Transport-Based Adaptive Correlation Module

    model/method

    The Adaptive Correlation Module (ACM) computes the matching score between the ii-th support feature fi,S\mathbf{f}_{i,S} and the jj-th query feature fj,Q\mathbf{f}_{j,Q} using an optimal transport (OT) formulation over the refined buoys FˉB\bar{\mathbf{F}}_B.

    Let SBB∈RK×K\mathbf{S}^{BB} \in \mathbb{R}^{K \times K} denote the cosine similarity matrix between buoys. The transport cost matrix is defined as (1−SBB)(\mathbf{1} - \mathbf{S}^{BB}). The marginal distribution constraints μ∈RK\boldsymbol{\mu} \in \mathbb{R}^K and ν∈RK\boldsymbol{\nu} \in \mathbb{R}^K represent the similarity distributions of fi,S\mathbf{f}_{i,S} and fj,Q\mathbf{f}_{j,Q} over the KK buoys, respectively.

    The entropic optimal transport problem is defined as:

    min⁡T∈TTr⁡(TT(1−SBB))+ϵH(T)\min_{\mathbf{T} \in \mathcal{T}} \operatorname{Tr}\left(\mathbf{T}^T (\mathbf{1} - \mathbf{S}^{BB})\right) + \epsilon H(\mathbf{T})

    subject to the constraint set:

    T={T∈R+K×K∣T1=μ,  TT1=ν}\mathcal{T} = \left\{\mathbf{T} \in \mathbb{R}_{+}^{K \times K} \mid \mathbf{T}\mathbf{1} = \boldsymbol{\mu}, \; \mathbf{T}^T \mathbf{1} = \boldsymbol{\nu}\right\}

    where H(T)=−∑k,lTk,llog⁡Tk,lH(\mathbf{T}) = -\sum_{k,l} T_{k,l} \log T_{k,l} is the entropy regularization term and ϵ>0\epsilon > 0 controls smoothness (solved via the Sinkhorn algorithm). The optimal transport plan Ti,j∗∈RK×K\mathbf{T}^*_{i,j} \in \mathbb{R}^{K \times K} represents the pixel-pair adaptive structural buoy contribution. The final buoy-correlation matching entry WR(i,j)\mathbf{W}_R(i, j) is computed as the Frobenius inner product:

    WR(i,j)=Ti,j∗⊙SBB\mathbf{W}_R(i, j) = \mathbf{T}^*_{i,j} \odot \mathbf{S}^{BB}

    where ⊙\odot denotes the sum of elementwise products.

  5. Knowl 5 — Few-Shot Segmentation Performance on PASCAL-5i

    data/table

    The following table compares the 1-shot and 5-shot segmentation performance of three baseline architectures (PFENet, CyCTR, DCAMA) with and without ABCNet on the PASCAL-5i5^i benchmark using ResNet-50 and ResNet-101 backbones. Performance is reported in terms of mean Intersection-over-Union (mIoU, %) across 4 cross-validation folds and Foreground-Background IoU (FBIoU, %).

    Method Backbone 1-shot 5-shot
    Fold-0 Fold-1 Fold-2 Fold-3 mIoU FBIoU Fold-0 Fold-1 Fold-2 Fold-3 mIoU FBIoU
    PFENet ResNet-50 61.7 69.5 55.4 56.3 60.7 73.3 63.1 70.7 55.8 57.9 61.9 73.9
    PFENet w/ ABCNet ResNet-50 62.5 70.8 57.2 58.1 62.2 74.1 64.7 73.0 57.1 59.5 63.6 74.2
    CyCTR ResNet-50 67.8 72.8 58.0 58.0 64.2 – 71.1 73.2 60.5 57.5 65.6 –
    CyCTR w/ ABCNet ResNet-50 67.8 74.3 59.2 59.4 65.2 73.8 72.6 74.4 61.3 59.0 66.8 76.2
    DCAMA ResNet-50 67.5 72.3 59.6 59.0 64.6 75.7 70.5 73.9 63.7 65.8 68.5 79.5
    DCAMA w/ ABCNet ResNet-50 68.8 73.4 62.3 59.5 66.0 76.0 71.7 74.2 65.4 67.0 69.6 80.0
    PFENet ResNet-101 60.5 69.4 54.4 55.9 60.1 72.9 62.8 70.4 54.9 57.6 61.4 73.5
    PFENet w/ ABCNet ResNet-101 62.7 70.0 55.1 57.5 61.3 73.7 63.4 71.8 56.4 57.7 62.3 74.0
    CyCTR ResNet-101 69.3 72.7 56.5 58.6 64.3 72.9 73.5 74.0 58.6 60.2 66.6 75.0
    CyCTR w/ ABCNet ResNet-101 71.2 73.0 57.9 60.2 65.6 74.6 74.2 73.0 60.2 62.1 67.4 76.6
    DCAMA ResNet-101 65.4 71.4 63.2 58.3 64.6 77.6 70.7 73.7 66.8 61.9 68.3 80.8
    DCAMA w/ ABCNet ResNet-101 65.3 72.9 65.0 59.3 65.6 78.5 71.4 75.0 68.2 63.1 69.4 80.8

    ABCNet consistently improves mIoU across all architectures and backbones: on ResNet-50, 1-shot mIoU improves by +1.5%+1.5\% on PFENet, +1.0%+1.0\% on CyCTR, and +1.4%+1.4\% on DCAMA; 5-shot mIoU improves by +1.7%+1.7\%, +1.2%+1.2\%, and +1.1%+1.1\%, respectively.

  6. Knowl 6 — Few-Shot Segmentation Performance on COCO-20i

    data/table

    The table below details 1-shot and 5-shot segmentation results on the COCO-20i20^i dataset across 4 cross-validation splits comparing baseline models with and without ABCNet. Metrics reported are mean Intersection-over-Union (mIoU, %) and Foreground-Background IoU (FBIoU, %).

    Method Backbone 1-shot 5-shot
    Fold-0 Fold-1 Fold-2 Fold-3 mIoU FBIoU Fold-0 Fold-1 Fold-2 Fold-3 mIoU FBIoU
    PFENet ResNet-101 34.3 33.0 32.3 30.1 32.4 58.6 38.5 38.6 38.2 34.3 37.4 61.9
    PFENet w/ ABCNet ResNet-101 36.5 35.7 34.7 31.4 34.6 59.2 40.1 40.1 39.0 35.9 38.8 62.8
    CyCTR ResNet-50 38.9 43.0 39.6 40.3 40.5 – 41.1 48.9 45.2 47.0 45.6 –
    CyCTR w/ ABCNet ResNet-50 40.7 45.9 41.6 40.6 42.2 66.7 43.2 50.8 45.8 47.1 46.7 62.8
    DCAMA ResNet-50 41.9 45.1 44.4 41.7 43.3 69.5 45.9 50.5 50.7 46.0 48.3 71.7
    DCAMA w/ ABCNet ResNet-50 42.3 46.2 46.0 42.0 44.1 69.9 45.5 51.7 52.6 46.4 49.1 72.7

    ABCNet yields consistent improvements across all baselines on COCO-20i20^i: PFENet (ResNet-101) improves by +2.2%+2.2\% in 1-shot and +1.4%+1.4\% in 5-shot mIoU; CyCTR (ResNet-50) improves by +1.8%+1.8\% in 1-shot and +1.1%+1.1\% in 5-shot mIoU; and DCAMA (ResNet-50) improves by +0.8%+0.8\% in both 1-shot and 5-shot mIoU.

  7. Knowl 7 — Ablation Analysis of Buoy Mining and Adaptive Correlation Components

    data/table

    An ablation study conducted on PASCAL-5i5^i (1-shot setting, average mIoU across 4 splits) using DCAMA with a ResNet-50 backbone evaluates the progressive inclusion of each submodule in the Buoys Mining Module (BMM) and the Adaptive Correlation Module (ACM).

    Buoys Mining Module (BMM) ACM mIoU (%)
    SVD Init Cross-Agg Self-Agg RD-Loss
    ✓ 63.0
    ✓ ✓ 64.1
    ✓ ✓ ✓ 64.9
    ✓ ✓ ✓ ✓ 65.3
    ✓ ✓ ✓ ✓ ✓ 66.0

    The baseline using raw SVD-initialized buoys with simple dot-product matching achieves 63.0% mIoU (lower than original DCAMA's 64.6%, showing that unrefined buoys can hurt matching). Introducing cross-aggregation improves mIoU by +1.1%+1.1\% (to 64.1%), self-aggregation adds +0.8%+0.8\% (to 64.9%), representation decay loss adds +0.4%+0.4\% (to 65.3%), and the optimal transport-based ACM contributes an additional +0.7%+0.7\%, reaching the full ABCNet performance of 66.0% mIoU.

  8. Knowl 8 — Impact of Buoy Initialization Strategies on Segmentation Performance

    data/table

    A comparative ablation on PASCAL-5i5^i (1-shot setting, 4-split average mIoU) using DCAMA with a ResNet-50 backbone evaluates different methods for initializing buoys prior to refinement:

    Initialization Method mIoU (%)
    Random (Rand) 64.3
    Top-kk Cumulative Affinity 65.2
    Learnable Parameters 65.4
    SVD-Based Projection 66.0
    • Random (Rand): Randomly samples KK pixel features from support and query feature maps as buoys.
    • Top-kk: Sums cumulative affinities from the initial pairwise similarity matrix along support and query dimensions and selects the top-KK indices.
    • Learnable: Uses 2K2K randomly initialized, task-shared trainable buoy vectors.
    • SVD-Based: Projects support and query features onto the top-KK singular vectors of the correlation matrix.

    SVD-based initialization outperforms all alternative strategies because it retains the principal correlation energy of the feature maps while keeping the number of buoys compact.

  9. Knowl 9 — Sensitivity to Buoy Count and Representation Decay Loss Weight

    empirical result

    Hyperparameter evaluations using DCAMA with a ResNet-101 backbone on split-2 of PASCAL-5i5^i demonstrate the following sensitivities:

    1. Number of Buoys (KK): Performance was tested across K∈{8,16,24,32,40,48}K \in \{8, 16, 24, 32, 40, 48\}. The segmentation mIoU increases steadily from under 62.0% at K=8K=8 to a peak at K=24K=24 (exceeding 65.5% mIoU), before declining as KK increases further due to feature redundancy.
    2. Representation Decay Loss Weight (λrd\lambda_{rd}): Evaluating weights over the range λrd∈[0.02,0.15]\lambda_{rd} \in [0.02, 0.15] shows that performance peaks at λrd=0.10\lambda_{rd} = 0.10 (achieving approximately 65.5% mIoU). Smaller weights insufficiently penalize information loss, whereas larger weights over-constrain the buoy correlation representations.

Coverage note — Qualitative visualization figures demonstrating pixel correspondence heatmaps, prior mask refinements, and individual buoy activation areas were omitted as their conclusions are fully reflected in the quantitative performance and ablation knowls.

References

  1. 1.Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. End-to-end object detection with transformers. In European conference on computer vision, pages 213–229. Springer, 2020.
  2. 2.Liang-Chieh Chen, George Papandreou, Iasonas Kokkinos, Kevin Murphy, and Alan L Yuille. Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs. IEEE transactions on pattern analysis and machine intelligence, 40(4):834–848, 2017.
  3. 3.Liang-Chieh Chen, Yukun Zhu, George Papandreou, Florian Schroff, and Hartwig Adam. Encoder-decoder with atrous separable convolution for semantic image segmentation. In Proceedings of the European conference on computer vision (ECCV), pages 801–818, 2018.
  4. 4.Bowen Cheng, Alex Schwing, and Alexander Kirillov. Per-pixel classification is not all you need for semantic segmentation. Advances in Neural Information Processing Systems, 34, 2021.
  5. 5.Marco Cuturi. Sinkhorn distances: Lightspeed computation of optimal transport. Advances in neural information processing systems, 26, 2013.
  6. 6.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929, 2020.
  7. 7.Mark Everingham, Luc Van Gool, Christopher KI Williams, John Winn, and Andrew Zisserman. The pascal visual object classes (voc) challenge. International journal of computer vision, 88(2):303–338, 2010.
  8. 8.Jun Fu, Jing Liu, Haijie Tian, Yong Li, Yongjun Bao, Zhiwei Fang, and Hanqing Lu. Dual attention network for scene segmentation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 3146–3154, 2019.
  9. 9.Bharath Hariharan, Pablo Arbelaez, Ross Girshick, and Jitendra Malik. Simultaneous detection and segmentation. In European conference on computer vision, pages 297–312. Springer, 2014.
  10. 10.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016.
  11. 11.Sunghwan Hong, Seokju Cho, Jisu Nam, Stephen Lin, and Seungryong Kim. Cost aggregation with 4d convolutional swin transformer for few-shot segmentation. In European Conference on Computer Vision, pages 108–126. Springer, 2022.
  12. 12.Zilong Huang, Xinggang Wang, Lichao Huang, Chang Huang, Yunchao Wei, and Wenyu Liu. Ccnet: Criss-cross attention for semantic segmentation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 603–612, 2019.
  13. 13.Alexander Kirillov, Ross Girshick, Kaiming He, and Piotr Dollar. Panoptic feature pyramid networks. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 6399–6408, 2019.
  14. 14.Chunbo Lang, Gong Cheng, Binfei Tu, and Junwei Han. Learning what not to segment: A new perspective on few-shot segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8057–8067, 2022.
  15. 15.Gen Li, Varun Jampani, Laura Sevilla-Lara, Deqing Sun, Jonghyun Kim, and Joongkyu Kim. Adaptive prototype learning and allocation for few-shot segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8334–8343, 2021.
  16. 16.Guosheng Lin, Anton Milan, Chunhua Shen, and Ian Reid. Refinenet: Multi-path refinement networks for high-resolution semantic segmentation. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 1925–1934, 2017.
  17. 17.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollar, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In European conference on computer vision, pages 740–755. Springer, 2014.
  18. 18.Yuanwei Liu, Nian Liu, Xiwen Yao, and Junwei Han. Intermediate prototype mining transformer for few-shot semantic segmentation. arXiv preprint arXiv:2210.06780, 2022.
  19. 19.Yongfei Liu, Xiangyi Zhang, Songyang Zhang, and Xuming He. Part-aware prototype network for few-shot semantic segmentation. In European Conference on Computer Vision, pages 142–158. Springer, 2020.
  20. 20.Jonathan Long, Evan Shelhamer, and Trevor Darrell. Fully convolutional networks for semantic segmentation. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 3431–3440, 2015.
  21. 21.Shervin Minaee, Yuri Y Boykov, Fatih Porikli, Antonio J Plaza, Nasser Kehtarnavaz, and Demetri Terzopoulos. Image segmentation using deep learning: A survey. IEEE transactions on pattern analysis and machine intelligence, 2021.
  22. 22.Khoi Nguyen and Sinisa Todorovic. Feature weighting and boosting for few-shot segmentation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 622–631, 2019.
  23. 23.Hyeonwoo Noh, Seunghoon Hong, and Bohyung Han. Learning deconvolution network for semantic segmentation. In Proceedings of the IEEE international conference on computer vision, pages 1520–1528, 2015.
  24. 24.Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, et al. Imagenet large scale visual recognition challenge. International journal of computer vision, 115(3):211–252, 2015.
  25. 25.Amirreza Shaban, Shray Bansal, Zhen Liu, Irfan Essa, and Byron Boots. One-shot learning for semantic segmentation. arXiv preprint arXiv:1709.03410, 2017.
  26. 26.Xinyu Shi, Dong Wei, Yu Zhang, Donghuan Lu, Munan Ning, Jiashun Chen, Kai Ma, and Yefeng Zheng. Dense cross-query-and-support attention weighted mask aggregation for few-shot segmentation. In European Conference on Computer Vision, pages 151–168. Springer, 2022.
  27. 27.Robin Strudel, Ricardo Garcia, Ivan Laptev, and Cordelia Schmid. Segmenter: Transformer for semantic segmentation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 7262–7272, 2021.
  28. 28.Zhuotao Tian, Hengshuang Zhao, Michelle Shu, Zhicheng Yang, Ruiyu Li, and Jiaya Jia. Prior guided feature enrichment network for few-shot segmentation. IEEE Transactions on Pattern Analysis & Machine Intelligence, (01):1–1, 2020.
  29. 29.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in neural information processing systems, 30, 2017.
  30. 30.Haochen Wang, Xudong Zhang, Yutao Hu, Yandan Yang, Xianbin Cao, and Xiantong Zhen. Few-shot semantic segmentation with democratic attention networks. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XIII 16, pages 730–746. Springer, 2020.
  31. 31.Kaixin Wang, Jun Hao Liew, Yingtian Zou, Daquan Zhou, and Jiashi Feng. Panet: Few-shot image semantic segmentation with prototype alignment. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 9197–9206, 2019.
  32. 32.Yuan Wang, Rui Sun, Zhe Zhang, and Tianzhu Zhang. Adaptive agent transformer for few-shot segmentation. In European Conference on Computer Vision, pages 36–52. Springer, 2022.
  33. 33.Zhonghua Wu, Xiangxi Shi, Guosheng Lin, and Jianfei Cai. Learning meta-class memory for few-shot semantic segmentation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 517–526, 2021.
  34. 34.Enze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar, Jose M Alvarez, and Ping Luo. Segformer: Simple and efficient design for semantic segmentation with transformers. Advances in Neural Information Processing Systems, 34, 2021.
  35. 35.Guo-Sen Xie, Jie Liu, Huan Xiong, and Ling Shao. Scale-aware graph neural network for few-shot semantic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5475–5484, 2021.
  36. 36.Boyu Yang, Chang Liu, Bohao Li, Jianbin Jiao, and Qixiang Ye. Prototype mixture models for few-shot semantic segmentation. In European Conference on Computer Vision, pages 763–778. Springer, 2020.
  37. 37.Bingfeng Zhang, Jimin Xiao, and Terry Qin. Self-guided and cross-guided learning for few-shot segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8312–8321, 2021.
  38. 38.Chi Zhang, Guosheng Lin, Fayao Liu, Jiushuang Guo, Qingyao Wu, and Rui Yao. Pyramid graph networks with connection attentions for region-based one-shot semantic segmentation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 9587–9595, 2019.
  39. 39.Chi Zhang, Guosheng Lin, Fayao Liu, Rui Yao, and Chunhua Shen. Canet: Class-agnostic segmentation networks with iterative refinement and attentive few-shot learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5217–5226, 2019.
  40. 40.Gengwei Zhang, Guoliang Kang, Yi Yang, and Yunchao Wei. Few-shot segmentation via cycle-consistent transformer. Advances in Neural Information Processing Systems, 34:21984–21996, 2021.
  41. 41.Jian-Wei Zhang, Yifan Sun, Yi Yang, and Wei Chen. Feature-proxy transformer for few-shot segmentation. arXiv preprint arXiv:2210.06908, 2022.
  42. 42.Xiaolin Zhang, Yunchao Wei, Yi Yang, and Thomas S Huang. Sg-one: Similarity guidance network for one-shot semantic segmentation. IEEE Transactions on Cybernetics, 50(9):3855–3865, 2020.
  43. 43.Yifan Zhang, Bo Pang, and Cewu Lu. Semantic segmentation by early region proxy. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1258–1268, 2022.
  44. 44.Hengshuang Zhao, Jianping Shi, Xiaojuan Qi, Xiaogang Wang, and Jiaya Jia. Pyramid scene parsing network. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2881–2890, 2017.

Citation

MLA
Wang, Y., et al. “Rethinking the Correlation in Few-Shot Segmentation: A Buoys View”. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023, pp. 7183–92, https://doi.org/10.1109/CVPR52729.2023.00694.
APA
Wang, Y., Sun, R., & Zhang, T. (2023). Rethinking the Correlation in Few-Shot Segmentation: A Buoys View. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 7183–7192. https://doi.org/10.1109/CVPR52729.2023.00694
Chicago
Wang, Y., R. Sun, and T. Zhang. 2023. “Rethinking the Correlation in Few-Shot Segmentation: A Buoys View”. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 7183–92. https://doi.org/10.1109/CVPR52729.2023.00694.
Harvard
Wang, Y., Sun, R. and Zhang, T. (2023) “Rethinking the Correlation in Few-Shot Segmentation: A Buoys View”, 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp. 7183–7192. Available at: https://doi.org/10.1109/CVPR52729.2023.00694.
Vancouver
1. Wang Y, Sun R, Zhang T (2023) Rethinking the Correlation in Few-Shot Segmentation: A Buoys View. In: 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp 7183–7192

BibTeX

@inproceedings{Wang_2023, title={Rethinking the Correlation in Few-Shot Segmentation: A Buoys View}, url={http://dx.doi.org/10.1109/CVPR52729.2023.00694}, DOI={10.1109/cvpr52729.2023.00694}, booktitle={2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)}, publisher={IEEE}, author={Wang, Yuan and Sun, Rui and Zhang, Tianzhu}, year={2023}, month=June, pages={7183–7192} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE