Matching Feature Sets for Few-Shot Image Classification

Arman AfrasiyabiHugo LarochelleJean-François LalondeChristian Gagné

article2022CVPR115 citations

Proposes SetFeat, an approach that inserts shallow self-attention mappers into standard CNN backbones to extract sets of multi-scale feature vectors paired with set-to-set matching metrics, setting new state-of-the-art accuracy across major few-shot classification benchmarks without increasing parameter complexity.

Listen

Modern computer vision systems typically require thousands of labeled examples to recognize objects accurately. In real-world applications where data is scarce, expensive to label, or constantly evolving, systems must rely on few-shot learning—the ability to classify new categories using only one or a few reference images. Standard methods compress an entire image into a single summary vector, which often loses critical visual details needed to recognize novel classes accurately.

To overcome this limitation, the article introduces SetFeat, an approach that represents each image as a set of multiple feature vectors extracted across different network depths and spatial scales. The primary objective is to demonstrate that extracting and matching diverse feature sets, rather than relying on a single monolithic representation, significantly enhances classification accuracy on few-shot visual recognition tasks without increasing the model's total parameter count.

The authors implemented this strategy by embedding shallow, lightweight self-attention modules called mappers throughout standard convolutional network architectures, such as Conv4 and ResNet-12. To ensure fair evaluation and prevent over-parameterization, the internal convolutional filters were adjusted so that the adapted models maintained roughly the same total parameter footprint as baseline networks. The framework was trained via a two-stage process combining initial pre-training on base categories with episodic meta-training. The authors evaluated three set-matching metrics—Match-sum, Min-min, and Sum-min—across standard benchmark datasets: miniImageNet, tieredImageNet, and the fine-grained CUB dataset, evaluating both 1-shot and 5-shot scenarios across 600 test episodes.

The experimental findings show that the set-based approach consistently outperforms existing state-of-the-art methods. Among the proposed comparison functions, the Sum-min metric performed best across almost all benchmarks. In 1-shot tasks, the method achieved accuracy gains over competitive baselines of approximately 1.83% on miniImageNet, 1.42% on tieredImageNet, and 1.83% to 2.04% on CUB. Ablation studies demonstrated that simply concatenating multi-scale features into a single large vector eliminates these accuracy gains, confirming that reasoning over distinct sets of features is the key driver of performance. Visual saliency and activation analyses further verified that the individual mappers attend to diverse, complementary image regions and remain consistently active across diverse object classes.

These results indicate that organizations deploying vision systems in data-constrained environments can achieve superior recognition performance without incurring higher computational costs or deploying larger, heavier models. By preserving multifaceted visual cues—such as fine textures and distinct object parts—set-based matching reduces classification error risks in low-data regimes while maintaining operational efficiency.

Organizations evaluating few-shot visual classification pipelines should consider adopting set-based feature extraction and non-linear set-to-set matching metrics rather than conventional single-vector pooling. Future development should explore expanding beyond the fixed 10-mapper configuration within larger backbones, investigating weighted set-matching functions, and applying set-feature extraction to self-supervised learning domains.

The findings are supported by consistent results across standard benchmarks and rigorous confidence intervals. However, readers should note that the evaluation was restricted to standard 4-block convolutional and ResNet backbones using a fixed set of ten attention mappers. Additional pilot testing is advisable when transferring this architecture to entirely different visual domains or scaling it to larger transformer-based backbones.

  • Paper: Prototypical Networks for Few-shot Learning, Jake Snell et al. (2017). This seminal paper introduces metric-based prototypical representations for few-shot image classification, establishing the standard single-vector embedding baseline that SetFeat replaces with sets of features.
  • Paper: Matching Networks for One Shot Learning, Oriol Vinyals et al. (2016). It introduces episodic meta-learning and attention-based matching mechanisms for few-shot learning, providing foundational concepts behind metric-based set matching.
  • Paper: Deep Sets, Manzil Zaheer et al. (2017). It provides the foundational theoretical and architectural framework for permutation-invariant deep neural networks on sets, informing set-based representations and set-to-set matching metrics.
  • Paper: A Closer Look at Few-shot Classification, Wei-Yu Chen et al. (2019). It provides standard evaluation protocols, benchmark datasets (miniImageNet and CUB), and baselines that SetFeat utilizes to benchmark its set-based few-shot classification approach.
  • Paper: Learning to Compare: Relation Network for Few-Shot Learning, Flood Sung et al. (2017). It introduces end-to-end learnable distance metrics for few-shot classification, which motivates designing specialized metric comparisons for richer feature representations.
  • Paper: TADAM: Task dependent adaptive metric for improved few-shot learning, Boris N. Oreshkin et al. (2018). It explores task-dependent feature conditioning and metric scaling in few-shot classification, offering valuable context on adapting feature extractors for few-shot tasks.
  • Paper: Meta-Learning With Differentiable Convex Optimization, Kwonjoon Lee et al. (2019). It establishes strong baseline methodologies and benchmarks for few-shot classification using differentiable convex optimization on standard architectures.
Cover for Matching Feature Sets for Few-Shot Image Classification

Abstract

In image classification, it is common practice to train deep networks to extract a single feature vector per input image. Few-shot classification methods also mostly follow this trend. In this work, we depart from this established direction and instead propose to extract sets of feature vectors for each image. We argue that a set-based representation intrinsically builds a richer representation of images from the base classes, which can subsequently better transfer to the few-shot classes. To do so, we propose to adapt existing feature extractors to instead produce sets of feature vectors from images. Our approach, dubbed SetFeat, embeds shallow self-attention mechanisms inside existing encoder architectures. The attention modules are lightweight, and as such our method results in encoders that have approximately the same number of parameters as their original versions. During training and inference, a set-to-set matching metric is used to perform image classification. The effectiveness of our proposed architecture and metrics is demonstrated via thorough experiments on standard few-shot datasets—namely miniImageNet, tieredImageNet, and CUB—in both the 1- and 5-shot scenarios. In all cases but one, our method outperforms the state-of-the-art.

Table of Contents

  • 1. Introduction
  • 2. Related work
  • 3. Preliminaries
  • 4. Set-based few-shot image classification
  • 4.1. Set-feature extractor
  • 4.2. Set-to-set matching metrics
  • 4.3. Inference
  • 4.4. Training procedure
  • 5. Evaluation
  • 5.1. Backbones
  • 5.2. Datasets and implementation details
  • 5.3. Quantitative and comparative evaluations
  • 6. Ablation
  • 6.1. Mapper configurations
  • 6.2. Over-parameterization of SetFeat4-64
  • 6.3. Probing the activation of mappers
  • 6.4. Topm analysis
  • 6.5. Visualizing mappers saliency
  • 7. Discussion
  • References

Knowls

  1. Knowl 1 — SetFeat Set-Feature Extractor Architecture

    model/method

    SetFeat is a set-feature extraction framework designed for few-shot image classification. Departing from conventional methods that extract a single monolithic feature vector per image, SetFeat maps an input image xx to a set of MM distinct feature vectors H={hm}m=1M\mathcal{H} = \{h_m\}_{m=1}^M across multiple stages of a deep convolutional backbone.

    Given a backbone network composed of BB sequential convolutional blocks with parameters θf={θbf}b=1B\theta^f = \{\theta^f_b\}_{b=1}^B, intermediate representations zb∈RP×Dpz_b \in \mathbb{R}^{P \times D_p} are computed at each block bb, where PP is the number of 1×11 \times 1 spatial patches and DpD_p is the channel dimension. Lightweight, unit-depth self-attention modules called mappers g(⋅∣θmg)g(\cdot \mid \theta_m^g) are embedded after convolutional blocks. Each mapper operates independently to convert its corresponding intermediate representation zbmz_{b_m} into a feature embedding hm∈RDh_m \in \mathbb{R}^D. To facilitate set-to-set matching, every mapper projects its output to the identical channel dimension DD equal to the output dimension of the final convolutional block of the backbone.

  2. Knowl 2 — Attention-Based Feature Mapper Formulation

    equation

    Let zbm∈RP×Dpz_{b_m} \in \mathbb{R}^{P \times D_p} denote the intermediate feature map extracted after block bmb_m of a convolutional backbone, treated as PP spatial patches of dimension DpD_p. The mm-th feature mapper g(zbm∣θmg)g(z_{b_m} \mid \theta_m^g) uses single-head self-attention parameterized by query q(⋅∣θmq)q(\cdot \mid \theta_m^q), key k(⋅∣θmk)k(\cdot \mid \theta_m^k), and value v(⋅∣θmv)v(\cdot \mid \theta_m^v) projections.

    The spatial attention score matrix βm∈RP×P\beta_m \in \mathbb{R}^{P \times P} is computed across patches as:

    βm=Softmax(q(zbm∣θmq) k(zbm∣θmk)⊤dk)\beta_m = \text{Softmax}\left(\frac{q(z_{b_m} \mid \theta_m^q) \, k(z_{b_m} \mid \theta_m^k)^\top}{\sqrt{d_k}}\right)

    where dk\sqrt{d_k} is the scaling factor. The attended representation am∈RP×Daa_m \in \mathbb{R}^{P \times D_a} is then obtained by:

    am=βm v(zbm∣θmv)a_m = \beta_m \, v(z_{b_m} \mid \theta_m^v)

    For ResNet backbones, a residual connection is added to form am+zbma_m + z_{b_m} (employing a 1×11 \times 1 convolution if Da≠DpD_a \neq D_p). The final feature vector hm∈RDah_m \in \mathbb{R}^{D_a} is computed by average pooling over all PP spatial patches:

    hm=1P∑p=1P(am+zbm)ph_m = \frac{1}{P} \sum_{p=1}^P (a_m + z_{b_m})_p

  3. Knowl 3 — Set-to-Set Matching Metrics

    equation

    In NN-way KK-shot image classification, let xqx_q be a query image with extracted feature set {hi(xq)}i=1M\{h_i(x_q)\}_{i=1}^M and let Sn={(xin,yi=n)}i=1KS^n = \{(x_i^n, y_i = n)\}_{i=1}^K be the support set for class nn. The class prototype centroid for mapper mm is defined as hˉm(Sn)=1∣Sn∣∑x∈Snhm(x)\bar{h}_m(S^n) = \frac{1}{|S^n|} \sum_{x \in S^n} h_m(x). Using the negative cosine similarity as the base distance metric d(u,v)=−u⊤v∥u∥2∥v∥2d(u, v) = -\frac{u^\top v}{\|u\|_2 \|v\|_2}, distances between query xqx_q and class support set SnS^n are computed via three set-to-set formulations:

    1. Match-sum (dmsd_{\text{ms}}): Aggregates distances between identical mapper indices:

    dms(xq,Sn)=∑i=1Md(hi(xq),hˉi(Sn))d_{\text{ms}}(x_q, S^n) = \sum_{i=1}^M d\left(h_i(x_q), \bar{h}_i(S^n)\right)

    1. Min-min (dmmd_{\text{mm}}): Finds the global minimum distance across all query and support centroid mapper pairs:

    dmm(xq,Sn)=min⁡i=1Mmin⁡j=1Md(hi(xq),hˉj(Sn))d_{\text{mm}}(x_q, S^n) = \min_{i=1}^M \min_{j=1}^M d\left(h_i(x_q), \bar{h}_j(S^n)\right)

    1. Sum-min (dsmd_{\text{sm}}): For each query mapper, finds the minimum distance to any support centroid mapper and sums these minimums:

    dsm(xq,Sn)=∑i=1Mmin⁡j=1Md(hi(xq),hˉj(Sn))d_{\text{sm}}(x_q, S^n) = \sum_{i=1}^M \min_{j=1}^M d\left(h_i(x_q), \bar{h}_j(S^n)\right)

  4. Knowl 4 — Set-Based Prototypical Inference

    equation

    Given a query sample xqx_q, an NN-way support set S=⋃n=1NSnS = \bigcup_{n=1}^N S^n, and a set-to-set metric dset∈{dms,dmm,dsm}d_{\text{set}} \in \{d_{\text{ms}}, d_{\text{mm}}, d_{\text{sm}}\}, the posterior probability that query xqx_q belongs to class n∈{1,…,N}n \in \{1, \dots, N\} is defined using a softmax over the negative set-to-set distances:

    p(y=n∣xq,S)=exp⁡(−dset(xq,Sn))∑j=1Nexp⁡(−dset(xq,Sj))p(y = n \mid x_q, S) = \frac{\exp\left(-d_{\text{set}}(x_q, S^n)\right)}{\sum_{j=1}^N \exp\left(-d_{\text{set}}(x_q, S^j)\right)}

    The predicted class label y^q\hat{y}_q is chosen by finding the support class that minimizes the set-to-set distance:

    y^q=arg⁡min⁡n∈{1,…,N}dset(xq,Sn)\hat{y}_q = \arg\min_{n \in \{1, \dots, N\}} d_{\text{set}}(x_q, S^n)

  5. Knowl 5 — SetFeat Pre-Training and Episodic Meta-Training Algorithm

    algorithm

    SetFeat is trained in two sequential stages:

    1. Standard Multi-Mapper Pre-Training: For each mapper m∈{1,…,M}m \in \{1, \dots, M\}, an auxiliary fully-connected linear layer omo_m outputs logits over the CC base classes. Over a mini-batch Xbatch\mathcal{X}_{\text{batch}}, all mappers are optimized jointly using the multi-head cross-entropy loss:

    ℓpre=−∑xi∈Xbatch∑m=1Mlog⁡exp⁡(om,yi(hm,i))∑c=1Cexp⁡(om,c(hm,i))\ell_{\text{pre}} = -\sum_{x_i \in \mathcal{X}_{\text{batch}}} \sum_{m=1}^M \log \frac{\exp(o_{m, y_i}(h_{m, i}))}{\sum_{c=1}^C \exp(o_{m, c}(h_{m, i}))}

    1. Episodic Meta-Training: The linear classification heads omo_m are discarded, and the network parameters θ={θf,θg}\theta = \{\theta^f, \theta^g\} are fine-tuned on sampled few-shot episodes using set-to-set distance matching:
    Input: Backbone parameters θf\theta^f, mapper parameters θg\theta^g, episodic training set Xtrain\mathcal{X}_{\text{train}}, validation set Xvalid\mathcal{X}_{\text{valid}}, maximum epochs tmax⁡t_{\max}
    Output: Best model parameters θbest={θbestf,θbestg}\theta_{\text{best}} = \{\theta^f_{\text{best}}, \theta^g_{\text{best}}\}
    Evalidbest←∞E_{\text{valid}}^{\text{best}} \leftarrow \infty
    θ←{θf,θg}\theta \leftarrow \{\theta^f, \theta^g\}
    for t=1,…,tmax⁡t = 1, \dots, t_{\max} do
        for (xq,S)∈Xtrain(x_q, S) \in \mathcal{X}_{\text{train}} do
            ℓt←−log⁡p(yq∣xq,S)\ell^t \leftarrow -\log p(y_q \mid x_q, S)
            Update θ\theta via gradient descent on ℓt\ell^t
        end for
        for (xq,S)∈Xvalid(x_q, S) \in \mathcal{X}_{\text{valid}} do
            y^q←arg⁡min⁡Sn∈Sdset(xq,Sn)\hat{y}_q \leftarrow \arg\min_{S^n \in S} d_{\text{set}}(x_q, S^n)
        end for
        Evalid←1∣Xvalid∣∑(xq,S)∈Xvalidℓ0−1(y^q,yq)E_{\text{valid}} \leftarrow \frac{1}{|\mathcal{X}_{\text{valid}}|} \sum_{(x_q, S) \in \mathcal{X}_{\text{valid}}} \ell_{0-1}(\hat{y}_q, y_q)
        if Evalid<EvalidbestE_{\text{valid}} < E_{\text{valid}}^{\text{best}} then
            Evalidbest←EvalidE_{\text{valid}}^{\text{best}} \leftarrow E_{\text{valid}}
            θbest←θ\theta_{\text{best}} \leftarrow \theta
        end if
    end for
    return θbest\theta_{\text{best}}
  6. Knowl 6 — Backbone Parameter Parity Design

    model/method

    To ensure that performance improvements are due to set-based multi-scale feature matching rather than network over-parameterization, the convolutional kernel counts of the backbones are reduced so that the adapted SetFeat models match the parameter budgets of standard backbones:

    • Conv4-512 (1.591M parameters): Standard filter channels are 96/128/256/512. The adapted SetFeat4-512 reduces filter channels to 96/128/160/200, resulting in 1.583M total parameters including its 10 mappers (mapper feature dimension D=200D = 200).
    • ResNet12 (12.424M parameters): Standard filter channels are 64/160/320/640. The adapted SetFeat12 adjusts filter channels to 128/150/180/512, resulting in 12.349M total parameters including its 10 mappers (mapper feature dimension D=512D = 512).
    • ResNet18 (11.511M parameters): SetFeat12∗^* adjusts filter channels to 128/150/196/480, resulting in 11.466M total parameters (mapper feature dimension D=480D = 480).
    • Conv4-64 (0.113M parameters): SetFeat4-64 employs lightweight fully connected attention layers yielding 0.238M parameters, as channel reduction in Conv4-64 causes training collapse.
  7. Knowl 7 — Few-Shot Classification Accuracy on miniImageNet

    data/table

    The table below presents 5-way 1-shot and 5-shot classification accuracy (mean ±\pm 95% confidence intervals over 600 episodes) on miniImageNet across Conv4-64, Conv4-512, and ResNet12 backbones.

    Method Backbone 1-shot 5-shot
    ProtoNet Conv4-64 49.42 ±\pm 0.78 68.20 ±\pm 0.66
    MAML Conv4-64 48.07 ±\pm 1.75 63.15 ±\pm 0.91
    FEAT Conv4-64 55.15 ±\pm 0.20 71.61 ±\pm 0.16
    MELR Conv4-64 55.35 ±\pm 0.43 72.27 ±\pm 0.35
    SetFeat (Match-sum) SF4-64 55.74 ±\pm 0.65 72.18 ±\pm 0.70
    SetFeat (Min-min) SF4-64 56.22 ±\pm 0.89 72.70 ±\pm 0.65
    SetFeat (Sum-min) SF4-64 57.18 ±\pm 0.89 73.67 ±\pm 0.71
    ProtoNet Conv4-512 53.52 ±\pm 0.43 73.34 ±\pm 0.36
    PN+rot Conv4-512 56.02 ±\pm 0.46 74.00 ±\pm 0.35
    MELR Conv4-512 57.54 ±\pm 0.44 74.37 ±\pm 0.34
    SetFeat (Match-sum) SF4-512 56.50 ±\pm 0.85 72.69 ±\pm 0.68
    SetFeat (Min-min) SF4-512 58.57 ±\pm 0.87 73.46 ±\pm 0.68
    SetFeat (Sum-min) SF4-512 59.10 ±\pm 0.87 74.97 ±\pm 0.66
    MetaOptNet ResNet12 62.64 ±\pm 0.61 78.63 ±\pm 0.46
    DeepEMD ResNet12 65.91 ±\pm 0.82 82.41 ±\pm 0.56
    FEAT ResNet12 66.78 82.05
    DMF ResNet12 67.76 ±\pm 0.46 82.71 ±\pm 0.31
    MELR ResNet12 67.40 ±\pm 0.43 83.40 ±\pm 0.28
    SetFeat (Match-sum) SF-12 67.41 ±\pm 0.64 81.79 ±\pm 0.55
    SetFeat (Min-min) SF-12 67.88 ±\pm 0.55 82.07 ±\pm 0.61
    SetFeat (Sum-min) SF-12 68.32 ±\pm 0.62 82.71 ±\pm 0.46

    The sum-min metric consistently yields the highest accuracy among the set metrics, outperforming previous methods by 1.83% (1-shot on Conv4-64), 1.56% (1-shot on Conv4-512), and 0.56% (1-shot on ResNet12) over the respective competitive baselines.

  8. Knowl 8 — Few-Shot Classification Accuracy on tieredImageNet and CUB

    data/table

    The table below summarizes 5-way 1-shot and 5-shot classification accuracy (with 95% confidence intervals over 600 episodes) on tieredImageNet and CUB datasets.

    Dataset Method Backbone 1-shot 5-shot
    tieredImageNet FEAT ResNet12 70.80 ±\pm 0.23 84.79 ±\pm 0.16
    tieredImageNet DeepEMD ResNet12 71.16 ±\pm 0.87 86.03 ±\pm 0.58
    tieredImageNet DMF ResNet12 71.89 ±\pm 0.52 85.96 ±\pm 0.35
    tieredImageNet MELR ResNet12 72.14 ±\pm 0.51 87.01 ±\pm 0.35
    tieredImageNet Distill ResNet12 72.21 ±\pm 0.90 87.08 ±\pm 0.58
    tieredImageNet SetFeat (Sum-min) SF12 73.63 ±\pm 0.88 87.59 ±\pm 0.57
    CUB ProtoNet Conv4-64 64.42 ±\pm 0.48 81.82 ±\pm 0.35
    CUB FEAT Conv4-64 68.87 ±\pm 0.22 82.90 ±\pm 0.15
    CUB MELR Conv4-64 70.26 ±\pm 0.50 85.01 ±\pm 0.32
    CUB SetFeat (Sum-min) SF4-64 72.09 ±\pm 0.92 87.05 ±\pm 0.58
    CUB ProtoNet ResNet18 71.88 ±\pm 0.9 86.64 ±\pm 0.5
    CUB Neg-Margin ResNet18 72.66 ±\pm 0.9 89.40 ±\pm 0.4
    CUB MixtFSL ResNet18 73.94 ±\pm 1.1 86.01 ±\pm 0.5
    CUB SetFeat (Sum-min) SF12∗^* 79.60 ±\pm 0.80 90.48 ±\pm 0.44

    On tieredImageNet, SetFeat (Sum-min) achieves 73.63% in 1-shot (a 1.42% gain over Distill) and 87.59% in 5-shot (a 0.51% gain). On fine-grained CUB classification, SetFeat achieves 72.09% in 1-shot on SF4-64 (+1.83% over MELR) and 79.60% in 1-shot on SF12∗^* (+5.66% over MixtFSL).

  9. Knowl 9 — Ablation of Mapper Placement and Set Matching vs Vector Concatenation

    empirical result

    An ablation study on the miniImageNet validation set evaluates mapper distributions across the 4 convolutional blocks (denoted as mappers per block b1-b2-b3-b4b_1\text{-}b_2\text{-}b_3\text{-}b_4) and compares set-based representation against flat vector concatenation using the sum-min metric:

    • End-only placement (0-0-0-10): Placing all 10 mappers at the final block achieves 52.90% (1-shot) / 69.49% (5-shot) on SetFeat4-64 and 55.36% / 71.59% on SetFeat4-512.
    • Uniform single-mapper (1-1-1-1, 4 mappers): Achieves 51.11% / 69.41% on SetFeat4-64 and 53.57% / 71.60% on SetFeat4-512.
    • Equal distribution (2-2-3-3, 10 mappers): Achieves 54.73% / 71.98% on SetFeat4-64 and 56.29% / 74.74% on SetFeat4-512.
    • Progressive growth (1-2-3-4, 10 mappers): Achieves 54.71% / 71.35% on SetFeat4-64 and 58.74% / 75.30% on SetFeat4-512.
    • Concatenation baseline (1-2-3-4 concat): Concatenating the outputs of all 10 mappers into a single monolithic feature vector yields 53.56% (1-shot) and 71.82% (5-shot) on SetFeat4-64, failing to improve over the ProtoNet Conv4-512 baseline (53.51% / 71.57%).

    These results demonstrate that multi-scale feature extraction provides substantial benefits only when paired with set-based matching rather than single vector concatenation, and that progressive distribution across depth outperforms placing mappers solely at the final stage.

  10. Knowl 10 — Ablation of Top-m Mapper Summation in Set Matching

    empirical result

    Evaluating top-mm mapper distance summation on the CUB validation set shows a monotonic improvement in few-shot classification accuracy as the number of aggregated mappers mm increases from 1 to 10:

    • top-1 (min-min): 70.15% (1-shot) / 84.94% (5-shot) on SetFeat4; 78.51% (1-shot) / 89.73% (5-shot) on SetFeat12∗^*.
    • top-2: 70.84% / 85.30% on SetFeat4; 77.92% / 89.87% on SetFeat12∗^*.
    • top-4: 70.34% / 85.95% on SetFeat4; 78.37% / 89.78% on SetFeat12∗^*.
    • top-8: 71.47% / 86.88% on SetFeat4; 79.56% / 90.03% on SetFeat12∗^*.
    • top-10 (sum-min): 72.09% / 87.05% on SetFeat4; 79.60% / 90.48% on SetFeat12∗^*.

    Summing minimum distances across all 10 mappers yields superior accuracy over taking the minimum over a subset, confirming that each mapper across the network captures distinct, complementary visual attributes.

Coverage note — Qualitative gradient saliency visualizations and t-SNE embedding plots were omitted as they serve as visual illustrations supporting the quantitative empirical findings.

References

  1. 1.Arman Afrasiyabi, Jean-Francois Lalonde, and Christian Gagne. Associative alignment for few-shot image classification. In European Conference on Computer Vision, 2020. 2
  2. 2.Arman Afrasiyabi, Jean-Francois Lalonde, and Christian Gagne. Mixture-based feature space learning for few-shot image classification. International Conference on Computer Vision, 2021. 5, 6
  3. 3.Kelsey R Allen, Evan Shelhamer, Hanul Shin, and Joshua B Tenenbaum. Infinite mixture prototypes for few-shot learning. International Conference on Machine Learning, PMLR, 2019. 5
  4. 4.Antreas Antoniou, Harrison Edwards, and Amos Storkey. How to train your maml. International Conference on Learning Representations, 2019. 2
  5. 5.Luca Bertinetto, Joao F. Henriques, Philip Torr, and Andrea Vedaldi. Meta-learning with differentiable closed-form solvers. In International Conference on Learning Representations, 2019. 2
  6. 6.Malik Boudiaf, Ziko Imtiaz Masud, Jerome Rony, Jose Dolz, Pablo Piantanida, and Ismail Ben Ayed. Transductive information maximization for few-shot learning. Neural Information Processing Systems, 2020. 2
  7. 7.Qi Cai, Yingwei Pan, Ting Yao, Chenggang Yan, and Tao Mei. Memory matching networks for one-shot image recognition. In Conference on Computer Vision and Pattern Recognition, 2018. 5
  8. 8.Wei-Yu Chen, Yen-Cheng Liu, Zsolt Kira, Yu-Chiang Frank Wang, and Jia-Bin Huang. A closer look at few-shot classification. International Conference on Learning Representations, 2019. 2, 5, 6, 8
  9. 9.Yinbo Chen, Zhuang Liu, Huijuan Xu, Trevor Darrell, and Xiaolong Wang. Meta-baseline: Exploring simple meta-learning for few-shot learning. In International Conference on Computer Vision, 2021. 5
  10. 10.Zitian Chen, Yanwei Fu, Yu-Xiong Wang, Lin Ma, Wei Liu, and Martial Hebert. Image deformation meta-networks for one-shot learning. In Conference on Computer Vision and Pattern Recognition, 2019. 2
  11. 11.Wen-Hsuan Chu, Yu-Jhe Li, Jing-Cheng Chang, and Yu-Chiang Frank Wang. Spot and learn: A maximum-entropy patch sampler for few-shot image classification. In Conference on Computer Vision and Pattern Recognition, 2019. 2
  12. 12.Guneet S Dhillon, Pratik Chaudhari, Avinash Ravichandran, and Stefano Soatto. A baseline for few-shot image classification. International Conference on Learning Representations, 2020. 2
  13. 13.Carl Doersch, Ankush Gupta, and Andrew Zisserman. Crosstransformers: spatially-aware few-shot transfer. In Neural Information Processing Systems, 2021. 2
  14. 14.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. International Conference on Learning Representations, 2020. 2, 3
  15. 15.Nikita Dvornik, Cordelia Schmid, and Julien Mairal. Diversity with cooperation: Ensemble methods for few-shot classification. In International Conference on Computer Vision, 2019. 6
  16. 16.Nanyi Fei, Zhiwu Lu, Tao Xiang, and Songfang Huang. Melr: Meta-learning via modeling episode-level relationships for few-shot learning. In International Conference on Learning Representations, 2021. 5, 6
  17. 17.Chelsea Finn, Pieter Abbeel, and Sergey Levine. Model-agnostic meta-learning for fast adaptation of deep networks. In International Conference on Machine Learning, 2017. 2, 5, 6
  18. 18.Chelsea Finn, Kelvin Xu, and Sergey Levine. Probabilistic model-agnostic meta-learning. In Neural Information Processing Systems, 2018. 5
  19. 19.Hang Gao, Zheng Shou, Alireza Zareian, Hanwang Zhang, and Shih-Fu Chang. Low-shot learning via covariance-preserving adversarial augmentation networks. In Neural Information Processing Systems, 2018. 2
  20. 20.Victor Garcia and Joan Bruna. Few-shot learning with graph neural networks. International Conference on Learning Representations, 2018. 2
  21. 21.Spyros Gidaris, Andrei Bursuc, Nikos Komodakis, Patrick Perez, and Matthieu Cord. Boosting few-shot visual learning with self-supervision. In International Conference on Computer Vision, 2019. 2, 5
  22. 22.Spyros Gidaris and Nikos Komodakis. Dynamic few-shot visual learning without forgetting. In Conference on Computer Vision and Pattern Recognition, 2018. 2, 5
  23. 23.Spyros Gidaris and Nikos Komodakis. Generating classification weights with gnn denoising autoencoders for few-shot learning. Conference on Computer Vision and Pattern Recognition, 2019. 2
  24. 24.Benjamin Graham, Alaaeldin El-Nouby, Hugo Touvron, Pierre Stock, Armand Joulin, Herve Jegou, and Matthijs Douze. Levit: A vision transformer in convnet's clothing for faster inference. In International Conference on Computer Vision (ICCV), 2021. 2
  25. 25.Bharath Hariharan and Ross Girshick. Low-shot visual recognition by shrinking and hallucinating features. In International Conference on Computer Vision, 2017. 2
  26. 26.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Conference on Computer Vision and Pattern Recognition, 2016. 3, 5
  27. 27.Xiaopeng Hong, Hong Chang, Shiguang Shan, Xilin Chen, and Wen Gao. Sigma set: A small second order statistical region descriptor. In Conference on Computer Vision and Pattern Recognition, 2009. 2
  28. 28.Ruibing Hou, Hong Chang, Bingpeng MA, Shiguang Shan, and Xilin Chen. Cross attention network for few-shot classification. In Neural Information Processing Systems, 2019. 2
  29. 29.Junsik Kim, Tae-Hyun Oh, Seokju Lee, Fei Pan, and In So Kweon. Variational prototyping-encoder: One-shot learning with prototypical images. In Conference on Computer Vision and Pattern Recognition, 2019. 2
  30. 30.Svetlana Lazebnik, Cordelia Schmid, and Jean Ponce. Beyond bags of features: Spatial pyramid matching for recognizing natural scene categories. In 2006 IEEE computer society conference on computer vision and pattern recognition (CVPR'06), 2006. 2
  31. 31.Kwonjoon Lee, Subhransu Maji, Avinash Ravichandran, and Stefano Soatto. Meta-learning with differentiable convex optimization. In Conference on Computer Vision and Pattern Recognition, 2019. 5, 6
  32. 32.Wenbin Li, Lei Wang, Jinglin Xu, Jing Huo, Yang Gao, and Jiebo Luo. Revisiting local descriptor based image-to-class measure for few-shot learning. In Conference on Computer Vision and Pattern Recognition, 2019. 2
  33. 33.Yann Lifchitz, Yannis Avrithis, Sylvaine Picard, and Andrei Bursuc. Dense classification and implanting for few-shot learning. In Conference on Computer Vision and Pattern Recognition, 2019. 2
  34. 34.Tsung-Yi Lin, Piotr Dollar, Ross Girshick, Kaiming He, Bharath Hariharan, and Serge Belongie. Feature pyramid networks for object detection. In Conference on Computer Vision and Pattern Recognition, 2017. 1, 2
  35. 35.Bin Liu, Yue Cao, Yutong Lin, Qi Li, Zheng Zhang, Mingsheng Long, and Han Hu. Negative margin matters: Understanding margin in few-shot classification. In European Conference on Computer Vision, 2020. 5, 6
  36. 36.Bin Liu, Zhirong Wu, Han Hu, and Stephen Lin. Deep metric transfer for label propagation with limited annotated data. In International Conference on Computer Vision Workshops, 2019. 2
  37. 37.Akshay Mehrotra and Ambedkar Dukkipati. Generative adversarial residual pairwise networks for one shot learning. arXiv preprint arXiv:1703.08033, 2017. 2
  38. 38.Tsendsuren Munkhdalai, Xingdi Yuan, Soroush Mehri, and Adam Trischler. Rapid adaptation with conditionally shifted neurons. In International Conference on Machine Learning, 2018. 5
  39. 39.Jaehoon Oh, Hyungjun Yoo, ChangHwan Kim, and Se-Young Yun. Boil: Towards representation change for few-shot learning. In International Conference on Learning Representations, 2021. 5
  40. 40.Boris Oreshkin, Pau Rodrıguez Lopez, and Alexandre Lacoste. Tadam: Task dependent adaptive metric for improved few-shot learning. In Neural Information Processing Systems, 2018. 2, 5
  41. 41.Hang Qi, Matthew Brown, and David G. Lowe. Low-shot learning with imprinted weights. In Conference on Computer Vision and Pattern Recognition, 2018. 2
  42. 42.Pedro Quelhas, Florent Monay, J-M Odobez, Daniel Gatica-Perez, Tinne Tuytelaars, and Luc Van Gool. Modeling scenes with local descriptors and latent aspects. In International Conference on Computer Vision, 2005. 2
  43. 43.Sachin Ravi and Hugo Larochelle. Optimization as a model for few-shot learning. In International Conference on Learning Representations, 2016. 2
  44. 44.Mengye Ren, Eleni Triantafillou, Sachin Ravi, Jake Snell, Kevin Swersky, Joshua B Tenenbaum, Hugo Larochelle, and Richard S Zemel. Meta-learning for semi-supervised few-shot classification. International Conference on Learning Representations, 2018. 2, 5
  45. 45.Mamshad Nayeem Rizve, Salman Khan, Fahad Shahbaz Khan, and Mubarak Shah. Exploring complementary strengths of invariant and equivariant representations for few-shot learning. In Conference on Computer Vision and Pattern Recognition, 2021. 2, 6
  46. 46.Eli Schwartz, Leonid Karlinsky, Joseph Shtok, Sivan Harary, Mattias Marder, Abhishek Kumar, Rogerio Feris, Raja Giryes, and Alex Bronstein. Delta-encoder: an effective sample synthesis method for few-shot object recognition. In Neural Information Processing Systems, 2018. 2
  47. 47.Christian Simon, Piotr Koniusz, Richard Nock, and Mehrtash Harandi. Adaptive subspaces for few-shot learning. In Conference on Computer Vision and Pattern Recognition, 2020. 6
  48. 48.Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman. Deep inside convolutional networks: Visualising image classification models and saliency maps. In In Workshop at International Conference on Learning Representations. Citeseer, 2014. 8
  49. 49.Jake Snell, Kevin Swersky, and Richard Zemel. Prototypical networks for few-shot learning. In Neural Information Processing Systems, 2017. 1, 2, 4, 5, 6, 7
  50. 50.Jong-Chyi Su, Subhransu Maji, and Bharath Hariharan. When does self-supervision improve few-shot learning? In European Conference on Computer Vision, 2020. 2
  51. 51.Qianru Sun, Yaoyao Liu, Tat-Seng Chua, and Bernt Schiele. Meta-transfer learning for few-shot learning. In Conference on Computer Vision and Pattern Recognition, 2019. 6
  52. 52.Flood Sung, Yongxin Yang, Li Zhang, Tao Xiang, Philip H.S. Torr, and Timothy M. Hospedales. Learning to compare: Relation network for few-shot learning. In Conference on Computer Vision and Pattern Recognition, 2018. 2, 5, 6
  53. 53.Kaihua Tang, Jianqiang Huang, and Hanwang Zhang. Long-tailed classification by keeping the good and removing the bad momentum causal effect. Neural Information Processing Systems, 2020. 6
  54. 54.Yonglong Tian, Yue Wang, Dilip Krishnan, Joshua B Tenenbaum, and Phillip Isola. Few-shot image classification: a good embedding is all you need. In European Conference on Computer Vision, 2020. 2, 5, 6
  55. 55.Hung-Yu Tseng, Hsin-Ying Lee, Jia-Bin Huang, and Ming-Hsuan Yang. Cross-domain few-shot classification via learned feature-wise transformation. International Conference on Learning Representations, 2020. 2
  56. 56.Laurens Van der Maaten and Geoffrey Hinton. Visualizing data using t-sne. Journal of machine learning research, 2008. 7
  57. 57.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Advances in neural information processing systems, 2017. 2, 3
  58. 58.Oriol Vinyals, Charles Blundell, Timothy Lillicrap, Daan Wierstra, et al. Matching networks for one shot learning. In Neural Information Processing Systems, 2016. 2, 3, 4, 5, 6
  59. 59.Catherine Wah, Steve Branson, Peter Welinder, Pietro Perona, and Serge Belongie. The Caltech-UCSD birds-200-2011 dataset, 2011. 5
  60. 60.Yu-Xiong Wang, Ross Girshick, Martial Hebert, and Bharath Hariharan. Low-shot learning from imaginary data. In Conference on Computer Vision and Pattern Recognition, 2018. 2
  61. 61.Yu-Xiong Wang and Martial Hebert. Learning from small sample sets by combining unsupervised meta-training with cnns. In Neural Information Processing Systems, 2016. 2
  62. 62.Davis Wertheimer and Bharath Hariharan. Few-shot learning with localization in realistic settings. In Conference on Computer Vision and Pattern Recognition, 2019. 2
  63. 63.Haiping Wu, Bin Xiao, Noel Codella, Mengchen Liu, Xiyang Dai, Lu Yuan, and Lei Zhang. Cvt: Introducing convolutions to vision transformers. arXiv preprint arXiv:2103.15808, 2021. 2, 5
  64. 64.Ziyang Wu, Yuwei Li, Lihua Guo, and Kui Jia. Parn: Position-aware relation networks for few-shot learning. In International Conference on Computer Vision, 2019. 5
  65. 65.Chengming Xu, Yanwei Fu, Chen Liu, Chengjie Wang, Jilin Li, Feiyue Huang, Li Zhang, and Xiangyang Xue. Learning dynamic alignment via meta-filter for few-shot learning. In Conference on Computer Vision and Pattern Recognition, 2021. 2, 4, 5, 6
  66. 66.Weijian Xu, Huaijin Wang, Zhuowen Tu, et al. Attentional constellation nets for few-shot learning. In International Conference on Learning Representations, 2020. 2
  67. 67.Han-Jia Ye, Hexiang Hu, De-Chuan Zhan, and Fei Sha. Few-shot learning via embedding adaptation with set-to-set functions. In Conference on Computer Vision and Pattern Recognition, 2020. 2, 4, 5, 6
  68. 68.Jaesik Yoon, Taesup Kim, Ousmane Dia, Sungwoong Kim, Yoshua Bengio, and Sungjin Ahn. Bayesian model-agnostic meta-learning. In Conference on Neural Information Processing Systems, 2018. 2
  69. 69.Sung Whan Yoon, Jun Seo, and Jaekyun Moon. Tapnet: Neural network augmented with task-adaptive projection for few-shot learning. International Conference on Machine Learning, 2019. 2, 6
  70. 70.Zhongjie Yu, Lin Chen, Zhongwei Cheng, and Jiebo Luo. Transmatch: A transfer-learning scheme for semi-supervised few-shot learning. In Conference on Computer Vision and Pattern Recognition, 2020. 2
  71. 71.Manzil Zaheer, Satwik Kottur, Siamak Ravanbakhsh, Barnabas Poczos, Ruslan Salakhutdinov, and Alexander Smola. Deep sets. In Advances in neural information processing systems, 2017. 2, 8
  72. 72.Baoquan Zhang, Xutao Li, Yunming Ye, Zhichao Huang, and Lisai Zhang. Prototype completion with primitive knowledge for few-shot learning. In Conference on Computer Vision and Pattern Recognition, 2021. 2
  73. 73.Chi Zhang, Yujun Cai, Guosheng Lin, and Chunhua Shen. Deepemd: Few-shot image classification with differentiable earth mover's distance and structured classifiers. In Conference on Computer Vision and Pattern Recognition, 2020. 2, 4, 5, 6
  74. 74.Hongyi Zhang, Moustapha Cisse, Yann N Dauphin, and David Lopez-Paz. mixup: Beyond empirical risk minimization. International Conference on Learning Representations, 2018. 2
  75. 75.Hongguang Zhang, Jing Zhang, and Piotr Koniusz. Few-shot learning via saliency-guided hallucination of samples. In Conference on Computer Vision and Pattern Recognition, 2019. 2
  76. 76.Jian Zhang, Chenglong Zhao, Bingbing Ni, Minghao Xu, and Xiaokang Yang. Variational few-shot learning. In International Conference on Computer Vision, 2019. 2
  77. 77.Manli Zhang, Jianhong Zhang, Zhiwu Lu, Tao Xiang, Mingyu Ding, and Songfang Huang. Iept: Instance-level and episode-level pretext tasks for few-shot learning. In International Conference on Learning Representations, 2020. 2

Citation

MLA
Afrasiyabi, A., et al. “Matching Feature Sets for Few-Shot Image Classification”. arXiv, 2022, http://arxiv.org/abs/2204.00949v1.
APA
Afrasiyabi, A., Larochelle, H., Lalonde, J.-F., & Gagné, C. (2022). Matching Feature Sets for Few-Shot Image Classification. arXiv. http://arxiv.org/abs/2204.00949v1
Chicago
Afrasiyabi, A., H. Larochelle, J.-F. Lalonde, and C. Gagné. 2022. “Matching Feature Sets for Few-Shot Image Classification”. arXiv. http://arxiv.org/abs/2204.00949v1.
Harvard
Afrasiyabi, A. et al. (2022) “Matching Feature Sets for Few-Shot Image Classification”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2204.00949v1.
Vancouver
1. Afrasiyabi A, Larochelle H, Lalonde J-F, Gagné C (2022) Matching Feature Sets for Few-Shot Image Classification. arXiv

BibTeX

@article{afrasiyabi2022matching,
  title = {Matching Feature Sets for Few-Shot Image Classification},
  author = {Afrasiyabi, Arman and Larochelle, Hugo and Lalonde, Jean-François and Gagné, Christian},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2204.00949v1},
  eprint = {2204.00949}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE