A Closer Look at Few-shot Classification

Wei-Yu ChenYen-Cheng LiuZsolt KiraYu-Chiang Frank WangJia-Bin Huang

article2019ICLR2,028 citations

Demonstrates through unified benchmarking that standard fine-tuning baselines with deeper backbones match or outperform complex meta-learning methods in few-shot classification, especially in cross-domain settings.

Listen

Modern computer vision models typically require thousands of labeled examples per category to achieve high accuracy, making data acquisition expensive and limiting deployment in data-scarce environments. To address this, few-shot classification aims to recognize novel categories using only a handful of labeled training examples. However, progress in this field has been obscured by inconsistent implementation details across studies, under-optimized standard baselines, and artificial testing setups that evaluate novel classes drawn from the exact same domain as the training data.

The article sets out to establish a rigorous, unified benchmark to fairly evaluate prominent few-shot algorithms against standard transfer learning baselines. It specifically investigates the impact of neural network depth and assesses model performance under realistic cross-domain shifts where test categories come from entirely different image distributions.

To conduct this evaluation, the researchers standardized experimental pipelines across four representative meta-learning algorithms—MatchingNet, ProtoNet, RelationNet, and MAML—alongside a standard transfer baseline and an enhanced variant, Baseline++, which uses cosine distance to minimize intra-class variation. The authors tested these methods across multiple standard datasets (mini-ImageNet and CUB-200-2011), varied the feature extractor depth from shallow four-layer networks to 34-layer residual networks, and introduced a cross-domain evaluation setting by training models on generic objects and testing them on fine-grained bird species.

The study revealed three major findings. First, adopting deeper neural network backbones significantly reduces the performance gap among all algorithms; when using deeper backbones on datasets with limited domain shifts, simple baselines perform on par with or better than complex meta-learning algorithms. Second, reducing intra-class variation via distance-based classification allows Baseline++ to achieve state-of-the-art accuracy under shallow architectures, improving the baseline 5-shot accuracy on fine-grained data from 64.16% to 79.34%. Third, under realistic cross-domain shifts, the standard baseline outperforms all evaluated meta-learning techniques, achieving 65.57% 5-shot accuracy compared to MAML's 51.34% and MatchingNet's 53.07%, because complex meta-learners overfit their training domain and struggle to adapt to distinct visual distributions.

These findings suggest that the perceived advantages of complex meta-learning frameworks are largely an artifact of using shallow architectures and domain-matched benchmarks. For practical engineering and deployment, organizations can achieve high accuracy and reduce model complexity, training overhead, and maintenance risks by using standard pre-training and fine-tuning rather than specialized meta-learning frameworks.

For future development, practitioners should prioritize standard transfer learning pipelines with deeper feature backbones as default solutions for few-shot tasks. When deploying across shifting data distributions, fine-tuning classifier layers on available target samples is recommended over fixed meta-learning inference. Researchers should refocus efforts on designing meta-learning algorithms that explicitly learn to adapt across domains rather than optimizing within static benchmarks.

The conclusions are drawn with high confidence under standard benchmark conditions across hundreds of standardized evaluation runs. However, readers should note that the cross-domain findings are evaluated primarily on image classification tasks with specific source-target pairs, meaning further validation on larger-scale industry datasets and broader multimodal tasks is advised before completely generalizing the results.

  • Paper: Prototypical Networks for Few-shot Learning, Jake Snell et al. (2017). This seminal work introduces Prototypical Networks, providing the foundational metric-based meta-learning formulation and standard benchmark evaluation protocols analyzed and re-evaluated in the source paper.
  • Paper: Matching Networks for One Shot Learning, Oriol Vinyals et al. (2016). This paper establishes the episodic training framework and the miniImageNet benchmark that form the basis of the few-shot learning comparison in the source paper.
  • Paper: Learning to Compare: Relation Network for Few-Shot Learning, Flood Sung et al. (2017). This work introduces Relation Networks for non-linear metric learning in few-shot tasks, serving as one of the primary representative meta-learning architectures evaluated in the source paper's benchmark.
  • Paper: Optimization as a Model for Few-Shot Learning, S. Ravi et al. (2017). This paper proposes an LSTM-based meta-learner optimization framework for few-shot learning, defining key optimization concepts that the source paper compares against simple fine-tuning baselines.
  • Paper: On First-Order Meta-Learning Algorithms, Alex Nichol et al. (2018). This study analyzes first-order gradient-based meta-learning algorithms, providing crucial context for why standard optimization and fine-tuning can compete with complex meta-learners.
  • Paper: Deep Residual Learning for Image Recognition, Kaiming He et al. (2016). This work introduces deep residual networks, which serve as the deeper backbone architectures that the source paper investigates to show performance convergence across few-shot methods.
  • Paper: Generalizing from a Few Examples, Yaqing Wang et al. (2019). This comprehensive survey contextualizes few-shot learning and incorporates insights on cross-domain transfer and empirical risk minimization that build upon the findings of the source paper.
  • Paper: Supervised Contrastive Learning, Prannay Khosla et al. (2020). This work extends representation learning by using supervised contrastive objectives to reduce intra-class variation and improve cross-domain transfer, directly relating to the feature properties highlighted in the source paper.
  • Paper: Visual Prompt Tuning, Menglin Jia et al. (2022). This paper advances parameter-efficient transfer and fine-tuning for vision backbones on low-data benchmarks, extending the baseline fine-tuning insights explored in the source paper.
  • Paper: Big Self-Supervised Models are Strong Semi-Supervised Learners, Ting Chen et al. (2020). This study demonstrates that scaling backbone capacity combined with simple pre-training and fine-tuning outperforms specialized low-data techniques, taking the backbone scaling conclusions of the source paper to larger models.
Cover for A Closer Look at Few-shot Classification

Abstract

Few-shot classification aims to learn a classifier to recognize unseen classes during training with limited labeled examples. While significant progress has been made, the growing complexity of network designs, meta-learning algorithms, and differences in implementation details make a fair comparison difficult. In this paper, we present 1) a consistent comparative analysis of several representative few-shot classification algorithms, with results showing that deeper backbones significantly reduce the performance differences among methods on datasets with limited domain differences, 2) a modified baseline method that surprisingly achieves competitive performance when compared with the state-of-the-art on both the \miniI and the CUB datasets, and 3) a new experimental setting for evaluating the cross-domain generalization ability for few-shot classification algorithms. Our results reveal that reducing intra-class variation is an important factor when the feature backbone is shallow, but not as critical when using deeper backbones. In a realistic cross-domain evaluation setting, we show that a baseline method with a standard fine-tuning practice compares favorably against other state-of-the-art few-shot learning algorithms.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Overview of Few-shot classification Algorithms
  • 3.1 Baseline
  • 3.2 Baseline++
  • 3.3 Meta-learning algorithms
  • 4 Experimental Results
  • 4.1 Experimental setup
  • 4.2 Evaluation using the Standard Setting
  • 4.3 Effect of increasing the network depth
  • 4.4 Effect of domain differences between base and novel classes
  • 4.5 Effect of further adaptation
  • 5 Conclusions
  • References
  • A1 Relationship between domain adaptation and few-shot classification
  • A2 Terminology difference
  • A3 Additional results on Omniglot and Omniglot→\rightarrowEMNIST
  • A4 Baseline with 1-NN classifier
  • A5 MAML and MAML with first-order approximation
  • A6 Intra-class variation and backbone depth
  • A7 Detailed statistics in effects of increasing backbone depth
  • A8 More-way in meta-testing stage

Knowls

  1. Knowl 1 — Baseline and Baseline++ Transfer Learning Formulations for Few-Shot Classification

    model/method

    The standard few-shot transfer learning framework comprises two training stages:

    1. Base Training Stage: Given a set of base class labeled images XbX_b over cc classes, a feature extractor fθ:X→Rdf_\theta: \mathcal{X} \to \mathbb{R}^d with parameters θ\theta and a base classifier C(⋅∣Wb)C(\cdot \mid W_b) parameterized by weight matrix Wb∈Rd×cW_b \in \mathbb{R}^{d \times c} are trained from scratch using standard cross-entropy loss LpredL_{\text{pred}}.

      • In the Baseline model, the classifier is a standard linear layer followed by softmax: y^i=σ(WbTfθ(xi))\hat{y}_i = \sigma(W_b^T f_\theta(x_i)) where σ(⋅)\sigma(\cdot) is the softmax function and xi∈Xbx_i \in X_b.
      • In Baseline++, the linear classifier is replaced by a distance-based prototype classifier designed to reduce intra-class variation. The classifier weight matrix is written as Wb=[w1,w2,…,wc]∈Rd×cW_b = [w_1, w_2, \dots, w_c] \in \mathbb{R}^{d \times c}, where each vector wj∈Rdw_j \in \mathbb{R}^d acts as a learned class prototype. For an input feature fθ(xi)f_\theta(x_i), class similarity scores si,js_{i,j} are computed via cosine similarity: si,j=fθ(xi)Twj∥fθ(xi)∥2∥wj∥2s_{i,j} = \frac{f_\theta(x_i)^T w_j}{\|f_\theta(x_i)\|_2 \|w_j\|_2} The similarity scores are multiplied by a class-wise learnable scalar to adjust the range [−1,1][-1, 1] before being passed to the softmax function σ([si,1,…,si,c]T)\sigma([s_{i,1}, \dots, s_{i,c}]^T).
    2. Fine-Tuning Stage: When novel classes XnX_n with limited support examples are provided, the feature extractor parameters θ\theta are fixed, and a new classifier C(⋅∣Wn)C(\cdot \mid W_n) is initialized and trained for 100 gradient steps with batch size 4 on the support set using cross-entropy loss (linear classifier for Baseline, cosine similarity prototype classifier for Baseline++).

  2. Knowl 2 — Evaluation of Few-Shot Algorithms on mini-ImageNet and CUB under Standard Conv-4 Backbone

    data/table

    Under a standard 5-way 1-shot and 5-shot setting using a four-layer convolutional backbone (Conv-4) with data augmentation (random cropping, horizontal flips, and color jitter), Baseline++ achieves performance competitive with or superior to meta-learning methods on both generic object recognition (mini-ImageNet) and fine-grained classification (CUB-200-2011). Furthermore, adding data augmentation to the Baseline model substantially improves its accuracy compared to earlier literature where Baseline was trained without augmentation.

    Method CUB mini-ImageNet
    1-shot 5-shot 1-shot 5-shot
    Baseline 47.12±0.74%47.12 \pm 0.74\% 64.16±0.71%64.16 \pm 0.71\% 42.11±0.71%42.11 \pm 0.71\% 62.53±0.69%62.53 \pm 0.69\%
    Baseline++ 60.53±0.83%60.53 \pm 0.83\% 79.34±0.61%79.34 \pm 0.61\% 48.24±0.75%48.24 \pm 0.75\% 66.43±0.63%66.43 \pm 0.63\%
    MatchingNet 60.52±0.88%60.52 \pm 0.88\% 75.29±0.75%75.29 \pm 0.75\% 48.14±0.78%48.14 \pm 0.78\% 63.48±0.66%63.48 \pm 0.66\%
    ProtoNet 50.46±0.88%50.46 \pm 0.88\% 76.39±0.64%76.39 \pm 0.64\% 44.42±0.84%44.42 \pm 0.84\% 64.24±0.72%64.24 \pm 0.72\%
    MAML 54.73±0.97%54.73 \pm 0.97\% 75.75±0.76%75.75 \pm 0.76\% 46.47±0.82%46.47 \pm 0.82\% 62.71±0.71%62.71 \pm 0.71\%
    RelationNet 62.34±0.94%62.34 \pm 0.94\% 77.84±0.68%77.84 \pm 0.68\% 49.31±0.85%49.31 \pm 0.85\% 66.60±0.69%66.60 \pm 0.69\%

    All results report the mean classification accuracy and 95% confidence intervals averaged over 600 randomly generated 5-way test episodes. On CUB, Baseline++ improves over Baseline by 13.41% in 1-shot and 15.18% in 5-shot, achieving top performance. On mini-ImageNet, Baseline++ outperforms or matches all evaluated meta-learning baselines in the 5-shot setting.

  3. Knowl 3 — Diminishing Performance Gaps and Reduced Intra-Class Variation with Deeper Backbones

    data/table

    Increasing the depth of the feature backbone (Conv-4, Conv-6, ResNet-10, ResNet-18, and ResNet-34) reduces intra-class feature variation (measured by the Davies-Bouldin index) and narrows the performance gaps between transfer learning baselines and meta-learning methods.

    Setting Method Conv-4 Conv-6 ResNet-10 ResNet-18 ResNet-34
    CUB 1-shot Baseline 47.12±0.74%47.12 \pm 0.74\% 55.77±0.86%55.77 \pm 0.86\% 63.34±0.91%63.34 \pm 0.91\% 65.51±0.87%65.51 \pm 0.87\% 67.96±0.89%67.96 \pm 0.89\%
    Baseline++ 60.53±0.83%60.53 \pm 0.83\% 66.00±0.89%66.00 \pm 0.89\% 69.55±0.89%69.55 \pm 0.89\% 67.02±0.90%67.02 \pm 0.90\% 68.00±0.83%68.00 \pm 0.83\%
    MatchingNet 60.52±0.88%60.52 \pm 0.88\% 66.47±0.93%66.47 \pm 0.93\% 71.29±0.87%71.29 \pm 0.87\% 73.49±0.89%73.49 \pm 0.89\% 73.49±0.89%73.49 \pm 0.89\%
    ProtoNet 50.46±0.88%50.46 \pm 0.88\% 66.36±1.00%66.36 \pm 1.00\% 73.22±0.92%73.22 \pm 0.92\% 72.99±0.88%72.99 \pm 0.88\% 72.94±0.91%72.94 \pm 0.91\%
    MAML 54.73±0.97%54.73 \pm 0.97\% 66.26±1.05%66.26 \pm 1.05\% 70.32±0.99%70.32 \pm 0.99\% 68.42±1.07%68.42 \pm 1.07\% 67.28±1.08%67.28 \pm 1.08\%
    RelationNet 62.34±0.94%62.34 \pm 0.94\% 64.38±0.94%64.38 \pm 0.94\% 70.47±0.99%70.47 \pm 0.99\% 68.58±0.94%68.58 \pm 0.94\% 69.72±0.98%69.72 \pm 0.98\%
    CUB 5-shot Baseline 64.16±0.71%64.16 \pm 0.71\% 73.07±0.71%73.07 \pm 0.71\% 81.27±0.57%81.27 \pm 0.57\% 82.85±0.55%82.85 \pm 0.55\% 84.27±0.53%84.27 \pm 0.53\%
    Baseline++ 79.34±0.61%79.34 \pm 0.61\% 82.02±0.55%82.02 \pm 0.55\% 85.17±0.50%85.17 \pm 0.50\% 83.58±0.54%83.58 \pm 0.54\% 84.50±0.51%84.50 \pm 0.51\%
    MatchingNet 75.29±0.75%75.29 \pm 0.75\% 77.92±0.65%77.92 \pm 0.65\% 83.47±0.58%83.47 \pm 0.58\% 84.45±0.58%84.45 \pm 0.58\% 86.51±0.52%86.51 \pm 0.52\%
    ProtoNet 76.39±0.64%76.39 \pm 0.64\% 82.03±0.59%82.03 \pm 0.59\% 85.01±0.52%85.01 \pm 0.52\% 86.64±0.51%86.64 \pm 0.51\% 87.86±0.47%87.86 \pm 0.47\%
    MAML 75.75±0.76%75.75 \pm 0.76\% 78.82±0.70%78.82 \pm 0.70\% 80.93±0.71%80.93 \pm 0.71\% 83.47±0.62%83.47 \pm 0.62\% 83.47±0.59%83.47 \pm 0.59\%
    RelationNet 77.84±0.68%77.84 \pm 0.68\% 80.16±0.64%80.16 \pm 0.64\% 83.70±0.55%83.70 \pm 0.55\% 84.05±0.56%84.05 \pm 0.56\% 83.18±0.54%83.18 \pm 0.54\%
    mini-Img 1-shot Baseline 42.11±0.71%42.11 \pm 0.71\% 45.82±0.74%45.82 \pm 0.74\% 52.37±0.79%52.37 \pm 0.79\% 51.75±0.80%51.75 \pm 0.80\% 49.82±0.73%49.82 \pm 0.73\%
    Baseline++ 48.24±0.75%48.24 \pm 0.75\% 48.29±0.72%48.29 \pm 0.72\% 53.97±0.79%53.97 \pm 0.79\% 51.87±0.77%51.87 \pm 0.77\% 52.65±0.83%52.65 \pm 0.83\%
    MatchingNet 48.14±0.78%48.14 \pm 0.78\% 50.47±0.86%50.47 \pm 0.86\% 54.49±0.81%54.49 \pm 0.81\% 52.91±0.88%52.91 \pm 0.88\% 53.20±0.78%53.20 \pm 0.78\%
    ProtoNet 44.42±0.84%44.42 \pm 0.84\% 50.37±0.83%50.37 \pm 0.83\% 51.98±0.84%51.98 \pm 0.84\% 54.16±0.82%54.16 \pm 0.82\% 53.90±0.83%53.90 \pm 0.83\%
    MAML 46.47±0.82%46.47 \pm 0.82\% 50.96±0.92%50.96 \pm 0.92\% 54.69±0.89%54.69 \pm 0.89\% 49.61±0.92%49.61 \pm 0.92\% 51.46±0.90%51.46 \pm 0.90\%
    RelationNet 49.31±0.85%49.31 \pm 0.85\% 51.84±0.88%51.84 \pm 0.88\% 52.19±0.83%52.19 \pm 0.83\% 52.48±0.86%52.48 \pm 0.86\% 51.74±0.83%51.74 \pm 0.83\%
    mini-Img 5-shot Baseline 62.53±0.69%62.53 \pm 0.69\% 66.42±0.67%66.42 \pm 0.67\% 74.69±0.64%74.69 \pm 0.64\% 74.27±0.63%74.27 \pm 0.63\% 73.45±0.65%73.45 \pm 0.65\%
    Baseline++ 66.43±0.63%66.43 \pm 0.63\% 68.09±0.69%68.09 \pm 0.69\% 75.90±0.61%75.90 \pm 0.61\% 75.68±0.63%75.68 \pm 0.63\% 76.16±0.63%76.16 \pm 0.63\%
    MatchingNet 63.48±0.66%63.48 \pm 0.66\% 63.19±0.70%63.19 \pm 0.70\% 68.82±0.65%68.82 \pm 0.65\% 68.88±0.69%68.88 \pm 0.69\% 68.32±0.66%68.32 \pm 0.66\%
    ProtoNet 64.24±0.72%64.24 \pm 0.72\% 67.33±0.67%67.33 \pm 0.67\% 72.64±0.64%72.64 \pm 0.64\% 73.68±0.65%73.68 \pm 0.65\% 74.65±0.64%74.65 \pm 0.64\%
    MAML 62.71±0.71%62.71 \pm 0.71\% 66.09±0.71%66.09 \pm 0.71\% 66.62±0.83%66.62 \pm 0.83\% 65.72±0.77%65.72 \pm 0.77\% 65.90±0.79%65.90 \pm 0.79\%
    RelationNet 66.60±0.69%66.60 \pm 0.69\% 64.55±0.70%64.55 \pm 0.70\% 70.20±0.66%70.20 \pm 0.66\% 69.83±0.68%69.83 \pm 0.68\% 69.61±0.67%69.61 \pm 0.67\%

    On the CUB dataset, methods that explicitly minimize intra-class variation (Baseline++, RelationNet, MatchingNet) show large advantages with a Conv-4 backbone, but this gap diminishes when deeper ResNet backbones are used because deeper representations naturally reduce intra-class variation. On mini-ImageNet 5-shot, Baseline and Baseline++ outperform all meta-learning methods when using ResNet-10, ResNet-18, and ResNet-34 backbones.

  4. Knowl 4 — Cross-Domain Few-Shot Evaluation: mini-ImageNet to CUB Benchmark

    empirical result

    In a cross-domain few-shot setting where base classes are sampled from a generic domain (64 base classes from mini-ImageNet) and novel testing classes are sampled from a fine-grained domain (50 novel classes from CUB-200-2011), transfer learning via the Baseline model outperforms all meta-learning algorithms.

    Evaluating 5-way 5-shot classification with a ResNet-18 feature backbone yields the following results:

    Method mini-ImageNet →\to CUB (5-shot)
    Baseline 65.57±0.70%65.57 \pm 0.70\%
    Baseline++ 62.04±0.76%62.04 \pm 0.76\%
    ProtoNet 62.02±0.70%62.02 \pm 0.70\%
    RelationNet 57.71±0.73%57.71 \pm 0.73\%
    MatchingNet 53.07±0.74%53.07 \pm 0.74\%
    MAML 51.34±0.72%51.34 \pm 0.72\%

    Meta-learning methods learn to condition on support sets under the assumption that support and query sets originate from the same task distribution encountered during meta-training. When a substantial domain shift occurs between base and novel categories, these conditioned models fail to generalize. In contrast, the Baseline method trains a new classifier layer directly on the novel support instances, enabling superior adaptation to domain shifts. Furthermore, Baseline outperforms Baseline++ under domain shift (65.57%65.57\% vs 62.04%62.04\%), suggesting that forcing intra-class variation reduction on base classes can compromise the transferability of learned representations to distant domains.

  5. Knowl 5 — Effect of Test-Time Adaptation on Meta-Learning Models under Domain Shift

    empirical result

    Augmenting meta-learning algorithms with test-time adaptation strategies on novel support instances improves their cross-domain performance:

    1. MatchingNet and ProtoNet: Fix the pre-trained feature extractor and train a new softmax classifier layer on the novel support set via cross-entropy optimization for 100 iterations.
    2. MAML: Extend support set gradient updates from a few inner-loop steps to 100 gradient updates on the classification layer.
    3. RelationNet: Split the novel support set into 3 support and 2 query instances per class and fine-tune the relation module for 100 epochs.

    Under cross-domain evaluation (mini-ImageNet →\to CUB with ResNet-18, 5-shot), test-time adaptation substantially increases accuracy for MatchingNet (from 53.07%53.07\% to over 60%60\%) and MAML (from 51.34%51.34\% to over 60%60\%), demonstrating that lack of domain adaptation is the primary reason meta-learning models lag behind the Baseline. However, when base and novel classes come from the same domain (e.g., CUB within-domain), test-time adaptation provides minimal improvement and can degrade ProtoNet accuracy due to optimization mismatch with meta-training.

  6. Knowl 6 — Performance of Few-Shot Methods on Mismatched N-Way Tasks at Test Time

    data/table

    When models meta-trained on 5-way tasks are evaluated on 10-way and 20-way classification at test time on mini-ImageNet (5-shot), Baseline++ demonstrates superior scaling compared to both Baseline and distance metric meta-learning methods across both Conv-4 and ResNet-18 backbones.

    Method Conv-4 ResNet-18
    5-way 10-way 20-way 5-way 10-way 20-way
    Baseline 62.53±0.69%62.53 \pm 0.69\% 46.44±0.41%46.44 \pm 0.41\% 32.27±0.24%32.27 \pm 0.24\% 74.27±0.63%74.27 \pm 0.63\% 55.00±0.46%55.00 \pm 0.46\% 42.03±0.25%42.03 \pm 0.25\%
    Baseline++ 66.43±0.63%66.43 \pm 0.63\% 52.26±0.40%52.26 \pm 0.40\% 38.03±0.24%38.03 \pm 0.24\% 75.68±0.63%75.68 \pm 0.63\% 63.40±0.44%63.40 \pm 0.44\% 50.85±0.25%50.85 \pm 0.25\%
    MatchingNet 63.48±0.66%63.48 \pm 0.66\% 47.61±0.44%47.61 \pm 0.44\% 33.97±0.24%33.97 \pm 0.24\% 68.88±0.69%68.88 \pm 0.69\% 52.27±0.46%52.27 \pm 0.46\% 36.78±0.25%36.78 \pm 0.25\%
    ProtoNet 64.24±0.68%64.24 \pm 0.68\% 48.77±0.45%48.77 \pm 0.45\% 34.58±0.23%34.58 \pm 0.23\% 73.68±0.65%73.68 \pm 0.65\% 59.22±0.44%59.22 \pm 0.44\% 44.96±0.26%44.96 \pm 0.26\%
    RelationNet 66.60±0.69%66.60 \pm 0.69\% 47.77±0.43%47.77 \pm 0.43\% 33.72±0.22%33.72 \pm 0.22\% 69.83±0.68%69.83 \pm 0.68\% 53.88±0.48%53.88 \pm 0.48\% 39.17±0.25%39.17 \pm 0.25\%

    MAML cannot be directly tested in this setting because its architecture hard-codes the classifier output dimension to the meta-training NN-way. Baseline++ maintains high performance at 10-way (63.40%63.40\%) and 20-way (50.85%50.85\%) with ResNet-18 because cosine prototype learning reduces intra-class dispersion without requiring large episodic batch sizes during training.

  7. Knowl 7 — Character Recognition Benchmarks on Omniglot and Omniglot-to-EMNIST

    data/table

    In character recognition experiments using a Conv-4 backbone without data augmentation, meta-learning methods outperform Baseline and Baseline++ in the 1-shot setting, but all approaches achieve comparable accuracy in the 5-shot setting on both within-domain (Omniglot) and cross-domain (Omniglot →\to EMNIST) benchmarks.

    Method Omniglot Omniglot →\to EMNIST
    1-shot 5-shot 1-shot 5-shot
    Baseline 94.89±0.45%94.89 \pm 0.45\% 99.12±0.13%99.12 \pm 0.13\% 63.94±0.87%63.94 \pm 0.87\% 86.00±0.59%86.00 \pm 0.59\%
    Baseline++ 95.41±0.39%95.41 \pm 0.39\% 99.38±0.10%99.38 \pm 0.10\% 64.74±0.82%64.74 \pm 0.82\% 87.31±0.58%87.31 \pm 0.58\%
    MatchingNet 97.78±0.30%97.78 \pm 0.30\% 99.37±0.11%99.37 \pm 0.11\% 72.71±0.79%72.71 \pm 0.79\% 87.60±0.56%87.60 \pm 0.56\%
    ProtoNet 98.01±0.30%98.01 \pm 0.30\% 99.15±0.12%99.15 \pm 0.12\% 70.43±0.80%70.43 \pm 0.80\% 87.04±0.55%87.04 \pm 0.55\%
    MAML 98.57±0.19%98.57 \pm 0.19\% 99.53±0.08%99.53 \pm 0.08\% 72.04±0.83%72.04 \pm 0.83\% 88.24±0.56%88.24 \pm 0.56\%
    RelationNet 97.22±0.33%97.22 \pm 0.33\% 99.30±0.10%99.30 \pm 0.10\% 75.55±0.87%75.55 \pm 0.87\% 88.94±0.54%88.94 \pm 0.54\%

    The 1-shot performance gap between meta-learning and baseline methods on character datasets arises because Omniglot characters (binary, centered, rotation-sensitive) preclude standard data augmentation during base training, causing transfer baselines to overfit base characters after very few epochs (selected at 5 epochs). When 5 labeled examples per class are available, the effect of overfitting is substantially reduced.

  8. Knowl 8 — Comparison of 1-Nearest-Neighbor and Softmax Classifiers for Baseline Adaptation

    empirical result

    When adapting a pre-trained feature extractor to novel classes on 5-way mini-ImageNet with a Conv-4 backbone, the choice between a 1-Nearest-Neighbor (1-NN) classifier using cosine distance and a gradient-trained softmax classifier presents a clear shot-dependent trade-off:

    Method 1-shot 5-shot
    Softmax 1-NN Softmax 1-NN
    Baseline 42.11±0.71%42.11 \pm 0.71\% 44.18±0.69%44.18 \pm 0.69\% 62.53±0.69%62.53 \pm 0.69\% 56.68±0.67%56.68 \pm 0.67\%
    Baseline++ 48.24±0.75%48.24 \pm 0.75\% 49.57±0.73%49.57 \pm 0.73\% 66.43±0.63%66.43 \pm 0.63\% 61.93±0.65%61.93 \pm 0.65\%

    In the 1-shot setting, a 1-NN cosine classifier outperforms fine-tuning a softmax layer (44.18%44.18\% vs 42.11%42.11\% for Baseline; 49.57%49.57\% vs 48.24%48.24\% for Baseline++). In the 5-shot setting, fine-tuning a softmax classifier via cross-entropy optimization for 100 gradient steps achieves substantially higher accuracy than 1-NN (62.53%62.53\% vs 56.68%56.68\% for Baseline; 66.43%66.43\% vs 61.93%61.93\% for Baseline++).

  9. Knowl 9 — Equivalence of First-Order Approximation and Full Second-Order MAML

    empirical result

    In Model-Agnostic Meta-Learning (MAML), applying a first-order gradient approximation by omitting second-order Hessian-vector product computations in the meta-optimization update yields asymptotic validation and test classification accuracy identical to full second-order MAML. On 5-way 5-shot Omniglot classification with a Conv-4 backbone, full second-order MAML converges in fewer meta-training episodes, but both first-order and second-order formulations reach the exact same final validation accuracy (≈99.5%≈ 99.5\%) while first-order MAML achieves substantial GPU memory savings.

Coverage note — Hallucination-based few-shot methods were deliberately omitted from the comparative study by the authors and thus are not included as knowls. All other core methodological contributions, comparative experiments, cross-domain evaluations, network depth analyses, multi-way scaling results, and ablation studies have been covered.

References

  1. 1.Antreas Antoniou, Amos Storkey, and Harrison Edwards. Data augmentation generative adversarial networks. In Proceedings of the International Conference on Learning Representations Workshops (ICLR Workshops), 2018. 1, 3
  2. 2.Luca Bertinetto, João F Henriques, Philip HS Torr, and Andrea Vedaldi. Meta-learning with differentiable closed-form solvers. In Proceedings of the International Conference on Learning Representations (ICLR), 2019. 3
  3. 3.Gregory Cohen, Saeed Afshar, Jonathan Tapson, and André van Schaik. Emnist: an extension of mnist to handwritten letters. arXiv preprint arXiv:1702.05373, 2017. 14
  4. 4.David L Davies and Donald W Bouldin. A cluster separation measure. IEEE Transactions on Pattern Analysis and Machine Intelligence, 1979. 15
  5. 5.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2009. 6
  6. 6.Nanqing Dong and Eric P Xing. Domain adaption in one-shot learning. In Joint European Conference on Machine Learning and Knowledge Discovery in Databases. Springer, 2018. 3, 13, 14
  7. 7.Chelsea Finn, Pieter Abbeel, and Sergey Levine. Model-agnostic meta-learning for fast adaptation of deep networks. In Proceedings of the International Conference on Machine Learning (ICML), 2017. 1, 2, 5, 7, 13
  8. 8.Chelsea Finn, Kelvin Xu, and Sergey Levine. Probabilistic model-agnostic meta-learning. In Advances in Neural Information Processing Systems (NIPS), 2018. 2
  9. 9.Yaroslav Ganin and Victor Lempitsky. Unsupervised domain adaptation by backpropagation. In Proceedings of the International Conference on Machine Learning (ICML), 2015. 3
  10. 10.Victor Garcia and Joan Bruna. Few-shot learning with graph neural networks. In Proceedings of the International Conference on Learning Representations (ICLR), 2018. 1, 3
  11. 11.Spyros Gidaris and Nikos Komodakis. Dynamic few-shot visual learning without forgetting. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018. 1, 2, 3, 4
  12. 12.Bharath Hariharan and Ross Girshick. Low-shot visual recognition by shrinking and hallucinating features. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), 2017. 1, 3
  13. 13.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016. 8
  14. 14.Nathan Hilliard, Lawrence Phillips, Scott Howland, Artëm Yankov, Courtney D Corley, and Nathan O Hodas. Few-shot learning with metric-agnostic conditional embeddings. arXiv preprint arXiv:1802.04376, 2018. 6
  15. 15.Yen-Chang Hsu, Zhaoyang Lv, and Zsolt Kira. Learning to cluster in order to transfer across domains and tasks. 2018. 3, 13
  16. 16.Junlin Hu, Jiwen Lu, and Yap-Peng Tan. Deep transfer metric learning. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2015. 4
  17. 17.Gregory Koch, Richard Zemel, and Ruslan Salakhutdinov. Siamese neural networks for one-shot image recognition. In Proceedings of the International Conference on Machine Learning Workshops (ICML Workshops), 2015. 2
  18. 18.Brenden Lake, Ruslan Salakhutdinov, Jason Gross, and Joshua Tenenbaum. One shot learning of simple visual concepts. In Cogsci, 2011. 14
  19. 19.Thomas Mensink, Jakob Verbeek, Florent Perronnin, and Gabriela Csurka. Metric learning for large scale image classification: Generalizing to new classes at near-zero cost. In Proceedings of the European Conference on Computer Vision (ECCV), 2012. 4
  20. 20.George A Miller. Wordnet: a lexical database for english. Communications of the ACM, 1995. 8
  21. 21.Saeid Motiian, Quinn Jones, Seyed Iranmanesh, and Gianfranco Doretto. Few-shot adversarial domain adaptation. In Advances in Neural Information Processing Systems (NIPS), 2017. 13
  22. 22.Tsendsuren Munkhdalai and Hong Yu. Meta networks. In Proceedings of the International Conference on Machine Learning (ICML), 2017. 2
  23. 23.Alex Nichol and John Schulman. Reptile: a scalable metalearning algorithm. arXiv preprint arXiv:1803.02999, 2018. 2
  24. 24.Sinno Jialin Pan, Qiang Yang, et al. A survey on transfer learning. IEEE Transactions on Knowledge and Data Engineering (TKDE), 2010. 3
  25. 25.Hang Qi, Matthew Brown, and David G Lowe. Low-shot learning with imprinted weights. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018. 1, 2, 3, 4
  26. 26.Sachin Ravi and Hugo Larochelle. Optimization as a model for few-shot learning. In Proceedings of the International Conference on Learning Representations (ICLR), 2017. 1, 2, 6, 13, 14
  27. 27.Andrei A Rusu, Dushyant Rao, Jakub Sygnowski, Oriol Vinyals, Razvan Pascanu, Simon Osindero, and Raia Hadsell. Meta-learning with latent embedding optimization. In Proceedings of the International Conference on Learning Representations (ICLR), 2019. 2
  28. 28.Jake Snell, Kevin Swersky, and Richard Zemel. Prototypical networks for few-shot learning. In Advances in Neural Information Processing Systems (NIPS), 2017. 1, 3, 4, 5, 6, 7, 13, 14, 17
  29. 29.Flood Sung, Yongxin Yang, Li Zhang, Tao Xiang, Philip HS Torr, and Timothy M Hospedales. Learning to compare: Relation network for few-shot learning. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018. 1, 3, 5, 7
  30. 30.Oriol Vinyals, Charles Blundell, Tim Lillicrap, Daan Wierstra, et al. Matching networks for one shot learning. In Advances in Neural Information Processing Systems (NIPS), 2016. 1, 3, 4, 5, 6, 7, 8, 13, 14
  31. 31.Catherine Wah, Steve Branson, Peter Welinder, Pietro Perona, and Serge Belongie. The caltech-ucsd birds-200-2011 dataset. 2011. 6
  32. 32.Yu-Xiong Wang, Ross Girshick, Martial Hebert, and Bharath Hariharan. Low-shot learning from imaginary data. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018. 1, 3, 13

Citation

MLA
Chen, W.-Y., et al. “A Closer Look at Few-shot Classification”. arXiv, 2019, http://arxiv.org/abs/1904.04232v2.
APA
Chen, W.-Y., Liu, Y.-C., Kira, Z., Wang, Y.-C. F., & Huang, J.-B. (2019). A Closer Look at Few-shot Classification. arXiv. http://arxiv.org/abs/1904.04232v2
Chicago
Chen, W.-Y., Y.-C. Liu, Z. Kira, Y.-C. F. Wang, and J.-B. Huang. 2019. “A Closer Look at Few-shot Classification”. arXiv. http://arxiv.org/abs/1904.04232v2.
Harvard
Chen, W.-Y. et al. (2019) “A Closer Look at Few-shot Classification”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1904.04232v2.
Vancouver
1. Chen W-Y, Liu Y-C, Kira Z, Wang Y-CF, Huang J-B (2019) A Closer Look at Few-shot Classification. arXiv

BibTeX

@article{chen2019closer,
  title = {A Closer Look at Few-shot Classification},
  author = {Chen, Wei-Yu and Liu, Yen-Cheng and Kira, Zsolt and Wang, Yu-Chiang Frank and Huang, Jia-Bin},
  year = {2019},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1904.04232v2},
  eprint = {1904.04232}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF

License: Authors