Meta-Learning for Semi-Supervised Few-Shot Classification

Mengye RenEleni TriantafillouSachin RaviJake SnellKevin SwerskyJoshua B. TenenbaumHugo LarochelleRichard S. Zemel

article2018ICLR1,453 citations

Proposes semi-supervised extensions of Prototypical Networks that exploit unlabeled data and filter out distractor classes during episodic few-shot classification.

Listen

Standard deep learning systems require vast collections of labeled data, making them expensive to deploy and poorly suited for real-world tasks where only a handful of labeled examples are available. While few-shot meta-learning has emerged to train models to adapt rapidly to new categories, it traditionally ignores the abundance of readily accessible unlabeled data. Furthermore, in realistic scenarios, unlabeled data pools are noisy and contain "distractor" examples from irrelevant categories, which standard learning algorithms struggle to filter out.

The article develops and evaluates novel semi-supervised meta-learning models that systematically leverage unlabeled examples during training and testing. It demonstrates that algorithms can learn to refine category representations using unlabeled items and maintain robustness even when the unlabeled data contains irrelevant distractor categories.

The authors extended Prototypical Networks—a metric-learning framework that classifies inputs based on distance to class averages (prototypes)—into a semi-supervised episodic paradigm. They introduced three prototype refinement methods: basic soft clustering (Soft k-Means), soft clustering with an explicit distractor cluster, and Masked Soft k-Means, which dynamically masks out irrelevant unlabeled items using a lightweight neural network. The models were evaluated on benchmark datasets (Omniglot and miniImageNet) as well as tieredImageNet, a newly introduced benchmark constructed with a category hierarchy (608 classes across 34 categories) to prevent overlap between training and testing classes.

The experiments produced several critical findings. First, incorporating unlabeled data consistently outperforms purely supervised baselines across all datasets. For example, on tieredImageNet 1-shot classification, semi-supervised models improved accuracy from 46.52% to over 51–52%. Second, Masked Soft k-Means proved to be the most robust architecture when distractor categories were present, achieving top performance across datasets (such as 97.30% on Omniglot and 69.08% on tieredImageNet 5-shot) by effectively filtering out irrelevant noise. Third, models demonstrated strong extrapolation capabilities; training on only 5 unlabeled items per class generalized smoothly to 20 or 25 unlabeled items at test time, resulting in steady gains in accuracy.

These findings indicate that organizations can significantly reduce data labeling costs and operational timelines by combining limited labeled data with raw, uncurated data pools. The success of the masking mechanism shows that systems do not require pristine unlabeled datasets to benefit from semi-supervised learning, lowering the operational risk of automated web scraping or noisy data collection pipelines.

Practitioners facing low-data constraints should adopt masked clustering refinements within meta-learning architectures rather than relying on strictly supervised few-shot learners. When distractor noise is present, teams should favor threshold-masking mechanisms over single catch-all distractor clusters. For future development, the article recommends investigating adaptive embedding representations (such as fast weights) to enable representations to condition dynamically on episode context.

The confidence in these findings is high across the evaluated image recognition tasks, supported by standardized episodic splits and multiple random trials. However, confidence should be tempered when considering non-vision domains or environments with highly extreme distractor noise beyond the balanced 1:1 distractor-to-target ratio tested in this work.

  • Paper: TADAM: Task dependent adaptive metric for improved few-shot learning, Boris N. Oreshkin et al. (2018). This paper enhances prototype-based metric meta-learning by introducing task-dependent metric conditioning and metric scaling on benchmarks established in the source.
  • Paper: Meta-Learning With Differentiable Convex Optimization, Kwonjoon Lee et al. (2019). This work advances beyond simple prototype averaging in meta-learning by embedding convex optimization and support vector machines directly as the base learner.
  • Paper: Meta-Learning with Latent Embedding Optimization, Andrei A. Rusu et al. (2018). This paper extends few-shot meta-learning by adapting model parameters within a low-dimensional latent embedding space rather than relying purely on metric prototypes.
  • Paper: PANet: Few-Shot Image Semantic Segmentation With Prototype Alignment, Kaixin Wang et al. (2019). This research extends prototype-based metric learning paradigms from few-shot classification to semantic image segmentation using prototype alignment.
  • Paper: A Closer Look at Few-shot Classification, Wei-Yu Chen et al. (2019). This study provides a critical comparative benchmarking of metric-based few-shot algorithms including ProtoNet variants across varying backbone depths and domain shifts.
  • Paper: Generalizing from a Few Examples, Yaqing Wang et al. (2019). This survey provides a comprehensive taxonomy of few-shot learning methodologies, contextualizing prototype-based and semi-supervised meta-learning models.
  • Paper: A survey on semi-supervised learning, Jesper E. van Engelen et al. (2019). This survey synthesizes modern semi-supervised learning paradigms, systematically organizing inductive and transductive methods across machine learning.
Cover for Meta-Learning for Semi-Supervised Few-Shot Classification

Abstract

In few-shot classification, we are interested in learning algorithms that train a classifier from only a handful of labeled examples. Recent progress in few-shot classification has featured meta-learning, in which a parameterized model for a learning algorithm is defined and trained on episodes representing different classification problems, each with a small labeled training set and its corresponding test set. In this work, we advance this few-shot classification paradigm towards a scenario where unlabeled examples are also available within each episode. We consider two situations: one where all unlabeled examples are assumed to belong to the same set of classes as the labeled examples of the episode, as well as the more challenging situation where examples from other distractor classes are also provided. To address this paradigm, we propose novel extensions of Prototypical Networks (Snell et al., 2017) that are augmented with the ability to use unlabeled examples when producing prototypes. These models are trained in an end-to-end way on episodes, to learn to leverage the unlabeled examples successfully. We evaluate these methods on versions of the Omniglot and miniImageNet benchmarks, adapted to this new framework augmented with unlabeled examples. We also propose a new split of ImageNet, consisting of a large set of classes, with a hierarchical structure. Our experiments confirm that our Prototypical Networks can learn to improve their predictions due to unlabeled examples, much like a semi-supervised algorithm would.

Table of Contents

  • 1 Introduction
  • 2 Background
  • 2.1 Few-shot learning
  • 2.2 Prototypical Networks
  • 3 Semi-Supervised Few-Shot Learning
  • 3.1 Semi-Supervised Prototypical Networks
  • 3.1.1 Prototypical Networks with Soft kk-Means
  • 3.1.2 Prototypical Networks with Soft kk-Means with a Distractor Cluster
  • 3.1.3 Prototypical Networks with Soft kk-Means and Masking
  • 4 Related Work
  • 5 Experiments
  • 5.1 Datasets
  • 5.2 Adapting the Datasets for Semi-Supervised Learning
  • 5.3 Results
  • 6 Conclusion
  • References
  • A Omniglot Dataset Details
  • B tieredImagenet Dataset Details
  • C Extra Experimental Results
  • C.1 Few-shot classification baselines
  • C.2 Number of unlabeled items
  • D Hyperparameter Details

Knowls

  1. Knowl 1 — Masked Soft k-Means Prototypical Networks

    model/method

    Masked Soft kk-Means is a semi-supervised meta-learning method designed to refine class prototypes using unlabeled examples while remaining robust to unlabeled distractor instances from classes outside the episode's support set.

    Given an embedding network h(⋅)h(\cdot), a support set S={(xi,yi)}i=1N×KS = \{(x_i, y_i)\}_{i=1}^{N \times K} of NN classes with KK examples each, and an unlabeled set R={x~1,…,x~Mtot}R = \{\tilde{x}_1, \dots, \tilde{x}_{M_{\text{tot}}}\}, initial prototypes pcp_c for each class c∈{1,…,N}c \in \{1, \dots, N\} are computed as the support class means:

    pc=∑ih(xi)zi,c∑izi,c,where zi,c=1[yi=c]p_c = \frac{\sum_{i} h(x_i) z_{i,c}}{\sum_i z_{i,c}}, \quad \text{where } z_{i,c} = \mathbf{1}[y_i = c]

    For each unlabeled example x~j\tilde{x}_j and prototype pcp_c, Euclidean distances dj,c=∥h(x~j)−pc∥22d_{j,c} = \|h(\tilde{x}_j) - p_c\|_2^2 are normalized relative to their per-prototype mean across the MtotM_{\text{tot}} unlabeled examples:

    d~j,c=dj,c1Mtot∑j′=1Mtotdj′,c\tilde{d}_{j,c} = \frac{d_{j,c}}{\frac{1}{M_{\text{tot}}} \sum_{j'=1}^{M_{\text{tot}}} d_{j',c}}

    A multi-layer perceptron (MLP) takes five summary statistics of the normalized distances for cluster cc across all unlabeled examples—specifically the minimum, maximum, variance, skewness, and kurtosis—and predicts a soft distance threshold βc∈R\beta_c \in \mathbb{R} and a slope parameter γc∈R\gamma_c \in \mathbb{R}:

    [βc,γc]=MLP(min⁡j(d~j,c),max⁡j(d~j,c),varj(d~j,c),skewj(d~j,c),kurtj(d~j,c))[\beta_c, \gamma_c] = \text{MLP}\left(\min_j(\tilde{d}_{j,c}), \max_j(\tilde{d}_{j,c}), \text{var}_j(\tilde{d}_{j,c}), \text{skew}_j(\tilde{d}_{j,c}), \text{kurt}_j(\tilde{d}_{j,c})\right)

    Gradients are stopped through the distance summary statistics during backpropagation to prevent numerical instability. The soft mask mj,c∈(0,1)m_{j,c} \in (0, 1) indicating whether unlabeled instance x~j\tilde{x}_j belongs to class cc is computed using the logistic sigmoid function σ\sigma:

    mj,c=σ(−γc(d~j,c−βc))m_{j,c} = \sigma\left(-\gamma_c (\tilde{d}_{j,c} - \beta_c)\right)

    Combining the soft assignment weights z~j,c=exp⁡(−∥h(x~j)−pc∥22)∑c′=1Nexp⁡(−∥h(x~j)−pc′∥22)\tilde{z}_{j,c} = \frac{\exp(-\|h(\tilde{x}_j) - p_c\|_2^2)}{\sum_{c'=1}^N \exp(-\|h(\tilde{x}_j) - p_{c'}\|_2^2)} with the soft masks mj,cm_{j,c}, the refined prototype p~c\tilde{p}_c is computed as:

    p~c=∑ih(xi)zi,c+∑jh(x~j)z~j,cmj,c∑izi,c+∑jz~j,cmj,c\tilde{p}_c = \frac{\sum_{i} h(x_i) z_{i,c} + \sum_{j} h(\tilde{x}_j) \tilde{z}_{j,c} m_{j,c}}{\sum_{i} z_{i,c} + \sum_{j} \tilde{z}_{j,c} m_{j,c}}

    The network is trained end-to-end using the negative log-probability of the labeled query set examples under a softmax over Euclidean distances to the refined prototypes {p~c}c=1N\{\tilde{p}_c\}_{c=1}^N.

  2. Knowl 2 — Soft k-Means Prototypical Networks

    model/method

    Soft kk-Means Prototypical Networks extend standard Prototypical Networks to the semi-supervised setting by performing differentiable cluster refinement using unlabeled data within each meta-learning episode.

    Let S={(xi,yi)}i=1N×KS = \{(x_i, y_i)\}_{i=1}^{N \times K} denote the labeled support set across NN classes, R={x~1,…,x~Mtot}R = \{\tilde{x}_1, \dots, \tilde{x}_{M_{\text{tot}}}\} denote the set of unlabeled examples, and h(x)h(x) denote a parameterized neural network embedding function. Initial prototypes pcp_c are calculated as the mean embedding of the labeled support examples for class cc:

    pc=∑i=1N×Kh(xi)zi,c∑i=1N×Kzi,c,where zi,c=1[yi=c]p_c = \frac{\sum_{i=1}^{N \times K} h(x_i) z_{i,c}}{\sum_{i=1}^{N \times K} z_{i,c}}, \quad \text{where } z_{i,c} = \mathbf{1}[y_i = c]

    Unlabeled examples x~j∈R\tilde{x}_j \in R receive soft cluster assignments z~j,c∈(0,1)\tilde{z}_{j,c} \in (0, 1) computed via softmax over negative squared Euclidean distances to the initial prototypes:

    z~j,c=exp⁡(−∥h(x~j)−pc∥22)∑c′=1Nexp⁡(−∥h(x~j)−pc′∥22)\tilde{z}_{j,c} = \frac{\exp\left(-\|h(\tilde{x}_j) - p_c\|_2^2\right)}{\sum_{c'=1}^N \exp\left(-\|h(\tilde{x}_j) - p_{c'}\|_2^2\right)}

    The refined prototype p~c\tilde{p}_c for each class cc is then updated as the weighted average of both support and unlabeled embeddings:

    p~c=∑i=1N×Kh(xi)zi,c+∑j=1Mtoth(x~j)z~j,c∑i=1N×Kzi,c+∑j=1Mtotz~j,c\tilde{p}_c = \frac{\sum_{i=1}^{N \times K} h(x_i) z_{i,c} + \sum_{j=1}^{M_{\text{tot}}} h(\tilde{x}_j) \tilde{z}_{j,c}}{\sum_{i=1}^{N \times K} z_{i,c} + \sum_{j=1}^{M_{\text{tot}}} \tilde{z}_{j,c}}

    Query examples x∗x^* are assigned class probabilities using the refined prototypes:

    p(c∣x∗,{p~c})=exp⁡(−∥h(x∗)−p~c∥22)∑c′=1Nexp⁡(−∥h(x∗)−p~c′∥22)p(c \mid x^*, \{\tilde{p}_c\}) = \frac{\exp\left(-\|h(x^*) - \tilde{p}_c\|_2^2\right)}{\sum_{c'=1}^N \exp\left(-\|h(x^*) - \tilde{p}_{c'}\|_2^2\right)}

    The model parameters are optimized end-to-end across training episodes by minimizing the query set negative log-likelihood with a single prototype refinement step.

  3. Knowl 3 — Soft k-Means Prototypical Networks with a Distractor Cluster

    model/method

    Soft kk-Means with a Distractor Cluster extends semi-supervised Prototypical Networks to accommodate unlabeled distractor instances by augmenting the NN class prototypes with an extra (N+1)(N+1)-th background/distractor cluster centered at the origin.

    The prototype set {p1,…,pN+1}\{p_1, \dots, p_{N+1}\} is initialized as:

    pc={∑ih(xi)zi,c∑izi,cfor c∈{1,…,N}0for c=N+1p_c = \begin{cases} \frac{\sum_i h(x_i) z_{i,c}}{\sum_i z_{i,c}} & \text{for } c \in \{1, \dots, N\} \\ \mathbf{0} & \text{for } c = N + 1 \end{cases}

    where h(⋅)h(\cdot) is the embedding network and zi,c=1[yi=c]z_{i,c} = \mathbf{1}[y_i = c]. To account for the broad variance of unrelated distractor examples, each cluster has a length-scale parameter rcr_c. The length-scales of legitimate classes are fixed to r1=⋯=rN=1r_1 = \dots = r_N = 1, while the distractor scale rN+1∈R+r_{N+1} \in \mathbb{R}^+ is learned as a meta-parameter. Soft assignments z~j,c\tilde{z}_{j,c} for each unlabeled example x~j∈R\tilde{x}_j \in R are defined as:

    z~j,c=exp⁡(−1rc2∥h(x~j)−pc∥22−A(rc))∑c′=1N+1exp⁡(−1rc′2∥h(x~j)−pc′∥22−A(rc′)),where A(r)=12log⁡(2π)+log⁡(r)\tilde{z}_{j,c} = \frac{\exp\left(-\frac{1}{r_c^2}\|h(\tilde{x}_j) - p_c\|_2^2 - A(r_c)\right)}{\sum_{c'=1}^{N+1} \exp\left(-\frac{1}{r_{c'}^2}\|h(\tilde{x}_j) - p_{c'}\|_2^2 - A(r_{c'})\right)}, \quad \text{where } A(r) = \frac{1}{2}\log(2\pi) + \log(r)

    Prototype refinement is performed solely on the NN target classes using the soft assignments:

    p~c=∑ih(xi)zi,c+∑jh(x~j)z~j,c∑izi,c+∑jz~j,cfor c∈{1,…,N}\tilde{p}_c = \frac{\sum_i h(x_i) z_{i,c} + \sum_j h(\tilde{x}_j) \tilde{z}_{j,c}}{\sum_i z_{i,c} + \sum_j \tilde{z}_{j,c}} \quad \text{for } c \in \{1, \dots, N\}

    Unlabeled examples far from all target class prototypes are assigned high probability mass in z~j,N+1\tilde{z}_{j, N+1}, isolating them from corrupting the target class prototypes p~c\tilde{p}_c.

  4. Knowl 4 — Semi-Supervised Few-Shot Classification Episodic Formulation

    definition

    Semi-supervised few-shot classification formalizes episode-based meta-learning in the presence of unlabeled inputs during both training and evaluation.

    Given a set of training classes Ctrain\mathcal{C}_{\text{train}} and a disjoint set of test classes Ctest\mathcal{C}_{\text{test}} (Ctrain∩Ctest=∅ \mathcal{C}_{\text{train}} \cap \mathcal{C}_{\text{test}} = \emptyset), an NN-way, KK-shot semi-supervised episode is constructed with three sets:

    1. A labeled support set S={(x1,y1),…,(xN×K,yN×K)}S = \{(x_1, y_1), \dots, (x_{N \times K}, y_{N \times K})\}, containing KK labeled examples for each of NN sampled classes, with xi∈RDx_i \in \mathbb{R}^D and yi∈{1,…,N}y_i \in \{1, \dots, N\}.
    2. An unlabeled set R={x~1,…,x~Mtot}R = \{\tilde{x}_1, \dots, \tilde{x}_{M_{\text{tot}}}\}, consisting of MM unlabeled instances per target class (Mtot=N×MM_{\text{tot}} = N \times M) in the standard semi-supervised setting, or Mtot=N×M+H×MM_{\text{tot}} = N \times M + H \times M in the setting with distractors, where HH additional distractor classes are sampled from the same split and their unlabeled instances are pooled into RR without ground-truth indicators.
    3. A query set Q={(x1∗,y1∗),…,(xT∗,yT∗)}Q = \{(x^*_1, y^*_1), \dots, (x^*_T, y^*_T)\}, containing labeled instances belonging strictly to the NN target classes, used to calculate classification loss during training and accuracy during testing.

    Meta-learning optimizes the learner to exploit both SS and RR within each episode to maximize classification performance on QQ across unseen classes.

  5. Knowl 5 — tieredImageNet Hierarchical Dataset Split for Few-Shot Classification

    definition

    The tieredImageNet dataset is a large-scale few-shot learning benchmark derived from ILSVRC-12 consisting of 608 classes and 779,165 images (84×8484 \times 84 resolution).

    To establish strict semantic separation between meta-training and meta-testing phases, classes in tieredImageNet are grouped into 34 high-level categories based on the WordNet hierarchy. Each high-level category contains between 10 and 30 classes (average 17.8 classes). Classes with multiple parent categories in the hierarchy are removed to avoid split contamination.

    The 34 categories are split into disjoint sets at the category level:

    • Training split: 20 categories, 351 classes, 448,695 images.
    • Validation split: 6 categories, 97 classes, 124,261 images.
    • Testing split: 8 categories, 160 classes, 206,209 images.

    This structure ensures that training and test classes come from distinct higher-level semantic categories (unlike miniImageNet, where semantically related classes such as different musical instruments or animal breeds can be split across train and test sets).

  6. Knowl 6 — Semi-Supervised Classification Performance on tieredImageNet

    data/table

    5-way 1-shot and 5-shot classification accuracy (mean ±\pm standard error over 10 random labeled/unlabeled splits) on tieredImageNet across models trained on episodes with M=5M=5 unlabeled examples per class and tested with M=20M=20 unlabeled examples per class (with H=5H=5 distractor classes for the distractor condition "w/ D"):

    Models 1-shot Acc. 5-shot Acc. 1-shot Acc. w/ D 5-shot Acc. w/ D
    Supervised 46.52±0.5246.52 \pm 0.52 66.15±0.2266.15 \pm 0.22 46.52±0.5246.52 \pm 0.52 66.15±0.2266.15 \pm 0.22
    Semi-Supervised Inference 50.74±0.7550.74 \pm 0.75 69.37±0.2669.37 \pm 0.26 48.67±0.6048.67 \pm 0.60 67.46±0.2467.46 \pm 0.24
    Soft kk-Means 51.52±0.3651.52 \pm 0.36 70.25±0.3170.25 \pm 0.31 49.88±0.5249.88 \pm 0.52 68.32±0.2268.32 \pm 0.22
    Soft kk-Means+Cluster 51.85±0.2551.85 \pm 0.25 69.42±0.1769.42 \pm 0.17 51.36±0.3151.36 \pm 0.31 67.56±0.1067.56 \pm 0.10
    Masked Soft kk-Means 52.39±0.4452.39 \pm 0.44 69.88±0.2069.88 \pm 0.20 51.38±0.3851.38 \pm 0.38 69.08±0.2569.08 \pm 0.25

    Semi-supervised meta-training consistently outperforms both purely supervised Prototypical Networks and post-hoc Semi-Supervised Inference (which applies soft kk-means at test time to a supervised embedding). When distractors are present, Masked Soft kk-Means maintains superior accuracy compared to standard Soft kk-Means.

  7. Knowl 7 — Semi-Supervised Classification Performance on miniImageNet

    data/table

    5-way 1-shot and 5-shot classification accuracy (mean ±\pm standard error over 10 random labeled/unlabeled splits, using 40% of class data as labeled and 60% as unlabeled) on miniImageNet:

    Models 1-shot Acc. 5-shot Acc. 1-shot Acc. w/ D 5-shot Acc. w/ D
    Supervised 43.61±0.2743.61 \pm 0.27 59.08±0.2259.08 \pm 0.22 43.61±0.2743.61 \pm 0.27 59.08±0.2259.08 \pm 0.22
    Semi-Supervised Inference 48.98±0.3448.98 \pm 0.34 63.77±0.2063.77 \pm 0.20 47.42±0.3347.42 \pm 0.33 62.62±0.2462.62 \pm 0.24
    Soft kk-Means 50.09±0.4550.09 \pm 0.45 64.59±0.2864.59 \pm 0.28 48.70±0.3248.70 \pm 0.32 63.55±0.2863.55 \pm 0.28
    Soft kk-Means+Cluster 49.03±0.2449.03 \pm 0.24 63.08±0.1863.08 \pm 0.18 48.86±0.3248.86 \pm 0.32 61.27±0.2461.27 \pm 0.24
    Masked Soft kk-Means 50.41±0.3150.41 \pm 0.31 64.39±0.2464.39 \pm 0.24 49.04±0.3149.04 \pm 0.31 62.96±0.1462.96 \pm 0.14

    In the absence of distractors, Soft kk-Means and Masked Soft kk-Means achieve the top 1-shot and 5-shot accuracies. In the distractor condition ("w/ D"), Masked Soft kk-Means attains the highest 1-shot accuracy (49.04±0.31%49.04 \pm 0.31\%), demonstrating resistance to cluster corruption from distractor examples.

  8. Knowl 8 — Semi-Supervised Classification Performance on Omniglot

    data/table

    5-way 1-shot classification accuracy (mean ±\pm standard error over 10 splits) on Omniglot (10% labeled split, 90% unlabeled split) evaluated with M=20M=20 unlabeled examples per class, with and without H=5H=5 distractor classes:

    Models Acc. Acc. w/ D
    Supervised 94.62±0.0994.62 \pm 0.09 94.62±0.0994.62 \pm 0.09
    Semi-Supervised Inference 97.45±0.0597.45 \pm 0.05 95.08±0.0995.08 \pm 0.09
    Soft kk-Means 97.25±0.1097.25 \pm 0.10 95.01±0.0995.01 \pm 0.09
    Soft kk-Means+Cluster 97.68±0.0797.68 \pm 0.07 97.17±0.0497.17 \pm 0.04
    Masked Soft kk-Means 97.52±0.0797.52 \pm 0.07 97.30±0.0897.30 \pm 0.08

    While ordinary Soft kk-Means drops in accuracy from 97.25%97.25\% to 95.01%95.01\% when distractors are added, Soft kk-Means+Cluster (97.17%97.17\%) and Masked Soft kk-Means (97.30%97.30\%) retain high accuracy, confirming their ability to filter out out-of-distribution unlabeled characters.

  9. Knowl 9 — Test-Time Unlabeled Sample Size Extrapolation in Semi-Supervised Meta-Learning

    empirical result

    Semi-supervised Prototypical Networks meta-trained with a fixed small number of unlabeled examples per class (M=5M=5) successfully extrapolate to larger unlabeled sample sizes at test time without retraining.

    When evaluating 1-shot and 5-shot classification on tieredImageNet with varying test unlabeled set sizes Mtest∈{0,1,2,5,10,15,20,25}M_{\text{test}} \in \{0, 1, 2, 5, 10, 15, 20, 25\}:

    • Accuracy increases monotonically as MtestM_{\text{test}} grows from 0 to 25 across all semi-supervised models, both with and without distractor classes.
    • For 1-shot Masked Soft kk-Means without distractors, accuracy improves from 46.41%46.41\% at M=0M=0 to 50.19%50.19\% at M=5M=5, and reaches 52.61%52.61\% at M=25M=25.
    • For 1-shot Masked Soft kk-Means with distractors, accuracy improves from 46.20%46.20\% at M=0M=0 to 49.87%49.87\% at M=5M=5, and reaches 51.49%51.49\% at M=25M=25.

    This demonstrates that meta-learning induces an embedding space where prototype refinement scales with the availability of unlabeled test data.

Coverage note — None was omitted; all primary contributions (the semi-supervised episodic learning framework, three prototype refinement methods, the tieredImageNet dataset specification, and the full experimental results across Omniglot, miniImageNet, and tieredImageNet) are covered.

References

  1. 1.Jimmy Ba, Geoffrey E. Hinton, Volodymyr Mnih, Joel Z. Leibo, and Catalin Ionescu. Using fast weights to attend to the recent past. In Advances in Neural Information Processing Systems 29: Annual Conference on Neural Information Processing Systems 2016, December 5-10, 2016, Barcelona, Spain, pp. 4331–4339, 2016.
  2. 2.Philip Bachman, Alessandro Sordoni, and Adam Trischler. Learning algorithms for active learning. 2017.
  3. 3.Olivier Chapelle, Bernhard Schölkopf, and Alexander Zien. Semi-Supervised Learning. The MIT Press, 1st edition, 2010. ISBN 0262514125, 9780262514125.
  4. 4.Sanjay Chawla and Aristides Gionis. k-means–: A unified approach to clustering and outlier detection. In Proceedings of the 2013 SIAM International Conference on Data Mining, pp. 189–197. SIAM, 2013.
  5. 5.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In Computer Vision and Pattern Recognition, 2009. CVPR 2009. IEEE Conference on, pp. 248–255. IEEE, 2009.
  6. 6.Chelsea Finn, Pieter Abbeel, and Sergey Levine. Model-agnostic meta-learning for fast adaptation of deep networks. In 34th International Conference on Machine Learning, 2017.
  7. 7.Yanwei Fu, Timothy M. Hospedales, Tao Xiang, and Shaogang Gong. Transductive multi-view zero-shot learning. IEEE Trans. Pattern Anal. Mach. Intell., 37(11):2332–2345, 2015.
  8. 8.Shalmoli Gupta, Ravi Kumar, Kefu Lu, Benjamin Moseley, and Sergei Vassilvitskii. Local search methods for k-means with outliers. Proceedings of the VLDB Endowment, 10(7):757–768, 2017.
  9. 9.Ville Hautamäki, Svetlana Cherednichenko, Ismo Kärkkäinen, Tomi Kinnunen, and Pasi Fränti. Improving k-means by outlier removal. In Scandinavian Conference on Image Analysis, pp. 978–987. Springer, 2005.
  10. 10.Sepp Hochreiter, A Steven Younger, and Peter R Conwell. Learning to learn using gradient descent. In International Conference on Artificial Neural Networks, pp. 87–94. Springer, 2001.
  11. 11.Thorsten Joachims. Transductive inference for text classification using support vector machines. In Proceedings of the Sixteenth International Conference on Machine Learning, 1999.
  12. 12.Diederik Kingma and Jimmy Ba. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014.
  13. 13.Gregory Koch, Richard Zemel, and Ruslan Salakhutdinov. Siamese neural networks for one-shot image recognition. In ICML Deep Learning Workshop, volume 2, 2015.
  14. 14.Brenden M. Lake, Ruslan Salakhutdinov, Jason Gross, and Joshua B. Tenenbaum. One shot learning of simple visual concepts. In Proceedings of the 33th Annual Meeting of the Cognitive Science Society, CogSci 2011, Boston, Massachusetts, USA, July 20-23, 2011, 2011.
  15. 15.Stuart Lloyd. Least squares quantization in pcm. IEEE transactions on information theory, 28(2): 129–137, 1982.
  16. 16.Nikhil Mishra, Mostafa Rohaninejad, Xi Chen, and Pieter Abbeel. Meta-learning with temporal convolutions. CoRR, abs/1707.03141, 2017.
  17. 17.Sachin Ravi and Hugo Larochelle. Optimization as a model for few-shot learning. In 5th International Conference on Learning Representations, 2017.
  18. 18.Chuck Rosenberg, Martial Hebert, and Henry Schneiderman. Semi-supervised self-training of object detection models. 2005.
  19. 19.Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, et al. Imagenet large scale visual recognition challenge. International Journal of Computer Vision, 115(3):211–252, 2015.
  20. 20.Adam Santoro, Sergey Bartunov, Matthew Botvinick, Daan Wierstra, and Timothy P. Lillicrap. One-shot learning with memory-augmented neural networks. In 33rd International Conference on Machine Learning, 2016.
  21. 21.Jake Snell, Kevin Swersky, and Richard S. Zemel. Prototypical networks for few-shot learning. In Advances in Neural Information Processing Systems 30, 2017.
  22. 22.Sebastian Thrun. Lifelong learning algorithms. In Learning to learn, pp. 181–209. Springer, 1998.
  23. 23.V.N. Vapnik. Statistical Learning Theory. Wiley, 1998. ISBN 9788126528929.
  24. 24.Oriol Vinyals, Charles Blundell, Tim Lillicrap, Koray Kavukcuoglu, and Daan Wierstra. Matching networks for one shot learning. In Advances in Neural Information Processing Systems 29, pp. 3630–3638, 2016.
  25. 25.David Yarowsky. Unsupervised word sense disambiguation rivaling supervised methods. In Proceedings of the 33rd annual meeting on Association for Computational Linguistics, pp. 189–196. Association for Computational Linguistics, 1995.
  26. 26.Xiaojin Zhu. Semi-supervised learning literature survey. 2005.

Citation

MLA
Ren, M., et al. “Meta-Learning for Semi-Supervised Few-Shot Classification”. arXiv, 2018, http://arxiv.org/abs/1803.00676v1.
APA
Ren, M., Triantafillou, E., Ravi, S., Snell, J., Swersky, K., Tenenbaum, J. B., Larochelle, H., & Zemel, R. S. (2018). Meta-Learning for Semi-Supervised Few-Shot Classification. arXiv. http://arxiv.org/abs/1803.00676v1
Chicago
Ren, M., E. Triantafillou, S. Ravi, et al. 2018. “Meta-Learning for Semi-Supervised Few-Shot Classification”. arXiv. http://arxiv.org/abs/1803.00676v1.
Harvard
Ren, M. et al. (2018) “Meta-Learning for Semi-Supervised Few-Shot Classification”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1803.00676v1.
Vancouver
1. Ren M, Triantafillou E, Ravi S, Snell J, Swersky K, Tenenbaum JB, Larochelle H, Zemel RS (2018) Meta-Learning for Semi-Supervised Few-Shot Classification. arXiv

BibTeX

@article{ren2018meta,
  title = {Meta-Learning for Semi-Supervised Few-Shot Classification},
  author = {Ren, Mengye and Triantafillou, Eleni and Ravi, Sachin and Snell, Jake and Swersky, Kevin and Tenenbaum, Joshua B. and Larochelle, Hugo and Zemel, Richard S.},
  year = {2018},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1803.00676v1},
  eprint = {1803.00676}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Published with permission