Generalizing from a Few Examples

Yaqing WangQuanming YaoJames KwokLionel M. Ni

article2019ACM Computing Surveys2,192 citations

Establishes a unified taxonomy for few-shot learning by categorizing methods according to how prior knowledge is applied across data, model, and algorithm perspectives to overcome the core challenge of unreliable empirical risk minimization.

arXiv: 1904.05046
Cover for Generalizing from a Few Examples

Abstract

Machine learning has been highly successful in data-intensive applications but is often hampered when the data set is small. Recently, Few-Shot Learning (FSL) is proposed to tackle this problem. Using prior knowledge, FSL can rapidly generalize to new tasks containing only a few samples with supervised information. In this paper, we conduct a thorough survey to fully understand FSL. Starting from a formal definition of FSL, we distinguish FSL from several relevant machine learning problems. We then point out that the core issue in FSL is that the empirical risk minimized is unreliable. Based on how prior knowledge can be used to handle this core issue, we categorize FSL methods from three perspectives: (i) data, which uses prior knowledge to augment the supervised experience; (ii) model, which uses prior knowledge to reduce the size of the hypothesis space; and (iii) algorithm, which uses prior knowledge to alter the search for the best hypothesis in the given hypothesis space. With this taxonomy, we review and discuss the pros and cons of each category. Promising directions, in the aspects of the FSL problem setups, techniques, applications and theories, are also proposed to provide insights for future research.

Table of Contents

  • 1 Introduction
  • 1.1 Organization of the Survey
  • 1.2 Notation and Terminology
  • 2 Overview
  • 2.1 Problem Definition
  • 2.2 Relevant Learning Problems
  • 2.3 Core Issue
  • 2.3.1 Empirical Risk Minimization
  • 2.3.2 Unreliable Empirical Risk Minimizer
  • 2.4 Taxonomy
  • 3 Data
  • 3.1 Transforming Samples from Dtrain{D}_{\text{train}}
  • 3.2 Transforming Samples from a Weakly Labeled or Unlabeled Data Set
  • 3.3 Transforming Samples from Similar Data Sets
  • 3.4 Discussion and Summary
  • 4 Model
  • 4.1 Multitask Learning
  • 4.1.1 Parameter Sharing
  • 4.1.2 Parameter Tying
  • 4.2 Embedding Learning
  • 4.2.1 Task-Specific Embedding Model
  • 4.2.2 Task-Invariant Embedding Model
  • 4.2.3 Hybrid Embedding Model
  • 4.3 Learning with External Memory
  • 4.3.1 Refining Representations
  • 4.3.2 Refining Parameters
  • 4.4 Generative Modeling
  • 4.4.1 Decomposable Components
  • 4.4.2 Groupwise Shared Prior
  • 4.4.3 Parameters of Inference Networks
  • 4.5 Discussion and Summary
  • 5 Algorithm
  • 5.1 Refining Existing Parameters
  • 5.1.1 Fine-Tuning Existing Parameter by Regularization
  • 5.1.2 Aggregating a Set of Parameters
  • 5.1.3 Fine-Tuning Existing Parameter with New Parameters
  • 5.2 Refining Meta-Learned Parameter
  • 5.3 Learning the Optimizer
  • 5.4 Discussion and Summary
  • 6 Future Works
  • 6.1 Problem Setups
  • 6.2 Techniques
  • 6.3 Applications
  • 6.3.1 Computer Vision
  • 6.3.2 Robotics
  • 6.3.3 Natural Language Processing
  • 6.3.4 Acoustic Signal Processing
  • 6.3.5 Others
  • 6.4 Theories
  • 7 Conclusion
  • A Appendix: Meta-Learning
  • References

Knowls

  1. Knowl 1 — Formal Definition of Few-Shot Learning

    definition

    In machine learning, a computer program is said to learn from experience EE with respect to a class of tasks TT and performance measure PP if its performance on TT, measured by PP, improves with EE.

    Few-Shot Learning (FSL) is a type of machine learning problem (specified by EE, TT, and PP) where the experience EE contains only a limited number of supervised examples directly relevant to the target task TT. In a standard supervised FSL task TT, the dataset D={Dtrain,Dtest}D = \{D_{\text{train}}, D_{\text{test}}\} consists of a small training set Dtrain={(xi,yi)}i=1ID_{\text{train}} = \{(x_i, y_i)\}_{i=1}^I, where II is small, and a test set Dtest={xtest}D_{\text{test}} = \{x^{\text{test}}\}.

    • In NN-way-KK-shot classification, DtrainD_{\text{train}} contains I=K×NI = K \times N labeled examples drawn from NN distinct classes, with exactly KK examples per class.
    • When K=1K = 1, the task is termed one-shot learning.
    • When K=0K = 0 (meaning EE contains no supervised training examples of target classes), the problem becomes zero-shot learning (ZSL), which requires EE to include auxiliary information from other modalities (such as semantic attributes, WordNet hierarchies, knowledge graphs, or word embeddings) to enable cross-modal knowledge transfer.
  2. Knowl 2 — Error Decomposition and the Core Issue of Few-Shot Learning

    theoretical result

    Let p(x,y)p(x, y) denote the ground-truth joint probability distribution over input xx and output yy, (y^,y)\ell(\hat{y}, y) denote a loss function measuring prediction error, and H\mathcal{H} denote a hypothesis space parameterized by θ\theta. The expected risk R(h)R(h) and empirical risk RI(h)R_I(h) evaluated on a training set Dtrain={(xi,yi)}i=1ID_{\text{train}} = \{(x_i, y_i)\}_{i=1}^I of size II are:

    R(h)=E(x,y)p[(h(x),y)],RI(h)=1Ii=1I(h(xi),yi).R(h) = \mathbb{E}_{(x, y) \sim p}[\ell(h(x), y)], \qquad R_I(h) = \frac{1}{I} \sum_{i=1}^I \ell(h(x_i), y_i).

    Let h^=argminhR(h)\hat{h} = \arg\min_h R(h) be the global optimal hypothesis, h=argminhHR(h)h^* = \arg\min_{h \in \mathcal{H}} R(h) be the optimal hypothesis constrained within H\mathcal{H}, and hI=argminhHRI(h)h_I = \arg\min_{h \in \mathcal{H}} R_I(h) be the empirical risk minimizer on DtrainD_{\text{train}}. The total expected error decomposes into approximation error Eapp(H)\mathcal{E}_{\text{app}}(\mathcal{H}) and estimation error Eest(H,I)\mathcal{E}_{\text{est}}(\mathcal{H}, I):

    E[R(hI)R(h^)]=E[R(h)R(h^)]Eapp(H)+E[R(hI)R(h)]Eest(H,I),\mathbb{E}[R(h_I) - R(\hat{h})] = \underbrace{\mathbb{E}[R(h^*) - R(\hat{h})]}_{\mathcal{E}_{\text{app}}(\mathcal{H})} + \underbrace{\mathbb{E}[R(h_I) - R(h^*)]}_{\mathcal{E}_{\text{est}}(\mathcal{H}, I)},

    where the expectation is taken over the random sampling of DtrainD_{\text{train}}.

    The core issue of FSL is that because the number of available supervised training samples II is extremely small, the empirical risk RI(h)R_I(h) is a poor proxy for the expected risk R(h)R(h). Consequently, the empirical risk minimizer hIh_I becomes unreliable and suffers from severe overfitting. To reduce total generalization error, FSL methods must leverage prior knowledge to augment data, constrain the complexity of H\mathcal{H}, or alter the algorithmic search for hh^*.

  3. Knowl 3 — Three-Perspective Taxonomy for Few-Shot Learning

    model/method

    To address the unreliability of the empirical risk minimizer hIh_I when the sample size II is small, FSL methods incorporate prior knowledge across three distinct perspectives:

    1. Data: Uses prior knowledge to augment the small training set DtrainD_{\text{train}}, expanding its size from II to I~\tilde{I} where I~I\tilde{I} \gg I. Standard supervised learning models and optimization algorithms can then be directly applied to the augmented dataset to obtain a reliable empirical risk minimizer hI~h_{\tilde{I}}.
    2. Model: Uses prior knowledge to constrain the capacity and structure of the hypothesis space H\mathcal{H} into a significantly smaller hypothesis space H~H\tilde{\mathcal{H}} \subset \mathcal{H}, removing hypotheses that are unlikely to contain the optimal hh^*. Within H~\tilde{\mathcal{H}}, the small dataset DtrainD_{\text{train}} contains sufficient supervision to identify a reliable empirical minimizer without overfitting.
    3. Algorithm: Uses prior knowledge to guide or alter the parameter search strategy in H\mathcal{H}. Prior knowledge provides an informed parameter initialization θ0\theta_0 or directly parameterizes a meta-learned optimizer that outputs update steps to find the optimal parameter vector θ\theta^* for hh^*.
  4. Knowl 4 — Data Augmentation Strategies for Few-Shot Learning

    model/method

    Data-perspective FSL methods use a transformer function t()t(\cdot) informed by prior knowledge to synthesize additional samples (x~,y~)(\tilde{x}, \tilde{y}), expanding the training set DtrainD_{\text{train}} via three primary data sources:

    1. Transforming samples from DtrainD_{\text{train}}: Learns transformation functions t(xi)t(x_i) from auxiliary classes (such as intra-class variation autoencoders, sample-pair difference transfer functions, or continuous attribute subspace regressors) and applies them to each (xi,yi)Dtrain(x_i, y_i) \in D_{\text{train}} to generate (t(xi),yi)(t(x_i), y_i).
    2. Transforming samples from weakly labeled or unlabeled datasets: Employs a model pre-trained or refined on DtrainD_{\text{train}} (e.g., exemplar SVMs, progressive pseudo-labeling, or label propagation on graph networks) to predict target labels t(xˉ)t(\bar{x}) for unlabeled inputs xˉ\bar{x}, augmenting DtrainD_{\text{train}} with pseudo-labeled pairs (xˉ,t(xˉ))(\bar{x}, t(\bar{x})).
    3. Transforming samples from similar datasets: Aggregates and adapts labeled samples {(x^j,y^j)}\{(\hat{x}_j, \hat{y}_j)\} from related large-scale datasets using similarity metrics or dual-generator generative adversarial networks (GANs) that map between few-shot and abundant class domains, synthesizing (t({x^j}),t({y^j}))(t(\{\hat{x}_j\}), t(\{\hat{y}_j\})).

    A limitation of data augmentation strategies is that transformation policies are frequently hand-crafted or tailored in an ad-hoc manner for specific image datasets, making transfer to complex structural data (such as text or audio) non-trivial.

  5. Knowl 5 — Hypothesis Space Constraining via Multitask Learning

    model/method

    Multitask learning constrains the few-shot hypothesis space by jointly learning CC related tasks T1,,TCT_1, \dots, T_C with corresponding datasets Dc={Dtrainc,Dtestc}D_c = \{D^c_{\text{train}}, D^c_{\text{test}}\}, where few-shot tasks serve as target tasks and data-abundant tasks serve as source tasks. Joint learning forces the parameters θc\theta_c of the task hypothesis hch_c to be constrained by information from other tasks via two main mechanisms:

    • Parameter Sharing: Directly shares generic network layers (such as early convolutional feature extractors or variational autoencoder encoders) across all tasks, while training task-specific output layers or classification heads.
    • Parameter Tying: Maintains separate parameter sets θc\theta_c for each task while adding explicit regularization terms to the joint objective to penalize parameter divergences (e.g., pairwise distance penalties between parameter vectors θcθc\|\theta_c - \theta_{c'}\| or layer-wise alignment terms).

    A limitation of multitask learning in FSL is that all tasks must be trained jointly; introducing a new few-shot task requires retraining the full multitask model, and source tasks with large sample counts may dominate the optimization.

  6. Knowl 6 — Hypothesis Space Constraining via Embedding Learning

    model/method

    Embedding learning projects input samples xiXRdx_i \in \mathcal{X} \subseteq \mathbb{R}^d into a lower-dimensional embedding space ziZRmz_i \in \mathcal{Z} \subseteq \mathbb{R}^m where intra-class samples cluster together and inter-class samples are well-separated, enabling the construction of a lower-complexity hypothesis space H~\tilde{\mathcal{H}}.

    The framework consists of an embedding function f(xtest)f(x^{\text{test}}) for test queries, an embedding function g(xi)g(x_i) for support samples xiDtrainx_i \in D_{\text{train}}, and a similarity metric s(f(xtest),g(xi))s(f(x^{\text{test}}), g(x_i)). Methods fall into three categories:

    1. Task-Specific Embedding: Learns embeddings using only task-specific sample pairs and ranking losses generated from DtrainD_{\text{train}}.
    2. Task-Invariant Embedding: Learns a fixed generic metric space across auxiliary tasks via episodic meta-learning without updating the embedding parameters on DtrainD_{\text{train}}:
      • Matching Networks: Compares f(xtest)f(x^{\text{test}}) to each g(xi)g(x_i) using cosine similarity over contextual bidirectional LSTM representations.
      • Prototypical Networks (ProtoNet): Computes class prototypes cn=1Ki=1Kg(xi)c_n = \frac{1}{K} \sum_{i=1}^K g(x_i) for each class nn and classifies xtestx^{\text{test}} using squared Euclidean distance: s(f(xtest),cn)=f(xtest)cn22s(f(x^{\text{test}}), c_n) = -\|f(x^{\text{test}}) - c_n\|_2^2.
      • Relation Networks / Graph Neural Networks: Parameterizes similarity metrics directly via deep neural networks or graph edge-update layers.
    3. Hybrid Embedding: Dynamically adapts task-invariant embedding functions using task-specific embeddings extracted from DtrainD_{\text{train}} (such as Learnet, TADAM, and dynamic conditional networks).
  7. Knowl 7 — Hypothesis Space Constraining via External Memory Networks

    model/method

    External memory networks constrain the few-shot hypothesis space by storing knowledge extracted from DtrainD_{\text{train}} into a key-value memory matrix MRb×mM \in \mathbb{R}^{b \times m} consisting of bb memory slots M(i)=(Mkey(i),Mvalue(i))M(i) = (M_{\text{key}}(i), M_{\text{value}}(i)). Test samples xtestx^{\text{test}} are represented as attention-weighted averages over retrieved memory values according to the similarity s(f(xtest),Mkey(i))s(f(x^{\text{test}}), M_{\text{key}}(i)).

    Methods operate under two primary designs:

    • Refining Representations: Stores support representations in memory (e.g., Memory-Augmented Neural Networks (MANN), surprise-based memory updating, and lifelong memory models). To protect rare/few-shot classes from being overwritten by abundant classes in lifelong settings, memory write operations reset the age of accessed slots to zero and selectively replace the oldest slot only when the predicted output is incorrect.
    • Refining Parameters: Uses memory slots to store meta-learned "slow weights" and rapid task-specific "fast weights" (e.g., Meta Networks, conditionally shifted neurons) to modulate or generate network parameters conditioned on DtrainD_{\text{train}}.

    A limitation is the computational and memory footprint of matrix key-value lookups, which limits external memory sizes in practical deployments.

  8. Knowl 8 — Hypothesis Space Constraining via Generative Modeling

    model/method

    Generative modeling methods constrain the hypothesis space H\mathcal{H} by modeling the data distribution p(x;θ)=p(xz;θ)p(z;γ)dzp(x; \theta) = \int p(x|z; \theta) p(z; \gamma) dz, where the latent prior distribution p(z;γ)p(z; \gamma) is pre-trained on auxiliary data to enforce structured inductive biases. The resulting posterior over latent parameters is constrained to a compact subspace H~\tilde{\mathcal{H}}.

    Generative FSL models are divided into three types:

    1. Decomposable Components: Models objects hierarchically as combinations of shared primitive sub-parts (e.g., Bayesian Program Learning decomposes handwritten characters into types, tokens, templates, parts, and strokes), restricting inference to valid component combinations.
    2. Groupwise Shared Priors: Discovers hierarchical task clusters via unsupervised learning on auxiliary datasets {Dc}\{D_c\} to learn group-level class prior probabilities, which are inherited by novel few-shot classes assigned to that cluster.
    3. Inference Networks: Approximates intractable true posteriors p(zx;θ,γ)=p(xz;θ)p(z;γ)p(xz;θ)p(z;γ)dzp(z|x; \theta, \gamma) = \frac{p(x|z; \theta)p(z; \gamma)}{\int p(x|z; \theta)p(z; \gamma)dz} via amortized variational inference using deep inference networks q(z;δ)q(z; \delta) (e.g., variational autoencoders, conditional autoregressive models, GANs).

    A limitation is that generative inference involves higher computational cost and more complex derivation compared to deterministic discriminative models.

  9. Knowl 9 — Search Strategy Alteration via Pre-trained Parameter Refining

    model/method

    This strategy uses an initial parameter vector θ0\theta_0 pre-trained on related tasks to provide a favorable starting point in the hypothesis space H\mathcal{H}, reducing the required optimization steps on the small target dataset DtrainD_{\text{train}}. To prevent severe overfitting during standard gradient descent on DtrainD_{\text{train}}, several regularization techniques are employed:

    • Fine-Tuning by Regularization: Prevents overfitting via early stopping on a split validation subset, selective updating (freezing most pre-trained layers and updating only specific filter strengths or top layers), grouping related weights (updating neuron clusters jointly), or learning a model regression network that maps few-shot parameter estimates to large-sample parameter values.
    • Aggregating Pre-trained Parameters: Reuses and combines multiple pre-trained model parameter vectors {θ0k}\{\theta_0^k\} learned from diverse auxiliary tasks by learning linear combination coefficients on DtrainD_{\text{train}}.
    • Parameter Expansion: Preserves pre-trained feature parameters θ0\theta_0 and introduces dedicated task-specific parameters δ\delta (e.g., learning a linear classifier head on fixed deep features, θ={θ0,δ}\theta = \{\theta_0, \delta\}).

    A trade-off of pre-trained parameter refining is that θ0\theta_0 is static and learned from disparate tasks, which can sacrifice task-specific precision for optimization speed.

  10. Knowl 10 — Search Strategy Alteration via Meta-Learned Parameter Initialization

    model/method

    Meta-learned parameter initialization optimizes an initial parameter vector θ0\theta_0 across a distribution of training tasks p(T)p(T) such that one or a few gradient descent steps on few-shot data DtrainsD^s_{\text{train}} yield high generalization performance on DtestsD^s_{\text{test}}.

    In the Model-Agnostic Meta-Learning (MAML) framework, given task Tsp(T)T_s \sim p(T), task adaptation is performed in an inner loop via gradient descent:

    ϕs=θ0αθ0Ltrains(θ0),\phi_s = \theta_0 - \alpha \nabla_{\theta_0} \mathcal{L}^s_{\text{train}}(\theta_0),

    where α\alpha is the inner step size and Ltrains\mathcal{L}^s_{\text{train}} is the empirical loss on DtrainsD^s_{\text{train}}. Across meta-training tasks, the meta-parameter θ0\theta_0 is optimized in an outer loop via feedback on test sets DtestsD^s_{\text{test}}:

    θ0θ0βθ0Tsp(T)Ltests(ϕs),\theta_0 \leftarrow \theta_0 - \beta \nabla_{\theta_0} \sum_{T_s \sim p(T)} \mathcal{L}^s_{\text{test}}(\phi_s),

    where β\beta is the meta-step size.

    Key extensions to MAML include:

    1. Conditioning initialization on task-specific subspaces or embeddings to avoid uniform initializations across heterogeneous tasks.
    2. Probabilistic and hierarchical Bayesian formulations to model epistemic parameter uncertainty resulting from tiny sample sizes.
    3. Regularizing inner gradient descent trajectories using model regression networks.
  11. Knowl 11 — Search Strategy Alteration via Learned Optimizers

    model/method

    Instead of using hand-designed gradient descent update rules to search the parameter space H\mathcal{H}, learned optimizer methods meta-learn an optimization model (such as a Recurrent Neural Network or LSTM) that directly computes and outputs parameter update steps Δϕt\Delta\phi_t.

    At iteration tt of learning a new task, the task-specific parameter ϕt\phi_t is updated as:

    ϕt=ϕt1+Δϕt1,\phi_t = \phi_{t-1} + \Delta\phi_{t-1},

    where the meta-learner receives the loss signal t(ϕt1)=(h(xt;ϕt1),yt)\ell_t(\phi_{t-1}) = \ell(h(x_t; \phi_{t-1}), y_t) and gradient information from the learner on support sample (xt,yt)Dtrain(x_t, y_t) \in D_{\text{train}}, and produces Δϕt1\Delta\phi_{t-1} directly. In LSTM-based optimizers, the learner's parameter ϕt\phi_t is mapped directly to the cell state of the LSTM meta-learner.

    The meta-optimizer's own parameters are optimized across tasks Tsp(T)T_s \sim p(T) using gradient descent on test loss Ltests(ϕT)\mathcal{L}^s_{\text{test}}(\phi_T) evaluated on DtestsD^s_{\text{test}}, eliminating manual tuning of step sizes and search directions.

  12. Knowl 12 — Episodic Meta-Learning Framework for Few-Shot Learning

    definition

    Meta-learning for few-shot learning operates over a task distribution p(T)p(T) using an episodic training and evaluation paradigm partitioned into meta-training and meta-testing phases:

    • Meta-Training: Samples a set of support tasks Tsp(T)T_s \sim p(T). Each task TsT_s provides an episode dataset Ds={Dtrains,Dtests}D_s = \{D^s_{\text{train}}, D^s_{\text{test}}\} with NN classes, where DtrainsD^s_{\text{train}} contains KK examples per class (NN-way KK-shot). The meta-learner optimizes generic meta-parameters θ0\theta_0 across tasks to minimize the expected test loss across learners:

    θ0=argminθ0ETsp(T)(xi,yi)Dtests(h(xi;θ0(Dtrains)),yi).\theta_0 = \arg\min_{\theta_0} \mathbb{E}_{T_s \sim p(T)} \sum_{(x_i, y_i) \in D^s_{\text{test}}} \ell(h(x_i; \theta_0(D^s_{\text{train}})), y_i).

    • Meta-Testing: Evaluates generalization using a disjoint set of test tasks Ttp(T)T_t \sim p(T) with datasets Dt={Dtraint,Dtestt}D_t = \{D^t_{\text{train}}, D^t_{\text{test}}\} drawn from target classes unseen during meta-training. The learner adapts using DtraintD^t_{\text{train}} and is evaluated on DtesttD^t_{\text{test}}, with the average loss across all TtT_t defining the meta-learning test performance.

Coverage note — Domain-specific benchmark applications (e.g., character recognition on Omniglot, image classification on miniImageNet, robotics, and NLP benchmarks) and citations of individual application papers were omitted as they exemplify application instances rather than the survey's foundational contributions (the formal problem definitions, error analysis, three-perspective taxonomy, and category-level methodological mechanisms).

References

  1. 1.N. Abdo, H. Kretzschmar, L. Spinello, and C. Stachniss. 2013. Learning manipulation actions from a few demonstrations. In International Conference on Robotics and Automation. 1268–1275.
  2. 2.Z. Akata, F. Perronnin, Z. Harchaoui, and C. Schmid. 2013. Label-embedding for attribute-based classification. In Conference on Computer Vision and Pattern Recognition. 819–826.
  3. 3.M. Al-Shedivat, T. Bansal, Y. Burda, I. Sutskever, I. Mordatch, and P. Abbeel. 2018. Continuous adaptation via meta-learning in nonstationary and competitive environments. In International Conference on Learning Representations.
  4. 4.H. Altae-Tran, B. Ramsundar, A. S. Pappu, and V. Pande. 2017. Low data drug discovery with one-shot learning. ACS Central Science 3, 4 (2017), 283–293.
  5. 5.M. Andrychowicz, M. Denil, S. Gomez, M. W. Hoffman, D. Pfau, T. Schaul, and N. de Freitas. 2016. Learning to learn by gradient descent by gradient descent. In Advances in Neural Information Processing Systems. 3981–3989.
  6. 6.S. Arik, J. Chen, K. Peng, W. Ping, and Y. Zhou. 2018. Neural voice cloning with a few samples. In Advances in Neural Information Processing Systems. 10019–10029.
  7. 7.S. Azadi, M. Fisher, V. G. Kim, Z. Wang, E. Shechtman, and T. Darrell. 2018. Multi-content GAN for few-shot font style transfer. In Conference on Computer Vision and Pattern Recognition. 7564–7573.
  8. 8.P. Bachman, A. Sordoni, and A. Trischler. 2017. Learning algorithms for active learning. In International Conference on Machine Learning. 301–310.
  9. 9.Bengio Y. Bahdanau D, Cho K. 2015. Neural machine translation by jointly learning to align and translate. In International Conference on Learning Representations.
  10. 10.E. Bart and S. Ullman. 2005. Cross-generalization: Learning novel classes from a single example by feature replacement. In Conference on Computer Vision and Pattern Recognition, Vol. 1. 672–679.
  11. 11.S. Ben-David, J. Blitzer, K. Crammer, and F. Pereira. 2007. Analysis of representations for domain adaptation. In Advances in Neural Information Processing Systems. 137–144.
  12. 12.S. Benaim and L. Wolf. 2018. One-shot unsupervised cross domain translation. In Advances in Neural Information Processing Systems. 2104–2114.
  13. 13.L. Bertinetto, J. F. Henriques, P. Torr, and A. Vedaldi. 2019. Meta-learning with differentiable closed-form solvers. In International Conference on Learning Representations.
  14. 14.L. Bertinetto, J. F. Henriques, J. Valmadre, P. Torr, and A. Vedaldi. 2016. Learning feed-forward one-shot learners. In Advances in Neural Information Processing Systems. 523–531.
  15. 15.C. M. Bishop. 2006. Pattern Recognition and Machine Learning. Springer.
  16. 16.J. Blitzer, K. Crammer, A. Kulesza, F. Pereira, and J. Wortman. 2008. Learning bounds for domain adaptation. In Advances in Neural Information Processing Systems. 129–136.
  17. 17.L. Bottou and O. Bousquet. 2008. The tradeoffs of large scale learning. In Advances in Neural Information Processing Systems. 161–168.
  18. 18.L. Bottou, F. E. Curtis, and J. Nocedal. 2018. Optimization methods for large-scale machine learning. SIAM Rev. 60, 2 (2018), 223–311.
  19. 19.A. Brock, T. Lim, J.M. Ritchie, and N. Weston. 2018. SMASH: One-shot model architecture search through hypernetworks. In International Conference on Learning Representations.
  20. 20.J. Bromley, I. Guyon, Y. LeCun, E. Säckinger, and R. Shah. 1994. Signature verification using a "siamese" time delay neural network. In Advances in Neural Information Processing Systems. 737–744.
  21. 21.S. Caelles, K.-K. Maninis, J. Pont-Tuset, L. Leal-Taixé, D. Cremers, and L. Van Gool. 2017. One-shot video object segmentation. In Conference on Computer Vision and Pattern Recognition. 221–230.
  22. 22.Q. Cai, Y. Pan, T. Yao, C. Yan, and T. Mei. 2018. Memory matching networks for one-shot image recognition. In Conference on Computer Vision and Pattern Recognition. 4080–4088.
  23. 23.R. Caruana. 1997. Multitask learning. Machine learning 28, 1 (1997), 41–75.
  24. 24.J. Choi, J. Krishnamurthy, A. Kembhavi, and A. Farhadi. 2018. Structured set matching networks for one-shot part labeling. In Conference on Computer Vision and Pattern Recognition. 3627–3636.
  25. 25.J. D. Co-Reyes, A. Gupta, S. Sanjeev, N. Altieri, J. DeNero, P. Abbeel, and S. Levine. 2019. Meta-learning language-guided policy learning. In International Conference on Learning Representations.
  26. 26.J. J. Craig. 2009. Introduction to Robotics: Mechanics and Control. Pearson Education India.
  27. 27.E. D. Cubuk, B. Zoph, D. Mane, V. Vasudevan, and Q. V. Le. 2019. AutoAugment: Learning augmentation policies from data. In Conference on Computer Vision and Pattern Recognition. 113–123.
  28. 28.T. Deleu and Y. Bengio. 2018. The effects of negative adaptation in Model-Agnostic Meta-Learning. arXiv preprint arXiv:1812.02159 (2018).
  29. 29.G. Denevi, C. Ciliberto, D. Stamos, and M. Pontil. 2018. Learning to learn around a common mean. In Advances in Neural Information Processing Systems. 10190–10200.
  30. 30.J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei. 2009. ImageNet: A large-scale hierarchical image database. In Conference on Computer Vision and Pattern Recognition. 248–255.
  31. 31.X. Dong, L. Zhu, D. Zhang, Y. Yang, and F. Wu. 2018. Fast parameter adaptation for few-shot image captioning and visual question answering. In ACM International Conference on Multimedia. 54–62.
  32. 32.M. Douze, A. Szlam, B. Hariharan, and H. Jégou. 2018. Low-shot learning with large-scale diffusion. In Conference on Computer Vision and Pattern Recognition. 3349–3358.
  33. 33.Y. Duan, M. Andrychowicz, B. Stadie, J. Ho, J. Schneider, I. Sutskever, P. Abbeel, and W. Zaremba. 2017. One-shot imitation learning. In Advances in Neural Information Processing Systems. 1087–1098.
  34. 34.H. Edwards and A. Storkey. 2017. Towards a neural statistician. In International Conference on Learning Representations.
  35. 35.L. Fei-Fei, R. Fergus, and P. Perona. 2006. One-shot learning of object categories. IEEE Transactions on Pattern Analysis and Machine Intelligence 28, 4 (2006), 594–611.
  36. 36.M. Fink. 2005. Object classification from a single example utilizing class relevance metrics. In Advances in Neural Information Processing Systems. 449–456.
  37. 37.C. Finn, P. Abbeel, and S. Levine. 2017. Model-agnostic meta-learning for fast adaptation of deep networks. In International Conference on Machine Learning. 1126–1135.
  38. 38.C. Finn and S. Levine. 2018. Meta-learning and universality: Deep representations and gradient descent can approximate any learning algorithm. In International Conference on Learning Representations.
  39. 39.C. Finn, K. Xu, and S. Levine. 2018. Probabilistic model-agnostic meta-learning. In Advances in Neural Information Processing Systems. 9537–9548.
  40. 40.L. Franceschi, P. Frasconi, S. Salzo, R. Grazzi, and M. Pontil. 2018. Bilevel programming for hyperparameter optimization and meta-learning. In International Conference on Machine Learning. 1563–1572.
  41. 41.J. Friedman, T. Hastie, and R. Tibshirani. 2001. The Elements of Statistical Learning. Vol. 1. Springer series in statistics New York.
  42. 42.H. Gao, Z. Shou, A. Zareian, H. Zhang, and S. Chang. 2018. Low-shot learning via covariance-preserving adversarial augmentation networks. In Advances in Neural Information Processing Systems. 983–993.
  43. 43.P. Germain, F. Bach, A. Lacoste, and S. Lacoste-Julien. 2016. PAC-Bayesian theory meets Bayesian inference. In Advances in Neural Information Processing Systems. 1884–1892.
  44. 44.S. Gidaris and N. Komodakis. 2018. Dynamic few-shot visual learning without forgetting. In Conference on Computer Vision and Pattern Recognition. 4367–4375.
  45. 45.I. Goodfellow, Y. Bengio, and A. Courville. 2016. Deep Learning. MIT Press.
  46. 46.I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio. 2014. Generative adversarial nets. In Advances in Neural Information Processing Systems. 2672–2680.
  47. 47.J. Gordon, J. Bronskill, M. Bauer, S. Nowozin, and R. Turner. 2019. Meta-learning probabilistic inference for prediction. In International Conference on Learning Representations.
  48. 48.E. Grant, C. Finn, S. Levine, T. Darrell, and T. Griffiths. 2018. Recasting gradient-based meta-learning as hierarchical Bayes. In International Conference on Learning Representations.
  49. 49.A. Graves, G. Wayne, and I. Danihelka. 2014. Neural Turing machines. arXiv preprint arXiv:1410.5401 (2014).
  50. 50.L.-Y. Gui, Y.-X. Wang, D. Ramanan, and J. Moura. 2018. Few-shot human motion prediction via meta-learning. In European Conference on Computer Vision. 432–450.
  51. 51.M. Hamaya, T. Matsubara, T. Noda, T. Teramae, and J. Morimoto. 2016. Learning assistive strategies from a few user-robot interactions: Model-based reinforcement learning approach. In International Conference on Robotics and Automation. 3346–3351.
  52. 52.X. Han, H. Zhu, P. Yu, Z. Wang, Y. Yao, Z. Liu, and M. Sun. 2018. FewRel: A large-scale supervised few-shot relation classification dataset with state-of-the-art evaluation. In Conference on Empirical Methods in Natural Language Processing. 4803–4809.
  53. 53.B. Hariharan and R. Girshick. 2017. Low-shot visual recognition by shrinking and hallucinating features. In International Conference on Computer Vision.
  54. 54.H. He and E. A. Garcia. 2008. Learning from imbalanced data. IEEE Transactions on Knowledge and Data Engineering 9 (2008), 1263–1284.
  55. 55.K. He, X. Zhang, S. Ren, and J. Sun. 2016. Deep residual learning for image recognition. In Conference on Computer Vision and Pattern Recognition. 770–778.
  56. 56.A. Herbelot and M. Baroni. 2017. High-risk learning: Acquiring new word vectors from tiny data. In Conference on Empirical Methods in Natural Language Processing. 304–309.
  57. 57.L. B. Hewitt, M. I. Nye, A. Gane, T. Jaakkola, and J. B. Tenenbaum. 2018. The variational homoencoder: Learning to learn high capacity generative models from few examples. In Uncertainty in Artificial Intelligence. 988–997.
  58. 58.S. Hochreiter and J. Schmidhuber. 1997. Long short-term memory. Neural Computation 9, 8 (1997), 1735–1780.
  59. 59.S. Hochreiter, A. S. Younger, and P. R. Conwell. 2001. Learning to learn using gradient descent. In International Conference on Artificial Neural Networks. 87–94.
  60. 60.J. Hoffman, E. Tzeng, J. Donahue, Y. Jia, K. Saenko, and T. Darrell. 2013. One-shot adaptation of supervised deep convolutional models. In International Conference on Learning Representations.
  61. 61.Z. Hu, X. Li, C. Tu, Z. Liu, and M. Sun. 2018. Few-shot charge prediction with discriminative legal attributes. In International Conference on Computational Linguistics. 487–498.
  62. 62.S. J. Hwang and L. Sigal. 2014. A unified semantic embedding: Relating taxonomies and attributes. In Advances in Neural Information Processing Systems. 271–279.
  63. 63.Y. Jia, E. Shelhamer, J. Donahue, S. Karayev, J. Long, R. Girshick, S. Guadarrama, and T. Darrell. 2014. Caffe: Convolutional architecture for fast feature embedding. In ACM International Conference on Multimedia. 675–678.
  64. 64.V. Joshi, M. Peters, and M. Hopkins. 2018. Extending a parser to distant domains using a few dozen partially annotated examples. In Annual Meeting of the Association for Computational Linguistics. 1190–1199.
  65. 65.Ł. Kaiser, O. Nachum, A. Roy, and S. Bengio. 2017. Learning to remember rare events. In International Conference on Learning Representations.
  66. 66.J. M. Kanter and K. Veeramachaneni. 2015. Deep feature synthesis: Towards automating data science endeavors. In International Conference on Data Science and Advanced Analytics. 1–10.
  67. 67.R. Keshari, M. Vatsa, R. Singh, and A. Noore. 2018. Learning structure and strength of CNN filters for small sample size training. In Conference on Computer Vision and Pattern Recognition. 9349–9358.
  68. 68.D. P. Kingma and M. Welling. 2014. Auto-encoding variational Bayes. In International Conference on Learning Representations.
  69. 69.J. Kirkpatrick, R. Pascanu, N. Rabinowitz, J. Veness, G. Desjardins, A. A. Rusu, K. Milan, J. Quan, T. Ramalho, A. Grabska-Barwinska, et al. 2017. Overcoming catastrophic forgetting in neural networks. National Academy of Sciences 114, 13 (2017), 3521–3526.
  70. 70.G. Koch. 2015. Siamese neural networks for one-shot image recognition. Ph.D. Dissertation. University of Toronto.
  71. 71.L. Kotthoff, C. Thornton, H. H. Hoos, F. Hutter, and K. Leyton-Brown. 2017. Auto-WEKA 2.0: Automatic model selection and hyperparameter optimization in WEKA. Journal of Machine Learning Research 18, 1 (2017), 826–830.
  72. 72.J. Kozerawski and M. Turk. 2018. CLEAR: Cumulative learning for one-shot one-class image recognition. In Conference on Computer Vision and Pattern Recognition. 3446–3455.
  73. 73.A. Krizhevsky, I. Sutskever, and G. E. Hinton. 2012. ImageNet classification with deep convolutional neural networks. In Advances in Neural Information Processing Systems. 1097–1105.
  74. 74.R. Kwitt, S. Hegenbart, and M. Niethammer. 2016. One-shot learning of scene locations via feature trajectory transfer. In Conference on Computer Vision and Pattern Recognition. 78–86.
  75. 75.B. Lake, C.-Y. Lee, J. Glass, and J. Tenenbaum. 2014. One-shot learning of generative speech concepts. In Annual Meeting of the Cognitive Science Society, Vol. 36.
  76. 76.B. M. Lake, R. Salakhutdinov, and J. B. Tenenbaum. 2015. Human-level concept learning through probabilistic program induction. Science 350, 6266 (2015), 1332–1338.
  77. 77.B. M. Lake, T. D. Ullman, J. B. Tenenbaum, and S. J. Gershman. 2017. Building machines that learn and think like people. Behavioral and Brain Sciences 40 (2017).
  78. 78.C. H. Lampert, H. Nickisch, and S. Harmeling. 2009. Learning to detect unseen object classes by between-class attribute transfer. In Conference on Computer Vision and Pattern Recognition. 951–958.
  79. 79.Y. Lee and S. Choi. 2018. Gradient-based meta-learning with learned layerwise metric and subspace. In International Conference on Machine Learning. 2933–2942.
  80. 80.K. Li and J. Malik. 2017. Learning to optimize. In International Conference on Learning Representations.
  81. 81.X.-L. Li, P. S. Yu, B. Liu, and S.-K. Ng. 2009. Positive unlabeled learning for data stream classification. In SIAM International Conference on Data Mining. 259–270.
  82. 82.B. Liu, X. Wang, M. Dixit, R. Kwitt, and N. Vasconcelos. 2018. Feature space transfer for data augmentation. In Conference on Computer Vision and Pattern Recognition. 9090–9098.
  83. 83.H. Liu, K. Simonyan, and Y. Yang. 2019. DARTS: Differentiable architecture search. In International Conference on Learning Representations.
  84. 84.Y. Liu, J. Lee, M. Park, S. Kim, E. Yang, S. Hwang, and Y Yang. 2019. Learning to propopagate labels: Transductive propagation network for few-shot learning. In International Conference on Learning Representations.
  85. 85.Z. Luo, Y. Zou, J. Hoffman, and L. Fei-Fei. 2017. Label efficient learning of transferable representations acrosss domains and tasks. In Advances in Neural Information Processing Systems. 165–177.
  86. 86.S. Mahadevan and P. Tadepalli. 1994. Quantifying prior determination knowledge using the PAC learning model. Machine Learning 17, 1 (1994), 69–105.
  87. 87.D. McNamara and M.-F. Balcan. 2017. Risk bounds for transferring representations with and without fine-tuning. In International Conference on Machine Learning. 2373–2381.
  88. 88.T. Mensink, E. Gavves, and C. Snoek. 2014. Costa: Co-occurrence statistics for zero-shot classification. In Conference on Computer Vision and Pattern Recognition. 2441–2448.
  89. 89.A. Miller, A. Fisch, J. Dodge, A.-H. Karimi, A. Bordes, and J. Weston. 2016. Key-value memory networks for directly reading documents. In Conference on Empirical Methods in Natural Language Processing. 1400–1409.
  90. 90.E. G. Miller, N. E. Matsakis, and P. A. Viola. 2000. Learning from one example through shared densities on transforms. In Conference on Computer Vision and Pattern Recognition, Vol. 1. 464–471.
  91. 91.N. Mishra, M. Rohaninejad, X. Chen, and P. Abbeel. 2018. A simple neural attentive meta-learner. In International Conference on Learning Representations.
  92. 92.M. T. Mitchell. 1997. Machine Learning. McGraw-Hill.
  93. 93.S. H. Mohammadi and T. Kim. 2018. Investigation of using disentangled and interpretable representations for one-shot cross-lingual voice conversion. In INTERSPEECH. 2833–2837.
  94. 94.M. Mohri, A. Rostamizadeh, and A. Talwalkar. 2018. Foundations of machine learning. MIT Press.
  95. 95.S. Motiian, Q. Jones, S. Iranmanesh, and G. Doretto. 2017. Few-shot adversarial domain adaptation. In Advances in Neural Information Processing Systems. 6670–6680.
  96. 96.T. Munkhdalai and H. Yu. 2017. Meta networks. In International Conference on Machine Learning. 2554–2563.
  97. 97.T. Munkhdalai, X. Yuan, S. Mehri, and A. Trischler. 2018. Rapid adaptation with conditionally shifted neurons. In International Conference on Machine Learning. 3661–3670.
  98. 98.A. Nagabandi, C. Finn, and S. Levine. 2018. Deep online learning via meta-learning: Continual adaptation for model-based RL. In International Conference on Learning Representations.
  99. 99.H. Nguyen and L. Zakynthinou. 2018. Improved algorithms for collaborative PAC learning. In Advances in Neural Information Processing Systems. 7631–7639.
  100. 100.B. Oreshkin, P. R. López, and A. Lacoste. 2018. TADAM: Task dependent adaptive metric for improved few-shot learning. In Advances in Neural Information Processing Systems. 719–729.
  101. 101.S. J. Pan and Q. Yang. 2010. A survey on transfer learning. IEEE Transactions on Knowledge and Data Engineering 10, 22 (2010), 1345–1359.
  102. 102.T. Pfister, J. Charles, and A. Zisserman. 2014. Domain-adaptive discriminative one-shot learning of gestures. In European Conference on Computer Vision. 814–829.
  103. 103.H. Qi, M. Brown, and D. G. Lowe. 2018. Low-shot learning with imprinted weights. In Conference on Computer Vision and Pattern Recognition. 5822–5830.
  104. 104.T. Ramalho and M. Garnelo. 2019. Adaptive posterior learning: Few-shot learning with a surprise-based memory module. In International Conference on Learning Representations.
  105. 105.S. Ravi and A. Beatson. 2019. Amortized Bayesian meta-learning. In International Conference on Learning Representations.
  106. 106.S. Ravi and H. Larochelle. 2017. Optimization as a model for few-shot learning. In International Conference on Learning Representations.
  107. 107.S. Reed, Y. Chen, T. Paine, A. van den Oord, S. M. A. Eslami, D. Rezende, O. Vinyals, and N. de Freitas. 2018. Few-shot autoregressive density estimation: Towards learning to learn distributions. In International Conference on Learning Representations.
  108. 108.M. Ren, S. Ravi, E. Triantafillou, J. Snell, K. Swersky, J. B. Tenenbaum, H. Larochelle, and R. S. Zemel. 2018. Meta-learning for semi-supervised few-shot classification. In International Conference on Learning Representations.
  109. 109.D. Rezende, I. Danihelka, K. Gregor, and D. Wierstra. 2016. One-shot generalization in deep generative models. In International Conference on Machine Learning. 1521–1529.
  110. 110.A. Rios and R. Kavuluru. 2018. Few-shot and zero-shot multi-label learning for structured label spaces. In Conference on Empirical Methods in Natural Language Processing. 3132.
  111. 111.A. A. Rusu, D. Rao, J. Sygnowski, O. Vinyals, R. Pascanu, S. Osindero, and R. Hadsell. 2019. Meta-learning with latent embedding optimization. In International Conference on Learning Representations.
  112. 112.R. Salakhutdinov and G. Hinton. 2009. Deep boltzmann machines. In International Conference on Artificial Intelligence and Statistics. 448–455.
  113. 113.R. Salakhutdinov, J. Tenenbaum, and A. Torralba. 2012. One-shot learning with a hierarchical nonparametric Bayesian model. In ICML Workshop on Unsupervised and Transfer Learning. 195–206.
  114. 114.A. Santoro, S. Bartunov, M. Botvinick, D. Wierstra, and T. Lillicrap. 2016. Meta-learning with memory-augmented neural networks. In International Conference on Machine Learning. 1842–1850.
  115. 115.V. G. Satorras and J. B. Estrach. 2018. Few-shot learning with graph neural networks. In International Conference on Learning Representations.
  116. 116.E. Schwartz, L. Karlinsky, J. Shtok, S. Harary, M. Marder, A. Kumar, R. Feris, R. Giryes, and A. Bronstein. 2018. Delta-encoder: An effective sample synthesis method for few-shot object recognition. In Advances in Neural Information Processing Systems. 2850–2860.
  117. 117.B. Settles. 2009. Active learning literature survey. Technical Report. University of Wisconsin-Madison Department of Computer Sciences.
  118. 118.J. Shu, Z. Xu, and D Meng. 2018. Small sample learning in big data era. arXiv preprint arXiv:1808.04572 (2018).
  119. 119.P. Shyam, S. Gupta, and A. Dukkipati. 2017. Attentive recurrent comparators. In International Conference on Machine Learning. 3173–3181.
  120. 120.D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. Van Den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, et al. 2016. Mastering the game of Go with deep neural networks and tree search. Nature 529, 7587 (2016), 484–489.
  121. 121.J. Snell, K. Swersky, and R. S. Zemel. 2017. Prototypical networks for few-shot learning. In Advances in Neural Information Processing Systems. 4077–4087.
  122. 122.M. D. Spivak. 1970. A Comprehensive Introduction to Differential Geometry. Publish or Perish.
  123. 123.R. K. Srivastava, K. Greff, and J. Schmidhuber. 2015. Training very deep networks. In Advances in Neural Information Processing Systems. 2377–2385.
  124. 124.S. Sukhbaatar, J. Weston, R. Fergus, et al. 2015. End-to-end memory networks. In Advances in Neural Information Processing Systems. 2440–2448.
  125. 125.J. Sun, S. Wang, and C. Zong. 2018. Memory, show the way: Memory based few shot word representation learning. In Conference on Empirical Methods in Natural Language Processing. 1435–1444.
  126. 126.F. Sung, Y. Yang, L. Zhang, T. Xiang, P. H. Torr, and T. M. Hospedales. 2018. Learning to compare: Relation network for few-shot learning. In Conference on Computer Vision and Pattern Recognition. 1199–1208.
  127. 127.K. D. Tang, M. F. Tappen, R. Sukthankar, and C. H. Lampert. 2010. Optimizing one-shot recognition with micro-set learning. In Conference on Computer Vision and Pattern Recognition. 3027–3034.
  128. 128.A. Tjandra, S. Sakti, and S. Nakamura. 2018. Machine speech chain with one-shot speaker adaptation. In INTERSPEECH. 887–891.
  129. 129.A. Torralba, J. B. Tenenbaum, and R. R. Salakhutdinov. 2011. Learning to learn with compound HD models. In Advances in Neural Information Processing Systems. 2061–2069.
  130. 130.E. Triantafillou, R. Zemel, and R. Urtasun. 2017. Few-shot learning through an information retrieval lens. In Advances in Neural Information Processing Systems. 2255–2265.
  131. 131.E. Triantafillou, T. Zhu, V. Dumoulin, P. Lamblin, K. Xu, R. Goroshin, C. Gelada, K. Swersky, P.-A. Manzagol, et al. 2019. Meta-dataset: A dataset of datasets for learning to learn from few examples. arXiv preprint arXiv:1903.03096 (2019).
  132. 132.Y.-H. Tsai, L.-K. Huang, and R. Salakhutdinov. 2017. Learning robust visual-semantic embeddings. In Conference on Computer Vision and Pattern Recognition. 3571–3580.
  133. 133.Y. H. Tsai and R. Salakhutdinov. 2017. Improving one-shot learning through fusing side information. arXiv preprint arXiv:1710.08347 (2017).
  134. 134.M. A. Turing. 1950. Computing machinery and intelligence. Mind 59, 236 (1950), 433–433.
  135. 135.A. Van den Oord, N. Kalchbrenner, L. Espeholt, O. Vinyals, A. Graves, et al. 2016. Conditional image generation with PixelCNN decoders. In Advances in Neural Information Processing Systems. 4790–4798.
  136. 136.V. N. Vapnik. 1992. Principles of risk minimization for learning theory. In Advances in Neural Information Processing Systems. 831–838.
  137. 137.M. Vartak, A. Thiagarajan, C. Miranda, J. Bratman, and H. Larochelle. 2017. A meta-learning perspective on cold-start recommendations for items. In Advances in Neural Information Processing Systems. 6904–6914.
  138. 138.O. Vinyals, C. Blundell, T. Lillicrap, D. Wierstra, et al. 2016. Matching networks for one shot learning. In Advances in Neural Information Processing Systems. 3630–3638.
  139. 139.P. Wang, L. Liu, C. Shen, Z. Huang, A. van den Hengel, and H. Tao Shen. 2017. Multi-attention network for one shot learning. In Conference on Computer Vision and Pattern Recognition. 2721–2729.
  140. 140.X. Wang, Y. Ye, and A. Gupta. 2018. Zero-shot recognition via semantic embeddings and knowledge graphs. In Conference on Computer Vision and Pattern Recognition. 6857–6866.
  141. 141.Y.-X. Wang, R. Girshick, M. Hebert, and B. Hariharan. 2018. Low-shot learning from imaginary data. In Conference on Computer Vision and Pattern Recognition. 7278–7286.
  142. 142.Y.-X. Wang and M. Hebert. 2016. Learning from small sample sets by combining unsupervised meta-training with CNNs. In Advances in Neural Information Processing Systems. 244–252.
  143. 143.Y.-X. Wang and M. Hebert. 2016. Learning to learn: Model regression networks for easy small sample learning. In European Conference on Computer Vision. 616–634.
  144. 144.J. Wei and K. Zou. 2019. EDA: Easy data augmentation techniques for boosting performance on text classification tasks. In Conference on Empirical Methods in Natural Language Processing and International Joint Conference on Natural Language Processing. 6383–6389.
  145. 145.J. Weston, S. Chopra, and A. Bordes. 2014. Memory networks. arXiv preprint arXiv:1410.3916 (2014).
  146. 146.M. Woodward and C. Finn. 2017. Active one-shot learning. arXiv preprint arXiv:1702.06559 (2017).
  147. 147.Y. Wu and Y. Demiris. 2010. Towards one shot learning by imitation for humanoid robots. In International Conference on Robotics and Automation. 2889–2894.
  148. 148.Y. Wu, Y. Lin, X. Dong, Y. Yan, W. Ouyang, and Y. Yang. 2018. Exploit the unknown gradually: One-shot video-based person re-identification by stepwise learning. In Conference on Computer Vision and Pattern Recognition. 5177–5186.
  149. 149.Z. Xu, L. Zhu, and Y. Yang. 2017. Few-shot object recognition from machine-labeled web images. In Conference on Computer Vision and Pattern Recognition. 1164–1172.
  150. 150.L. Yan, Y. Zheng, and J. Cao. 2018. Few-shot learning for short text classification. Multimedia Tools and Applications 77, 22 (2018), 29799–29810.
  151. 151.W. Yan, J. Yap, and G. Mori. 2015. Multi-task transfer methods to improve one-shot learning for multimedia event detection. In British Machine Vision Conference.
  152. 152.H. Yang, X. He, and F. Porikli. 2018. One-shot action localization by learning sequence matching network. In Conference on Computer Vision and Pattern Recognition. 1450–1459.
  153. 153.Q. Yao, M. Wang, E. H. Jair, I. Guyon, Y.-Q. Hu, Y.-F. Li, W.-W. Tu, Q. Yang, and Y. Yu. 2018. Taking human out of learning applications: A survey on automated machine learning. arXiv preprint arXiv:1810.13306 (2018).
  154. 154.Q. Yao, J. Xu, W.-W. Tu, and Z. Zhu. 2020. Efficient neural architecture search via proximal iterations. In AAAI Conference on Artificial Intelligence.
  155. 155.D. Yoo, H. Fan, V. N. Boddeti, and K. M. Kitani. 2018. Efficient k-shot learning with regularized deep networks. In AAAI Conference on Artificial Intelligence.
  156. 156.J. Yoon, T. Kim, O. Dia, S. Kim, Y. Bengio, and S. Ahn. 2018. Bayesian model-agnostic meta-learning. In Advances in Neural Information Processing Systems. 7343–7353.
  157. 157.M. Yu, X. Guo, J. Yi, S. Chang, S. Potdar, Y. Cheng, G. Tesauro, H. Wang, and B. Zhou. 2018. Diverse few-shot text classification with multiple metrics. In Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 1206–1215.
  158. 158.C. Zhang, J. Butepage, H. Kjellstrom, and S. Mandt. 2019. Advances in variational inference. IEEE Transactions on Pattern Analysis and Machine Intelligence 41, 8 (2019), 2008–2026.
  159. 159.R. Zhang, T. Che, Z. Ghahramani, Y. Bengio, and Y. Song. 2018. MetaGAN: An adversarial approach to few-shot learning. In Advances in Neural Information Processing Systems. 2371–2380.
  160. 160.Y. Zhang, H. Tang, and K. Jia. 2018. Fine-grained visual categorization using meta-learning optimization with sample selection of auxiliary data. In European Conference on Computer Vision. 233–248.
  161. 161.Y. Zhang and Q. Yang. 2017. A survey on multi-task learning. arXiv preprint arXiv:1707.08114 (2017).
  162. 162.F. Zhao, J. Zhao, S. Yan, and J. Feng. 2018. Dynamic conditional networks for few-shot learning. In European Conference on Computer Vision.
  163. 163.Z.-H. Zhou. 2017. A brief introduction to weakly supervised learning. National Science Review 5, 1 (2017), 44–53.
  164. 164.L. Zhu and Y. Yang. 2018. Compound memory networks for few-shot video classification. In European Conference on Computer Vision. 751–766.
  165. 165.X. J. Zhu. 2005. Semi-supervised learning literature survey. Technical Report. University of Wisconsin-Madison Department of Computer Sciences.
  166. 166.B. Zoph and Q. V. Le. 2017. Neural architecture search with reinforcement learning. In International Conference on Learning Representations.

Citation

MLA
Wang, Y., et al. “Generalizing from a Few Examples: A Survey on Few-Shot Learning”. arXiv, 2019, http://arxiv.org/abs/1904.05046v3.
APA
Wang, Y., Yao, Q., Kwok, J., & Ni, L. M. (2019). Generalizing from a Few Examples: A Survey on Few-Shot Learning. arXiv. http://arxiv.org/abs/1904.05046v3
Chicago
Wang, Y., Q. Yao, J. Kwok, and L. M. Ni. 2019. “Generalizing from a Few Examples: A Survey on Few-Shot Learning”. arXiv. http://arxiv.org/abs/1904.05046v3.
Harvard
Wang, Y. et al. (2019) “Generalizing from a Few Examples: A Survey on Few-Shot Learning”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1904.05046v3.
Vancouver
1. Wang Y, Yao Q, Kwok J, Ni LM (2019) Generalizing from a Few Examples: A Survey on Few-Shot Learning. arXiv

BibTeX

@article{wang2019generalizing,
  title = {Generalizing from a Few Examples: A Survey on Few-Shot Learning},
  author = {Wang, Yaqing and Yao, Quanming and Kwok, James and Ni, Lionel M.},
  year = {2019},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1904.05046v3},
  eprint = {1904.05046}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF