Meta-Learning with Latent Embedding Optimization

Andrei A. RusuDushyant RaoJakub SygnowskiOriol VinyalsRazvan PascanuSimon OsinderoRaia Hadsell

article2018ICLR1,530 citations

Proposes Latent Embedding Optimization to overcome the scaling limitations of few-shot meta-learning by executing gradient-based model adaptation within a low-dimensional generative latent space rather than high-dimensional parameter space.

Listen

Modern artificial intelligence systems struggle to rapidly adapt to new concepts when only a handful of examples are available, in contrast to humans who learn new tasks efficiently from limited experience. Traditional optimization-based meta-learning methods, or algorithms that learn how to learn, attempt to address this few-shot learning challenge by finding a single shared parameter initialization across all tasks and fine-tuning it directly using gradient steps. However, performing gradient-based optimization directly within high-dimensional parameter spaces on tiny data samples often leads to overfitting, poor generalization, and severe optimization instability.

The article evaluates whether decoupling the gradient adaptation process from high-dimensional parameter space into a learned, low-dimensional latent embedding can overcome these limitations. The primary objective is to demonstrate a framework called Latent Embedding Optimization, which generates data-dependent model initializations and executes adaptation directly within a compact latent space to achieve superior few-shot learning performance.

The approach was evaluated across synthetic regression tasks and large-scale image classification benchmarks, specifically miniImageNet and tieredImageNet under one-shot and five-shot settings. Credibility was established by training on top of frozen visual feature representations extracted from a deep Wide Residual Network, running evaluations over 50,000 problem instances across multiple independent random seeds, and conducting systematic ablation studies to isolate individual architectural components.

The analysis produced several vital findings. First, the proposed approach established new state-of-the-art accuracy benchmarks on few-shot image classification, achieving 61.76% on one-shot and 77.59% on five-shot miniImageNet, as well as 66.33% on one-shot and 81.44% on five-shot tieredImageNet, noticeably outperforming previous leading methods and direct parameter optimization baselines like Meta-SGD. Second, curvature and coverage measurements showed that small gradient steps in the compact latent space induced substantial, structured movements in the parameter space, with latent space curvature two orders of magnitude higher than parameter space curvature. Third, the framework successfully modeled parametric uncertainty and captured distinct solution types in ambiguous regression tests where data could be explained by multiple distinct functional forms. Finally, ablation experiments proved that both the data-dependent initialization and the inner-loop latent space adaptation were essential, as removing either component significantly degraded accuracy.

These findings demonstrate that restructuring optimization-based meta-learning into a low-dimensional bottleneck significantly improves model generalization while reducing computational sensitivity. By separating the high-cost feature representation pre-training from the lightweight latent adaptation training, organizations can achieve rapid task adaptation with highly efficient meta-training cycles that require only hours on standard hardware rather than extensive compute clusters.

For technical leaders seeking to deploy rapid-adaptation machine learning systems, the evidence supports adopting latent-space optimization combined with modular pre-trained feature extractors over end-to-end direct parameter tuning. Future development should evaluate joint end-to-end meta-learning of feature representations, as well as testing latent optimization techniques on sequential and reinforcement learning domains to expand practical applicability.

Confidence in these findings is high for standard few-shot computer vision benchmarks, but stakeholders should note key limitations. The image classification results rely heavily on the quality of a pre-trained feature extractor, and empirical evaluations were constrained to controlled academic datasets, meaning performance may vary on messy, real-world operational data distributions.

arXiv: 1807.05960
  • Paper: A Closer Look at Few-shot Classification, Wei-Yu Chen et al. (2019). It critically analyzes the standardized few-shot benchmarks used by methods like LEO, evaluating how backbone capacity and domain shifts impact meta-learning performance.
  • Paper: On First-Order Meta-Learning Algorithms, Alex Nichol et al. (2018). It provides a theoretical and algorithmic analysis of first-order gradient meta-learning approximations like Reptile as an alternative method for scaling optimization-based few-shot learners.
Cover for Meta-Learning with Latent Embedding Optimization

Abstract

Gradient-based meta-learning techniques are both widely applicable and proficient at solving challenging few-shot learning and fast adaptation problems. However, they have practical difficulties when operating on high-dimensional parameter spaces in extreme low-data regimes. We show that it is possible to bypass these limitations by learning a data-dependent latent generative representation of model parameters, and performing gradient-based meta-learning in this low-dimensional latent space. The resulting approach, latent embedding optimization (LEO), decouples the gradient-based adaptation procedure from the underlying high-dimensional space of model parameters. Our evaluation shows that LEO can achieve state-of-the-art performance on the competitive miniImageNet and tieredImageNet few-shot classification tasks. Further analysis indicates LEO is able to capture uncertainty in the data, and can perform adaptation more effectively by optimizing in latent space.

Table of Contents

  • 1 Introduction
  • 2 Model
  • 2.1 Problem Definition
  • 2.2 Model-Agnostic Meta-Learning
  • 2.3 Latent Embedding Optimization for Meta-Learning
  • 2.3.1 Model overview
  • 2.3.2 Initialization: Generating Parameters Conditioned on a Few Examples
  • 2.3.3 Adaptation by Latent Embedding Optimization (LEO) (The “Inner Loop”)
  • 2.3.4 Meta-Training Strategy (The “Outer Loop”)
  • 2.3.5 Beyond Classification and Linear Output Layers
  • 3 Related Work
  • 4 Evaluation
  • 4.1 Few-shot Regression
  • 4.2 Few-shot Classification
  • 4.2.1 Datasets
  • 4.2.2 Pre-trained Features
  • 4.2.3 Fine-tuning
  • 4.3 Results
  • 4.4 Ablation Study
  • 4.5 Latent Embedding Visualization
  • 4.6 Curvature and coverage analysis
  • 5 Conclusions and Future Work
  • References
  • A Experimental setup - regression
  • A.1 Regression task description
  • A.2 LEO Network Architecture
  • B Experimental setup - classification
  • B.1 Data preparation
  • B.2 Feature Pre-Training
  • B.3 LEO Network Architecture
  • B.4 Optimization
  • B.5 Hyper-parameters
  • B.6 Training time
  • B.7 Overview of the training procedure
  • B.8 Overview of the evaluation procedure

Knowls

  1. Knowl 1 — Latent Embedding Optimization Framework

    model/method

    Latent Embedding Optimization (LEO) is an optimization-based meta-learning methodology designed for few-shot learning tasks that bypasses the instability and high variance of computing gradient updates directly in high-dimensional model parameter spaces Θ\Theta. Instead of optimizing a single shared initialization θ∈Θ\theta \in \Theta across tasks as in Model-Agnostic Meta-Learning (MAML), LEO learns a low-dimensional stochastic latent embedding space Z=Rnz\mathcal{Z} = \mathbb{R}^{n_z} (where nz≪dim⁡(Θ)n_z \ll \dim(\Theta)) conditioned on the task's training data Dtr\mathcal{D}^{tr}, from which high-dimensional parameters θ\theta are generated by a decoder.

    Task adaptation in LEO proceeds through three interconnected phases:

    1. Data-Dependent Initialization: Given task training data Dtr\mathcal{D}^{tr}, an encoder combined with a relation network maps the inputs into a class-conditional distribution over latent codes z∈Zz \in \mathcal{Z}. Sampling from this distribution produces an initial latent representation zz, which is then decoded by a generative decoder network into the initial task model parameters θ\theta.
    2. Latent Inner-Loop Optimization: Rather than performing gradient descent on θ\theta, the task-specific training loss Ltr(fθ)\mathcal{L}^{tr}(f_\theta) is backpropagated through the decoder to compute gradients directly with respect to the latent code zz. Gradient steps are performed in latent space (z′=z−α∇zLtr(fθ)z' = z - \alpha \nabla_z \mathcal{L}^{tr}(f_\theta)), and the adapted latent code z′z' is decoded to generate updated task parameters θ′\theta'.
    3. Outer-Loop Meta-Training: The performance of the adapted model fθ′f_{\theta'} is evaluated on the task validation set Dval\mathcal{D}^{val}, and the outer-loop meta-objective differentiates through the inner adaptation steps to optimize the parameters of the encoder, relation network, decoder, and meta-learned inner learning rate α\alpha.
  2. Knowl 2 — Data-Dependent Latent Encoding and Parameter Generation in LEO

    equation

    For an NN-way KK-shot classification problem where class n∈{1,…,N}n \in \{1, \dots, N\} has training instances Dntr={(xnk,ynk)∣k=1,…,K}\mathcal{D}_n^{tr} = \{(x_n^k, y_n^k) \mid k=1, \dots, K\}, LEO constructs class-dependent latent vectors zn∈Rnzz_n \in \mathbb{R}^{n_z} and softmax classifier weights wn∈Rnxw_n \in \mathbb{R}^{n_x} using an encoder gϕe:Rnx→Rnhg_{\phi_e}: \mathbb{R}^{n_x} \to \mathbb{R}^{n_h}, a relation network gϕr:R2nh→R2nzg_{\phi_r}: \mathbb{R}^{2n_h} \to \mathbb{R}^{2n_z}, and a decoder gϕd:Rnz→R2nxg_{\phi_d}: \mathbb{R}^{n_z} \to \mathbb{R}^{2n_x}.

    The encoding distribution parameters μne,σne∈Rnz\mu_n^e, \sigma_n^e \in \mathbb{R}^{n_z} for class nn are computed by processing all (NK)2(NK)^2 cross-example pairs through the relation network and averaging within each class group:

    μne,σne=1NK2∑kn=1K∑m=1N∑km=1Kgϕr(gϕe(xnkn),gϕe(xmkm))\mu_n^e, \sigma_n^e = \frac{1}{N K^2} \sum_{k_n=1}^K \sum_{m=1}^N \sum_{k_m=1}^K g_{\phi_r}\left(g_{\phi_e}(x_n^{k_n}), g_{\phi_e}(x_m^{k_m})\right)

    zn∼q(zn∣Dntr)=N(μne,diag⁡((σne)2))z_n \sim q(z_n \mid \mathcal{D}_n^{tr}) = \mathcal{N}\left(\mu_n^e, \operatorname{diag}\left((\sigma_n^e)^2\right)\right)

    The full task latent representation is the concatenation z=[z1,z2,…,zN]z = [z_1, z_2, \dots, z_N].

    The parameters wnw_n of the NN-way linear softmax classifier fθ(x)=softmax⁡([w1⋅x,…,wN⋅x]T)f_\theta(x) = \operatorname{softmax}([w_1 \cdot x, \dots, w_N \cdot x]^T) are independently generated from each class latent code znz_n via the decoder gϕdg_{\phi_d}:

    μnd,σnd=gϕd(zn)\mu_n^d, \sigma_n^d = g_{\phi_d}(z_n)

    wn∼p(wn∣zn)=N(μnd,diag⁡((σnd)2))w_n \sim p(w_n \mid z_n) = \mathcal{N}\left(\mu_n^d, \operatorname{diag}\left((\sigma_n^d)^2\right)\right)

    where θ={w1,…,wN}\theta = \{w_1, \dots, w_N\}, nxn_x is the input feature dimension, nhn_h is the intermediate hidden code dimension, and nzn_z is the latent bottleneck dimension (nz≪nxn_z \ll n_x).

  3. Knowl 3 — LEO Meta-Learning Objective and Latent Regularization

    equation

    The parameters ϕ={ϕe,ϕr,ϕd,α}\phi = \{\phi_e, \phi_r, \phi_d, \alpha\} of LEO (encoder, relation network, decoder, and inner learning rate) are meta-trained over task instances Ti=(Ditr,Dival)∼p(T)\mathcal{T}_i = (\mathcal{D}_i^{tr}, \mathcal{D}_i^{val}) \sim p(\mathcal{T}) by minimizing the outer-loop objective:

    min⁡ϕe,ϕr,ϕd,α∑Ti∼p(T)[LTival(fθi′)+βDKL(q(zn∣Dntr) ∥ p(zn))+γ∥stopgrad⁡(zn′)−zn∥22]+R\min_{\phi_e, \phi_r, \phi_d, \alpha} \sum_{\mathcal{T}_i \sim p(\mathcal{T})} \left[ \mathcal{L}_{\mathcal{T}_i}^{val}(f_{\theta'_i}) + \beta D_{KL}\left(q(z_n \mid \mathcal{D}_n^{tr}) \,\|\, p(z_n)\right) + \gamma \left\|\operatorname{stopgrad}(z'_n) - z_n\right\|_2^2 \right] + R

    where:

    • LTival(fθi′)=∑(x,y)∈Dival[−wy′⋅x+log⁡∑j=1Newj′⋅x]\mathcal{L}_{\mathcal{T}_i}^{val}(f_{\theta'_i}) = \sum_{(x,y) \in \mathcal{D}_i^{val}} \left[ - w'_y \cdot x + \log \sum_{j=1}^N e^{w'_j \cdot x} \right] is the cross-entropy loss of the adapted classifier parameters θi′={w1′,…,wN′}\theta'_i = \{w'_1, \dots, w'_N\} on validation set Dival\mathcal{D}_i^{val}.
    • p(zn)=N(0,Inz)p(z_n) = \mathcal{N}(0, I_{n_z}) is a standard isotropic Gaussian prior on class latent codes, weighted by β>0\beta > 0 to encourage disentangled representations.
    • γ>0\gamma > 0 weights a consistency penalty encouraging the initial encoder output znz_n to predict an initialization close to the adapted code zn′z'_n, where stopgrad⁡(⋅)\operatorname{stopgrad}(\cdot) prevents gradients from updating zn′z'_n through this term.
    • RR is a regularization term combining parameter L2L_2 weight decay and a soft layer-wise orthogonality penalty on the decoder weights ϕd\phi_d:

    R=λ1(∥ϕe∥22+∥ϕr∥22+∥ϕd∥22)+λ2∥Cd−I∥2R = \lambda_1 \left( \|\phi_e\|_2^2 + \|\phi_r\|_2^2 + \|\phi_d\|_2^2 \right) + \lambda_2 \|C_d - I\|_2

    where CdC_d is the correlation matrix between the rows of the decoder weight matrix ϕd\phi_d, II is the identity matrix, and λ1,λ2≥0\lambda_1, \lambda_2 \ge 0 are regularization hyperparameters.

  4. Knowl 4 — Latent Embedding Optimization Meta-Training Algorithm

    algorithm

    The LEO meta-training procedure executes nested inner and outer optimization loops over batches of few-shot learning tasks.

    Input: Task distribution p(T)p(\mathcal{T}), outer learning rate η\eta, initial inner learning rates, number of inner adaptation steps NadaptN_{\text{adapt}}
    Output: Meta-learned parameters ϕ={ϕe,ϕr,ϕd,α}\phi = \{\phi_e, \phi_r, \phi_d, \alpha\}
    Randomly initialize ϕe,ϕr,ϕd\phi_e, \phi_r, \phi_d
    Initialize ϕ={ϕe,ϕr,ϕd,α}\phi = \{\phi_e, \phi_r, \phi_d, \alpha\}
    while not converged do
        for each task instance Ti=(Ditr,Dival)\mathcal{T}_i = (\mathcal{D}_i^{tr}, \mathcal{D}_i^{val}) in batch sampled from p(T)p(\mathcal{T}) do
            Encode Ditr\mathcal{D}_i^{tr} to class latent distribution (μne,σne)(\mu_n^e, \sigma_n^e) using gϕeg_{\phi_e} and gϕrg_{\phi_r}
            Sample initial latent codes zn∼N(μne,diag⁡((σne)2))z_n \sim \mathcal{N}(\mu_n^e, \operatorname{diag}((\sigma_n^e)^2))
            Decode z=[z1,…,zN]z = [z_1, \dots, z_N] to initial parameter distribution (μnd,σnd)(\mu_n^d, \sigma_n^d) using gϕdg_{\phi_d}
            Sample initial classifier parameters wn∼N(μnd,diag⁡((σnd)2))w_n \sim \mathcal{N}(\mu_n^d, \operatorname{diag}((\sigma_n^d)^2)), forming θi\theta_i
            Initialize z′←z,θi′←θiz' \leftarrow z, \theta'_i \leftarrow \theta_i
            for step = 1 to NadaptN_{\text{adapt}} do
                Compute training cross-entropy loss LTitr(fθi′)\mathcal{L}_{\mathcal{T}_i}^{tr}(f_{\theta'_i}) on Ditr\mathcal{D}_i^{tr}
                Compute gradient step in latent space: z′←z′−α∇z′LTitr(fθi′)z' \leftarrow z' - \alpha \nabla_{z'} \mathcal{L}_{\mathcal{T}_i}^{tr}(f_{\theta'_i})
                Decode z′z' to updated classifier parameters θi′\theta'_i via gϕdg_{\phi_d}
            end for
            (Optional) Perform fine-tuning gradient steps on θi′\theta'_i in parameter space using Ditr\mathcal{D}_i^{tr}
            Compute validation loss LTival(fθi′)\mathcal{L}_{\mathcal{T}_i}^{val}(f_{\theta'_i}) on Dival\mathcal{D}_i^{val}
        end for
        Compute total meta-loss Lmeta=∑Ti[LTival(fθi′)+βDKL(q(zn∣Dntr)∥p(zn))+γ∥stopgrad⁡(zn′)−zn∥22]+R\mathcal{L}_{\text{meta}} = \sum_{\mathcal{T}_i} [\mathcal{L}_{\mathcal{T}_i}^{val}(f_{\theta'_i}) + \beta D_{KL}(q(z_n \mid \mathcal{D}_n^{tr}) \| p(z_n)) + \gamma \|\operatorname{stopgrad}(z'_n) - z_n\|_2^2] + R
        Update meta-parameters: ϕ←ϕ−η∇ϕLmeta\phi \leftarrow \phi - \eta \nabla_{\phi} \mathcal{L}_{\text{meta}}
    end while

    In standard implementations, inner adaptation uses Nadapt=5N_{\text{adapt}} = 5 steps in latent space followed by 5 optional fine-tuning steps directly in parameter space. Outer optimization utilizes Adam with meta-gradient clipping (gradient values and norms capped at 0.10.1).

  5. Knowl 5 — Curvature and Parameter Space Expansion in Latent Adaptation

    empirical result

    Analysis of the optimization geometry in LEO reveals the underlying mechanisms that enable rapid and expressive adaptation in low-dimensional latent space:

    1. Higher Curvature in Latent Space: The curvature of the loss surface with respect to latent codes zz (measured by the absolute eigenvalues of the Hessian ∇z2L\nabla_z^2 \mathcal{L}) is approximately two orders of magnitude larger (10−510^{-5} to 10−710^{-7}) than the curvature with respect to raw classifier parameters θ\theta (10−710^{-7} to 10−1010^{-10}). Consequently, small steps in latent space induce much more drastic functional transformations than equivalent steps in parameter space.
    2. Parameter Space Expansion by Decoder: Singular value decomposition of the linear decoder mapping gϕd:Z→Θg_{\phi_d}: \mathcal{Z} \to \Theta demonstrates that singular values exceed 1.01.0, expanding vectors by at least one order of magnitude across principal projection directions.
    3. Longer Trajectory Travel per Gradient Step: Gradient steps taken in zz during inner-loop optimization induce displacement norms in parameter space θ\theta that are substantially larger than those produced by direct parameter-space gradient descent methods like Meta-SGD. This enables LEO to transport model parameters over greater distances across task landscapes within only 5 adaptation steps.
  6. Knowl 6 — Few-Shot Image Classification Benchmark Performance

    data/table

    LEO was evaluated on standard 5-way 1-shot and 5-shot classification benchmarks on miniImageNet (100 classes) and tieredImageNet (608 classes across 34 high-level hierarchy nodes). Accuracies are reported as mean ±\pm standard deviation over 5 runs with 50,000 task evaluations per run.

    Model miniImageNet (1-shot) miniImageNet (5-shot)
    Matching Networks 43.56±0.84%43.56 \pm 0.84\% 55.31±0.73%55.31 \pm 0.73\%
    Meta-Learner LSTM 43.44±0.77%43.44 \pm 0.77\% 60.60±0.71%60.60 \pm 0.71\%
    MAML 48.70±1.84%48.70 \pm 1.84\% 63.11±0.92%63.11 \pm 0.92\%
    LLAMA 49.40±1.83%49.40 \pm 1.83\% –
    REPTILE 49.97±0.32%49.97 \pm 0.32\% 65.99±0.58%65.99 \pm 0.58\%
    PLATIPUS 50.13±1.86%50.13 \pm 1.86\% –
    Meta-SGD (WRN-28-10 features) 54.24±0.03%54.24 \pm 0.03\% 70.86±0.04%70.86 \pm 0.04\%
    SNAIL 55.71±0.99%55.71 \pm 0.99\% 68.88±0.92%68.88 \pm 0.92\%
    Gidaris Komodakis (2018) 56.20±0.86%56.20 \pm 0.86\% 73.00±0.64%73.00 \pm 0.64\%
    Bauer et al. (2017) 56.30±0.40%56.30 \pm 0.40\% 73.90±0.30%73.90 \pm 0.30\%
    Munkhdalai et al. (2017) 57.10±0.70%57.10 \pm 0.70\% 70.04±0.63%70.04 \pm 0.63\%
    DEML + Meta-SGD 58.49±0.91%58.49 \pm 0.91\% 71.28±0.69%71.28 \pm 0.69\%
    TADAM 58.50±0.30%58.50 \pm 0.30\% 76.70±0.30%76.70 \pm 0.30\%
    Qiao et al. (2017) 59.60±0.41%59.60 \pm 0.41\% 73.74±0.19%73.74 \pm 0.19\%
    LEO (ours) 61.76±0.08%\mathbf{61.76 \pm 0.08\%} 77.59±0.12%\mathbf{77.59 \pm 0.12\%}
    Model tieredImageNet (1-shot) tieredImageNet (5-shot)
    MAML (deeper network) 51.67±1.81%51.67 \pm 1.81\% 70.30±0.08%70.30 \pm 0.08\%
    Prototypical Networks 53.31±0.89%53.31 \pm 0.89\% 72.69±0.74%72.69 \pm 0.74\%
    Relation Network 54.48±0.93%54.48 \pm 0.93\% 71.32±0.78%71.32 \pm 0.78\%
    Transductive Propagation Nets 57.41±0.94%57.41 \pm 0.94\% 71.55±0.74%71.55 \pm 0.74\%
    Meta-SGD (WRN-28-10 features) 62.95±0.03%62.95 \pm 0.03\% 79.34±0.06%79.34 \pm 0.06\%
    LEO (ours) 66.33±0.05%\mathbf{66.33 \pm 0.05\%} 81.44±0.09%\mathbf{81.44 \pm 0.09\%}

    LEO established state-of-the-art performance across all four benchmark settings. When evaluated on the multi-view feature embedding representation of Qiao et al. (2017), LEO achieved 63.97±0.20%63.97 \pm 0.20\% (1-shot) and 79.49±0.70%79.49 \pm 0.70\% (5-shot) on miniImageNet.

  7. Knowl 7 — Ablation Analysis of LEO Architecture Components

    data/table

    An ablation study evaluated the relative contributions of data-dependent conditioning, latent adaptation, stochasticity, and parameter-space fine-tuning using identical pre-trained WRN-28-10 features across all configurations.

    Model Variant miniImageNet tieredImageNet
    1-shot 5-shot 1-shot 5-shot
    Meta-SGD (direct θ\theta adaptation) 54.24±0.03%54.24 \pm 0.03\% 70.86±0.04%70.86 \pm 0.04\% 62.95±0.03%62.95 \pm 0.03\% 79.34±0.06%79.34 \pm 0.06\%
    Conditional generator only 60.33±0.11%60.33 \pm 0.11\% 74.53±0.11%74.53 \pm 0.11\% 65.17±0.15%65.17 \pm 0.15\% 78.77±0.03%78.77 \pm 0.03\%
    Conditional generator + fine-tuning 60.62±0.31%60.62 \pm 0.31\% 76.42±0.09%76.42 \pm 0.09\% 65.74±0.28%65.74 \pm 0.28\% 80.65±0.07%80.65 \pm 0.07\%
    Previous State-of-the-Art 59.60±0.41%59.60 \pm 0.41\% 76.70±0.30%76.70 \pm 0.30\% 57.41±0.94%57.41 \pm 0.94\% 72.69±0.74%72.69 \pm 0.74\%
    LEO (random prior p(zn)p(z_n) init) 61.01±0.12%61.01 \pm 0.12\% 77.27±0.05%77.27 \pm 0.05\% 65.39±0.10%65.39 \pm 0.10\% 80.83±0.13%80.83 \pm 0.13\%
    LEO (deterministic) 61.48±0.05%61.48 \pm 0.05\% 76.53±0.24%76.53 \pm 0.24\% 66.18±0.17%66.18 \pm 0.17\% 82.06±0.08%82.06 \pm 0.08\%
    LEO (no fine-tuning) 61.62±0.15%61.62 \pm 0.15\% 77.46±0.12%77.46 \pm 0.12\% 66.14±0.17%66.14 \pm 0.17\% 80.89±0.11%80.89 \pm 0.11\%
    LEO (full model) 61.76±0.08%61.76 \pm 0.08\% 77.59±0.12%77.59 \pm 0.12\% 66.33±0.05%66.33 \pm 0.05\% 81.44±0.09%81.44 \pm 0.09\%

    Key takeaways:

    1. Bottleneck effect: The largest performance gap is between Meta-SGD (which adapts directly in parameter space Θ\Theta) and all latent embedding variants, confirming the essential value of the low-dimensional information bottleneck.
    2. Latent adaptation vs feedforward generation: The conditional generator without latent adaptation underperforms full LEO by 1.4–3.1%1.4\text{--}3.1\%; parameter fine-tuning alone cannot recover this gap.
    3. Data-dependent initialization: Replacing the relational encoder with a random sample from prior p(zn)p(z_n) degrades performance, showing the importance of contextual parameter initialization.
    4. Stochasticity and Fine-tuning: Stochasticity benefits miniImageNet more than the larger tieredImageNet. Parameter fine-tuning provides only minor gains (statistically significant primarily on 5-shot tieredImageNet).
  8. Knowl 8 — Visual Feature Pre-Training and Model Architecture Configuration

    experimental setup

    For few-shot classification on miniImageNet and tieredImageNet, visual feature representations and LEO parameter generators are configured as follows:

    1. Feature Extractor Pre-Training: A 28-layer Wide Residual Network (WRN-28-10) with dropout (pkeep=0.5p_{\text{keep}} = 0.5) is trained with supervised classification on 80×8080 \times 80 images strictly from the training meta-set (64 classes for miniImageNet, 351 classes for tieredImageNet). Features of dimension nx=640n_x = 640 are extracted from intermediate layer 21 with spatial average pooling (prior to final distribution-specific layers).
    2. LEO Generator Architecture:
      • Encoder (gϕeg_{\phi_e}): A linear mapping without biases from nx=640n_x = 640 to nh=64n_h = 64.
      • Relation Network (gϕrg_{\phi_r}): A 3-layer MLP with 128 units per layer, ReLU activations, and no biases. It processes concatenated pairs of (NK)2(NK)^2 intermediate codes and outputs vectors of dimension 2nz=1282 n_z = 128 representing the class-conditional Gaussian mean and variance parameters for latent bottleneck dimension nz=64n_z = 64.
      • Decoder (gϕdg_{\phi_d}): A linear mapping without biases from nz=64n_z = 64 to 2×6402 \times 640, parameterizing the mean and diagonal variance of the class linear classifier weights wn∈R640w_n \in \mathbb{R}^{640}.
    3. Optimization Setup: Inner loop uses 5 latent steps (initial α=1.0\alpha = 1.0) and 5 fine-tuning steps (initial learning rate 0.0010.001). Outer meta-training uses Adam with meta-batches of 12 task instances, running up to 100,000 steps with early stopping based on meta-validation accuracy.
  9. Knowl 9 — Multimodal Parameter Distribution Modeling in Ambiguous Few-Shot Regression

    empirical result

    In 1D few-shot regression under uncertainty, LEO was evaluated on a synthetic multimodal task distribution where tasks are sampled equally from sinusoids (y=Asin⁡(x−ϕ)y = A \sin(x - \phi), with A∈[0.1,5],ϕ∈[0,π]A \in [0.1, 5], \phi \in [0, \pi]) and lines (y=ax+by = ax + b, with a,b∈[−3,3]a, b \in [-3, 3]), with inputs x∈[−5,5]x \in [-5, 5] and Gaussian label noise with standard deviation σ=0.3\sigma = 0.3.

    Given 5 training points (K=5K=5), the underlying model fθf_\theta is a 3-layer MLP with 40 units per layer whose entire parameter tensor θ\theta (1761 parameters) is generated from a single nz=16n_z = 16 latent vector zz.

    The results demonstrate:

    1. Uncertainty Capture in Data-Sparse Regions: Sampling multiple latent vectors z∼q(z∣Dtr)z \sim q(z \mid \mathcal{D}^{tr}) produces tight function fits near observed data points, while variance expands across unobserved regions of the input space.
    2. Multimodal Hypothesis Generation: In ambiguous problem instances where the 5 training points can be equally well fit by a line or a sine wave, LEO samples distinct functional solutions belonging to both separate modal families (lines and sinusoids) from the same conditioning data.

Coverage note — None was omitted.

References

  1. 1.Marcin Andrychowicz, Misha Denil, Sergio Gomez, Matthew W Hoffman, David Pfau, Tom Schaul, Brendan Shillingford, and Nando De Freitas. Learning to learn by gradient descent by gradient descent. In Advances in Neural Information Processing Systems, pp. 3981–3989, 2016.
  2. 2.Jimmy Ba, Geoffrey E Hinton, Volodymyr Mnih, Joel Z Leibo, and Catalin Ionescu. Using fast weights to attend to the recent past. In Advances in Neural Information Processing Systems, pp. 4331–4339, 2016.
  3. 3.M. Bauer, M. Rojas-Carulla, J. Bartłomiej Świątkowski, B. Schölkopf, and R. E. Turner. Discriminative k-shot learning using probabilistic models. ArXiv e-prints, June 2017.
  4. 4.Kyunghyun Cho, Bart van Merrienboer, Çaglar Gölçehre, Fethi Bougares, Holger Schwenk, and Yoshua Bengio. Learning phrase representations using RNN encoder-decoder for statistical machine translation. CoRR, abs/1406.1078, 2014. URL http://arxiv.org/abs/1406.1078.
  5. 5.J. Deng, W. Dong, R. Socher, L. Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In 2009 IEEE Conference on Computer Vision and Pattern Recognition, pp. 248–255, June 2009a.
  6. 6.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In Computer Vision and Pattern Recognition, 2009. CVPR 2009. IEEE Conference on, pp. 248–255. IEEE, 2009b.
  7. 7.Li Fei-Fei, Rob Fergus, and Pietro Perona. One-shot learning of object categories. IEEE transactions on pattern analysis and machine intelligence, 28(4):594–611, 2006.
  8. 8.Chelsea Finn, Pieter Abbeel, and Sergey Levine. Model-agnostic meta-learning for fast adaptation of deep networks. In International Conference on Machine Learning, pp. 1126–1135, 2017.
  9. 9.Chelsea Finn, Kelvin Xu, and Sergey Levine. Probabilistic model-agnostic meta-learning. arXiv preprint arXiv:1806.02817, 2018.
  10. 10.Marta Garnelo, Dan Rosenbaum, Christopher Maddison, Tiago Ramalho, David Saxton, Murray Shanahan, Yee Whye Teh, Danilo Rezende, and S. M. Ali Eslami. Conditional neural processes. In Jennifer Dy and Andreas Krause (eds.), Proceedings of the 35th International Conference on Machine Learning, volume 80 of Proceedings of Machine Learning Research, pp. 1704–1713, Stockholmsmssan, Stockholm Sweden, 10–15 Jul 2018a. PMLR. URL http://proceedings.mlr.press/v80/garnelo18a.html.
  11. 11.Marta Garnelo, J. Schwarz, D. Rosenbaum, F. Viola, D. J. Rezende, S. M. A. Eslami, and Y. Whye Teh. Neural Processes. ArXiv e-prints, July 2018b.
  12. 12.Spyros Gidaris and Nikos Komodakis. Dynamic few-shot visual learning without forgetting. CoRR, abs/1804.09458, 2018. URL http://arxiv.org/abs/1804.09458.
  13. 13.Erin Grant, Chelsea Finn, Sergey Levine, Trevor Darrell, and Thomas Griffiths. Recasting gradient-based meta-learning as hierarchical bayes. In Proceedings of the 6th International Conference on Learning Representations (ICLR), 2018.
  14. 14.David Ha, Andrew Dai, and Quoc V Le. Hypernetworks. arXiv preprint arXiv:1609.09106, 2016.
  15. 15.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 770–778, 2016.
  16. 16.Irina Higgins, Loic Matthey, Arka Pal, Christopher Burgess, Xavier Glorot, Matthew Botvinick, Shakir Mohamed, and Alexander Lerchner. Beta-VAE: Learning basic visual concepts with a constrained variational framework. ICLR, 2017.
  17. 17.Geoffrey E Hinton and David C Plaut. Using fast weights to deblur old memories. In Proceedings of the ninth annual conference of the Cognitive Science Society, pp. 177–186, 1987.
  18. 18.Sepp Hochreiter, A. Steven Younger, and Peter R. Conwell. Learning to learn using gradient descent. In Proceedings of the International Conference on Artificial Neural Networks, ICANN ’01, pp. 87–94, London, UK, UK, 2001. Springer-Verlag. ISBN 3-540-42486-5. URL http://dl.acm.org/citation.cfm?id=646258.684281.
  19. 19.T. Kim, J. Yoon, O. Dia, S. Kim, Y. Bengio, and S. Ahn. Bayesian Model-Agnostic Meta-Learning. ArXiv e-prints, June 2018.
  20. 20.Diederik P. Kingma and Jimmy Ba. Adam: A method for stochastic optimization. CoRR, abs/1412.6980, 2014. URL http://arxiv.org/abs/1412.6980.
  21. 21.Gregory Koch, Richard Zemel, and Ruslan Salakhutdinov. Siamese neural networks for one-shot image recognition. In ICML Deep Learning Workshop, volume 2, 2015.
  22. 22.D. Krueger, C.-W. Huang, R. Islam, R. Turner, A. Lacoste, and A. Courville. Bayesian Hypernetworks. ArXiv e-prints, October 2017.
  23. 23.A. Lacoste, B. Oreshkin, W. Chung, T. Boquet, N. Rostamzadeh, and D. Krueger. Uncertainty in Multitask Transfer Learning. ArXiv e-prints, June 2018.
  24. 24.Brenden Lake, Ruslan Salakhutdinov, Jason Gross, and Joshua Tenenbaum. One shot learning of simple visual concepts. In Proceedings of the Annual Meeting of the Cognitive Science Society, volume 33, 2011.
  25. 25.Yann LeCun, Yoshua Bengio, and Geoffrey Hinton. Deep learning. nature, 521(7553):436, 2015.
  26. 26.Y. Lee and S. Choi. Gradient-Based Meta-Learning with Learned Layerwise Metric and Subspace. ArXiv e-prints, January 2018.
  27. 27.Zhenguo Li, Fengwei Zhou, Fei Chen, and Hang Li. Meta-sgd: Learning to learn quickly for few shot learning. CoRR, abs/1707.09835, 2017. URL http://arxiv.org/abs/1707.09835.
  28. 28.Yanbin Liu, Juho Lee, Minseop Park, Saehoon Kim, and Yi Yang. Transductive propagation network for few-shot learning. arXiv preprint arXiv:1805.10002, 2018.
  29. 29.Nikhil Mishra, Mostafa Rohaninejad, Xi Chen, and Pieter Abbeel. A simple neural attentive meta-learner. In Proceedings of the 6th International Conference on Learning Representations (ICLR), 2018.
  30. 30.Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A. Rusu, Joel Veness, Marc G. Bellemare, Alex Graves, Martin Riedmiller, Andreas K. Fidjeland, Georg Ostrovski, Stig Petersen, Charles Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Shane Legg, and Demis Hassabis. Human-level control through deep reinforcement learning. Nature, 518(7540):529–533, February 2015. ISSN 00280836. URL http://dx.doi.org/10.1038/nature14236.
  31. 31.Tsendsuren Munkhdalai, Xingdi Yuan, Soroush Mehri, Tong Wang, and Adam Trischler. Learning rapid-temporal adaptations. CoRR, abs/1712.09926, 2017. URL http://arxiv.org/abs/1712.09926.
  32. 32.Alex Nichol and John Schulman. Reptile: a scalable metalearning algorithm. arXiv preprint arXiv:1803.02999, 2018.
  33. 33.B. N. Oreshkin, P. Rodriguez, and A. Lacoste. TADAM: Task dependent adaptive metric for improved few-shot learning. ArXiv e-prints, May 2018.
  34. 34.Siyuan Qiao, Chenxi Liu, Wei Shen, and Alan Yuille. Few-shot image recognition by predicting parameters from activations. arXiv preprint arXiv:1706.03466, 2017.
  35. 35.Sachin Ravi and Hugo Larochelle. Optimization as a model for few-shot learning. In Proceedings of the 5th International Conference on Learning Representations (ICLR), 2017.
  36. 36.Mengye Ren, Sachin Ravi, Eleni Triantafillou, Jake Snell, Kevin Swersky, Josh B. Tenenbaum, Hugo Larochelle, and Richard S. Zemel. Meta-learning for semi-supervised few-shot classification. In International Conference on Learning Representations, 2018. URL https://openreview.net/forum?id=HJcSzz-CZ.
  37. 37.Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael S. Bernstein, Alexander C. Berg, and Fei-Fei Li. Imagenet large scale visual recognition challenge. CoRR, abs/1409.0575, 2014. URL http://arxiv.org/abs/1409.0575.
  38. 38.Adam Santoro, Sergey Bartunov, Matthew Botvinick, Daan Wierstra, and Timothy Lillicrap. Meta-learning with memory-augmented neural networks. In International conference on machine learning, pp. 1842–1850, 2016.
  39. 39.Jürgen Schmidhuber. Evolutionary principles in self-referential learning, or on learning how to learn: the meta-meta-... hook. PhD thesis, Technische Universität München, 1987.
  40. 40.Jürgen Schmidhuber. Deep learning in neural networks: An overview. Neural networks, 61:85–117, 2015.
  41. 41.David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, Yutian Chen, Timothy Lillicrap, Fan Hui, Laurent Sifre, George van den Driessche, Thore Graepel, and Demis Hassabis. Mastering the game of go without human knowledge. Nature, 550:354–, October 2017. URL http://dx.doi.org/10.1038/nature24270.
  42. 42.Karen Simonyan and Andrew Zisserman. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556, 2014.
  43. 43.Jake Snell, Kevin Swersky, and Richard S. Zemel. Prototypical networks for few-shot learning. CoRR, abs/1703.05175, 2017. URL http://arxiv.org/abs/1703.05175.
  44. 44.Flood Sung, Yongxin Yang, Li Zhang, Tao Xiang, Philip H. S. Torr, and Timothy M. Hospedales. Learning to compare: Relation network for few-shot learning. CoRR, abs/1711.06025, 2017. URL http://arxiv.org/abs/1711.06025.
  45. 45.Ilya Sutskever, Oriol Vinyals, and Quoc V. Le. Sequence to sequence learning with neural networks. CoRR, abs/1409.3215, 2014. URL http://arxiv.org/abs/1409.3215.
  46. 46.Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich. Going deeper with convolutions. In 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR).
  47. 47.Sebastian Thrun and Lorien Pratt. Learning to learn: Introduction and overview. In Learning to learn, pp. 3–17. Springer, 1998.
  48. 48.Oriol Vinyals, Charles Blundell, Tim Lillicrap, Daan Wierstra, et al. Matching networks for one shot learning. In Advances in Neural Information Processing Systems, pp. 3630–3638, 2016.
  49. 49.T. Wu, J. Peurifoy, I. L. Chuang, and M. Tegmark. Meta-learning autoencoders for few-shot prediction. ArXiv e-prints, July 2018.
  50. 50.Jason Yosinski, Jeff Clune, Yoshua Bengio, and Hod Lipson. How transferable are features in deep neural networks? CoRR, abs/1411.1792, 2014. URL http://arxiv.org/abs/1411.1792.
  51. 51.Sergey Zagoruyko and Nikos Komodakis. Wide residual networks. In British Machine Vision Conference, 2016a.
  52. 52.Sergey Zagoruyko and Nikos Komodakis. Wide residual networks. CoRR, abs/1605.07146, 2016b. URL http://arxiv.org/abs/1605.07146.
  53. 53.Fengwei Zhou, Bin Wu, and Zhenguo Li. Deep meta-learning: Learning to learn in the concept space. CoRR, abs/1802.03596, 2018. URL http://arxiv.org/abs/1802.03596.

Citation

MLA
Rusu, A. A., et al. “Meta-Learning with Latent Embedding Optimization”. arXiv, 2018, http://arxiv.org/abs/1807.05960v3.
APA
Rusu, A. A., Rao, D., Sygnowski, J., Vinyals, O., Pascanu, R., Osindero, S., & Hadsell, R. (2018). Meta-Learning with Latent Embedding Optimization. arXiv. http://arxiv.org/abs/1807.05960v3
Chicago
Rusu, A. A., D. Rao, J. Sygnowski, et al. 2018. “Meta-Learning with Latent Embedding Optimization”. arXiv. http://arxiv.org/abs/1807.05960v3.
Harvard
Rusu, A.A. et al. (2018) “Meta-Learning with Latent Embedding Optimization”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1807.05960v3.
Vancouver
1. Rusu AA, Rao D, Sygnowski J, Vinyals O, Pascanu R, Osindero S, Hadsell R (2018) Meta-Learning with Latent Embedding Optimization. arXiv

BibTeX

@article{rusu2018meta,
  title = {Meta-Learning with Latent Embedding Optimization},
  author = {Rusu, Andrei A. and Rao, Dushyant and Sygnowski, Jakub and Vinyals, Oriol and Pascanu, Razvan and Osindero, Simon and Hadsell, Raia},
  year = {2018},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1807.05960v3},
  eprint = {1807.05960}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Published with permission