A Simple Neural Attentive Meta-Learner

Nikhil MishraMostafa RohaninejadXi ChenPieter Abbeel

article2017ICLR1,404 citations

Introduces SNAIL, a versatile meta-learning architecture combining temporal convolutions and soft attention that achieves state-of-the-art performance across both supervised and reinforcement learning benchmarks without relying on hand-engineered algorithmic constraints.

Listen

Modern artificial intelligence models perform well with massive datasets but struggle when data is scarce or when tasks change rapidly. To address this, meta-learning trains models across a distribution of related tasks so they can learn overarching problem-solving strategies. However, existing approaches often rely on hand-crafted, domain-specific designs or enforce rigid optimization mechanics, such as gradient descent at test time, which constrain adaptability and require significant manual engineering.

The article introduces and evaluates the Simple Neural Attentive Learner (SNAIL), a versatile, general-purpose meta-learning architecture. The objective is to demonstrate that a single unified model can discover optimal strategies across diverse domains without relying on hand-designed components or hard-coded algorithmic priors.

The researchers designed an architecture combining two sequence-processing mechanisms: temporal convolutions to aggregate local contextual history and causal attention to precisely retrieve specific past experiences from arbitrary distances. The architecture was tested across several standard benchmarks spanning supervised learning—such as few-shot character and image classification on Omniglot and mini-ImageNet—and reinforcement learning, including multi-armed bandits, simulated decision-making environments, robotic locomotion, and 3D visual maze navigation.

The evaluation produced four primary findings. First, the proposed model attained state-of-the-art accuracy in supervised few-shot image classification, scoring 55.71% in 1-shot and 68.88% in 5-shot mini-ImageNet, significantly surpassing established baselines. Second, in reinforcement learning, the model scaled effectively to long sequences and complex visual inputs, solving 3D mazes in substantially fewer steps than recurrent networks (averaging 105.9 steps on large mazes compared to 150.6 steps for the baseline). Third, in robotic continuous control tasks, the model adapted within just a few timesteps of a single episode, exploiting task structures where gradient-adaptation baselines required dozens of rollouts. Fourth, ablation studies confirmed that temporal convolutions and attention mechanisms are mutually necessary; removing temporal convolutions caused models to fail in sequential environments, while removing attention degraded long-term memory capacity.

These findings demonstrate that organizations do not need custom, heavily-engineered architectures for different learning paradigms. A single generic sequence-to-sequence model can independently discover effective algorithms directly from data. This structural flexibility can reduce development timelines, lower model maintenance overhead, and enhance autonomous adaptation in dynamic real-world environments such as robotics and recommendation engines.

Organizations developing adaptive systems should consider adopting hybrid architectures combining temporal convolutions and attention instead of relying on purely recurrent networks or rigid gradient-based meta-learners. Further research is recommended to expand sequence-to-sequence meta-learners into general sequential problems like natural language translation and to develop lifelong memory mechanisms that retain experience across an agent's entire deployment lifecycle.

While the empirical results show high confidence and statistical significance across benchmarks, the experiments were conducted in controlled simulated environments and curated vision datasets. Leaders should recognize that real-world deployment on unconstrained, open-ended tasks may introduce unexpected distribution shifts or computational overhead, warranting targeted pilot evaluations before large-scale implementation.

arXiv: 1707.03141
  • Paper: Meta-Learning with Memory-Augmented Neural Networks, Adam Santoro et al. (2016). Introduces memory-augmented neural networks for rapid few-shot learning, providing key conceptual groundwork for recurrent memory-based meta-learners that SNAIL builds upon and enhances.
  • Paper: Matching Networks for One Shot Learning, Oriol Vinyals et al. (2016). Pioneers the use of attention mechanisms over support sets in episodic meta-learning, directly motivating SNAIL's integration of soft attention for pinpointing past experiences.
  • Paper: Model-Agnostic Meta-Learning for Fast Adaptation of Deep Networks, Chelsea Finn et al. (2017). Establishes gradient-based Model-Agnostic Meta-Learning (MAML), serving as a core baseline and foundational formulation of few-shot supervised and reinforcement learning tasks that SNAIL benchmarks against.
  • Paper: Optimization as a Model for Few-Shot Learning, Sachin Ravi et al. (2017). Frames few-shot meta-learning through recurrent neural network states, establishing standard Mini-ImageNet evaluation protocols and architectural baselines surpassed by SNAIL.
  • Paper: Learning to learn by gradient descent by gradient descent, Marcin Andrychowicz et al. (2016). Demonstrates training recurrent neural networks to act as general-purpose meta-optimizers, laying foundational ideas for black-box meta-learning models.
Cover for A Simple Neural Attentive Meta-Learner

Abstract

Deep neural networks excel in regimes with large amounts of data, but tend to struggle when data is scarce or when they need to adapt quickly to changes in the task. In response, recent work in meta-learning proposes training a meta-learner on a distribution of similar tasks, in the hopes of generalization to novel but related tasks by learning a high-level strategy that captures the essence of the problem it is asked to solve. However, many recent meta-learning approaches are extensively hand-designed, either using architectures specialized to a particular application, or hard-coding algorithmic components that constrain how the meta-learner solves the task. We propose a class of simple and generic meta-learner architectures that use a novel combination of temporal convolutions and soft attention; the former to aggregate information from past experience and the latter to pinpoint specific pieces of information. In the most extensive set of meta-learning experiments to date, we evaluate the resulting Simple Neural AttentIve Learner (or SNAIL) on several heavily-benchmarked tasks. On all tasks, in both supervised and reinforcement learning, SNAIL attains state-of-the-art performance by significant margins.

Table of Contents

  • 1 INTRODUCTION
  • 2 META-LEARNING PRELIMINARIES
  • 3.1 MODULAR BUILDING BLOCKS
  • 4 RELATED WORK
  • 5 EXPERIMENTS
  • 5.1 FEW-SHOT IMAGE CLASSIFICATION
  • 5.2 REINFORCEMENT LEARNING
  • 5.2.1 MULTI-ARMED BANDITS
  • 5.2.2 TABULAR MDPS
  • 5.2.3 CONTINUOUS CONTROL
  • 5.2.4 VISUAL NAVIGATION
  • 6 CONCLUSION AND FUTURE WORK
  • REFERENCES
  • APPENDIX
  • A FEW-SHOT CLASSIFICATION ARCHITECTURES
  • B FEW-SHOT CLASSIFICATION: ABLATIONS
  • C REINFORCEMENT LEARNING
  • C.1 MULTI-ARMED BANDIT AND TABULAR MDP ARCHITECTURES
  • C.2 CONTINUOUS CONTROL ARCHITECTURES
  • C.3 VISUAL NAVIGATION ARCHITECTURES
  • C.4 ADDITIONAL REINFORCEMENT LEARNING HYPERPARAMETERS
  • D REINFORCEMENT LEARNING: ABLATIONS

Knowls

  1. Knowl 1 — Meta-Learning Formulation as a Sequence-to-Sequence Objective

    equation

    The meta-learning problem is formalized as an episodic sequence-to-sequence learning task over a distribution of related tasks T=P(Ti)\mathcal{T} = P(\mathcal{T}_i). Each task Ti\mathcal{T}_i is defined by a sequence of inputs xtx_t, actions or outputs ata_t, a task loss function Li(xt,at)\mathcal{L}_i(x_t, a_t), an environment transition distribution Pi(xt∣xt−1,at−1)P_i(x_t \mid x_{t-1}, a_{t-1}), and a horizon length HiH_i.

    A meta-learner parameterized by θ\theta models the conditional distribution π(at∣x1,…,xt;θ)\pi(a_t \mid x_1, \dots, x_t; \theta) over the input history. The meta-training objective minimizes expected cumulative loss across the task distribution:

    min⁡θETi∼T[∑t=0HiLi(xt,at)]\min_{\theta} \mathbb{E}_{\mathcal{T}_i \sim \mathcal{T}} \left[ \sum_{t=0}^{H_i} \mathcal{L}_i(x_t, a_t) \right]

    where xt∼Pi(xt∣xt−1,at−1)x_t \sim P_i(x_t \mid x_{t-1}, a_{t-1}) and at∼π(at∣x1,…,xt;θ)a_t \sim \pi(a_t \mid x_1, \dots, x_t; \theta). At test time, the trained parameters θ\theta are evaluated on unseen tasks sampled from a related task distribution T~=P(T~i)\tilde{\mathcal{T}} = P(\tilde{\mathcal{T}}_i).

  2. Knowl 2 — Dense Block and Temporal Convolution (TC) Block in SNAIL

    algorithm

    The Simple Neural Attentive Learner (SNAIL) aggregates contextual information across sequential inputs using causal 1D temporal convolutions (TC). A Dense Block applies two parallel causal 1D convolutions (each with kernel size 2, dilation rate RR, and DD filters) to produce filter activations xfx_f and gate activations xgx_g. It computes gated activations via an element-wise product tanh⁡(xf)⊙sigmoid⁡(xg)\tanh(x_f) \odot \operatorname{sigmoid}(x_g) and concatenates the result with the original inputs along the feature dimension.

    A TC Block stacks multiple dense blocks with exponentially increasing dilation rates R=21,22,…,2⌈log⁡2T⌉R = 2^1, 2^2, \dots, 2^{\lceil \log_2 T \rceil} until the receptive field spans the sequence length TT:

    function DenseBlock(inputs, dilation_rate R, num_filters D):
        xf, xg = CausalConv(inputs, R, D), CausalConv(inputs, R, D)
        activations = tanh(xf) * sigmoid(xg)
        return concat(inputs, activations)
    function TCBlock(inputs, sequence_length T, num_filters D):
        for i in 1, ..., ceil(log2(T)):
            inputs = DenseBlock(inputs, 2^i, D)
        return inputs
  3. Knowl 3 — Causal Attention Block in SNAIL

    algorithm

    SNAIL provides pinpoint access to specific timesteps across an arbitrary context length using a causal key-value attention mechanism. Given an input sequence matrix of shape T×CT \times C (sequence length TT, channel dimension CC), affine transformations project the inputs into query and key matrices of dimension KK, and a value matrix of dimension VV. Logits are computed via scaled dot product, masked causally to prevent attending to future positions, normalized via softmax, and multiplied by values before concatenating with the input:

    function AttentionBlock(inputs, key_size K, value_size V):
        keys = affine(inputs, K)
        query = affine(inputs, K)
        logits = matmul(query, transpose(keys))
        probs = CausallyMaskedSoftmax(logits / sqrt(K))
        values = affine(inputs, V)
        read = matmul(probs, values)
        return concat(inputs, read)

    Here, CausallyMaskedSoftmax⁡(⋅)\operatorname{CausallyMaskedSoftmax}(\cdot) sets all matrix entries (t,t′)(t, t') where t′>tt' > t to −∞-\infty before applying softmax normalization along rows.

  4. Knowl 4 — SNAIL Architecture and Sequence Formulations for SL and RL

    model/method

    SNAIL combines Temporal Convolution (TC) blocks (which aggregate sequential context) with Causal Attention blocks (which pinpoint specific past values) in an interleaved architecture trained end-to-end.

    In supervised few-shot classification (NN-way, KK-shot), the model processes a sequence of length T=NK+1T = NK + 1 consisting of labeled examples (x1,y1),…,(xNK,yNK)(x_1, y_1), \dots, (x_{NK}, y_{NK}) followed by an unlabeled test example (xNK+1,−)(x_{NK+1}, -). An embedding network extracts image feature vectors, which are concatenated with one-hot labels (or zero vectors for unlabeled queries) and passed through: AttentionBlock(64,32)→TCBlock(T,128)→AttentionBlock(256,128)→TCBlock(T,128)→AttentionBlock(512,256)→1×1 Conv(N)\text{AttentionBlock}(64, 32) \to \text{TCBlock}(T, 128) \to \text{AttentionBlock}(256, 128) \to \text{TCBlock}(T, 128) \to \text{AttentionBlock}(512, 256) \to 1\times 1\text{ Conv}(N)

    In meta-reinforcement learning, the model processes observation-action-reward tuples (o1,−,−),(o2,a1,r1),…,(ot,at−1,rt−1)(o_1, -, -), (o_2, a_1, r_1), \dots, (o_t, a_{t-1}, r_{t-1}) alongside an episode-termination indicator flag. Crucially, SNAIL preserves its internal activations across episode boundaries within each task trial, maintaining memory across multi-episode rollouts. Training is conducted using Trust Region Policy Optimization (TRPO) with Generalized Advantage Estimation (GAE).

  5. Knowl 5 — Few-Shot Classification Benchmarks on Omniglot and mini-ImageNet

    data/table

    SNAIL was evaluated on few-shot classification using Omniglot and mini-ImageNet. Models were trained on variable shot sizes K∼U{1,5}K \sim \mathcal{U}\{1, 5\} within each NN-way episode rather than separate models per shot value. Accuracies are reported with 95% confidence intervals:

    Could not parse LaTeX table
    Could not parse LaTeX table

    SNAIL outperformed domain-specific metric learning and gradient-based adaptation baselines across all setups, achieving a gain of +6.50% on 1-shot and +3.11% on 5-shot mini-ImageNet over prior state of the art.

  6. Knowl 6 — Meta-RL Performance on Multi-Armed Bandits and Tabular MDPs

    data/table

    SNAIL was evaluated on Bernoulli multi-armed bandits (KK arms, horizon NN) and randomly generated tabular Markov Decision Processes (10 states, 5 actions, Gaussian rewards, Dirichlet transition priors, evaluated over NN episodes of 10 steps each).

    For bandits, the table below reports mean reward per episode alongside the Bayes-optimal Gittins index:

    Could not parse LaTeX table

    For tabular MDPs, performance is normalized by the value-iteration upper bound:

    Could not parse LaTeX table

    SNAIL matched or surpassed asymptotic human-designed RL strategies (such as PSRL and UCRL2) as well as recurrent meta-learners, scaling efficiently to sequence lengths up to 1,000 steps.

  7. Knowl 7 — Meta-RL in Visual Navigation and Continuous Control Locomotion

    empirical result

    SNAIL was tested on two complex reinforcement learning settings:

    1. Visual Navigation in Mazes: The agent navigates randomly-generated 3D mazes using 30×4030 \times 40 visual image inputs across 2 consecutive episodes per maze. The goal location is fixed across both episodes. In Small Mazes (episode horizon 250), SNAIL achieved an average target-finding step count of 50.3±0.350.3 \pm 0.3 in Episode 1 and 34.8±0.234.8 \pm 0.2 in Episode 2 (compared to LSTM's 52.4±1.3→39.1±0.952.4 \pm 1.3 \to 39.1 \pm 0.9). In Large Mazes (episode horizon 1000), SNAIL achieved 140.5±4.2140.5 \pm 4.2 in Episode 1 and 105.9±2.4105.9 \pm 2.4 in Episode 2 (compared to LSTM's 180.1±6.0→150.6±5.9180.1 \pm 6.0 \to 150.6 \pm 5.9). The policy meta-learned an explicit exploratory strategy on Episode 1 followed by direct navigation to the remembered goal on Episode 2.

    2. Continuous Control Locomotion: Simulated Planar Cheetah and 3D Ant robots were tasked with adapting to unknown target velocities or directions ({cheetah, ant} ×\times {goal velocity, goal direction}). Because SNAIL incorporates past actions and rewards into internal activations, it identified the goal within the initial timesteps of the first episode and immediately tracked the oracle policy, whereas MAML required dozens of rollout episodes and multiple explicit policy gradient update steps to achieve equivalent performance.

  8. Knowl 8 — Ablations on Temporal Convolutions vs. Causal Attention

    empirical result

    Ablation experiments across supervised learning and reinforcement learning demonstrate that temporal convolutions and causal attention are complementary and both necessary for peak performance:

    1. Few-Shot Classification (5-Way mini-ImageNet):

      • Full SNAIL: 55.71% (1-shot), 68.88% (5-shot).
      • SNAIL without Attention (TC layers only): 55.1% (1-shot), 61.2% (5-shot). Dilated convolutions have coarse access to older inputs; without attention, accuracy drops substantially on longer context lengths (5-shot, T=26T=26).
      • SNAIL without TC (Attention layers only): 49.9% (1-shot), 63.9% (5-shot). Removing local temporal aggregation degrades feature extraction quality.
      • Stacked LSTM baseline: Failed to train on mini-ImageNet (and achieved only 78.1% / 90.8% on 5-way Omniglot).
    2. Reinforcement Learning (Tabular MDPs, normalized return):

      • SNAIL without Attention (TC-only WaveNet variant): Achieved only 0.616 (N=10N=10) and saturated at 0.728 (N=100N=100), compared to 0.766 and 0.941 for full SNAIL, due to bounded capacity over multi-episode memory.
      • SNAIL without TC (Transformer attention with sinusoidal positional encodings): Failed completely and produced random-level performance (0.482), because isolated attention lookups cannot directly aggregate contiguous state-action-reward transition tuples.
  9. Knowl 9 — Cross-Dataset Transferability of Meta-Learned SNAIL Parameters

    empirical result

    The algorithm learned by SNAIL is modular and transferable between distinct visual domains without re-training the core meta-learner:

    • Omniglot SNAIL to mini-ImageNet: A SNAIL architecture trained exclusively on 5-way Omniglot had its weights frozen. Training only a newly initialized visual embedding network on mini-ImageNet achieved 50.62% (1-shot) and 62.34% (5-shot) on 5-way mini-ImageNet.
    • mini-ImageNet SNAIL to Omniglot: A SNAIL trained exclusively on mini-ImageNet had its weights frozen. Training only a new visual embedding on Omniglot achieved 98.66% (1-shot) and 99.56% (5-shot) on 5-way Omniglot.
    • Linear Adapter Transfer: Freezing an Omniglot-trained embedding network and a mini-ImageNet-trained SNAIL model, while training only a single intermediate linear projection layer between them, achieved 98.5% (1-shot) and 99.5% (5-shot) on 5-way Omniglot.

Coverage note — No substantial contributed material was omitted; all core model components, mathematical formulations, benchmarks in few-shot classification and reinforcement learning, ablation analyses, and transferability results are fully covered.

References

  1. 1.Marcin Andrychowicz, Misha Denil, Sergio Gomez, Matthew W Hoffman, David Pfau, Tom Schaul, and Nando de Freitas. Learning to learn by gradient descent by gradient descent. In Advances in Neural Information Processing Systems (NIPS), 2016.
  2. 2.Samy Bengio, Yoshua Bengio, Jocelyn Cloutier, and Jan Gecsei. On the optimization of a synaptic learning rule. In Optimality in Artificial and Biological Neural Networks, pp. 6–8. Univ. of Texas, 1992.
  3. 3.Yan Duan, John Schulman, Xi Chen, Peter L. Bartlett, Ilya Sutskever, and Pieter Abbeel. Rl∧2^{\wedge}2: Fast reinforcement learning via slow reinforcement learning. arXiv preprint arXiv:1611.02779, 2016.
  4. 4.Chelsea Finn, Pieter Abbeel, and Sergy Levine. Model-agnostic meta learning. International Conference on Machine Learning (ICML), 2017.
  5. 5.J.C. Gittins. Bandit processes and dynamic allocation indices. Journal of the Royal Statistical Society. Series B (Methodological), 1979.
  6. 6.Alex Graves, Greg Wayne, and Ivo Danihelka. Neural turing machines. arXiv preprint arXiv:1410.5401, 2014.
  7. 7.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Conference on Computer Vision and Pattern Recognition (CVPR), 2016.
  8. 8.Sepp Hochreiter, A Younger, and Peter Conwell. Learning to learn using gradient descent. Artificial Neural Networks, ICANN, 2001.
  9. 9.Gao Huang, Zhuang Liu, Kilian Q Weinberger, and Laurens van der Maaten. Densely connected convolutional networks. arXiv preprint arXiv:1608.06993, 2016.
  10. 10.Sergey Ioffe and Christian Szegedy. Batch normalization: Accelerating deep network training by reducing internal covariate shift. International Conference on Machine Learning (ICML), 2015.
  11. 11.Thomas Jaksch, Ronald Ortner, and Peter Auer. Near-optimal regret bounds for reinforcement learning. Journal of Machine Learning Research, 2010.
  12. 12.Diederik Kingma and Jimmy Ba. Adam: A method for stochastic optimization. In International Conference on Learning Representations (ICLR), 2015.
  13. 13.Gregory Koch. Siamese neural networks for one-shot image recognition. PhD thesis, University of Toronto, 2015.
  14. 14.Brenden M Lake, Ruslan Salakhutdinov, Jason Gross, and Joshua B Tenenbaum. One shot learning of simple visual concepts. In CogSci, 2011.
  15. 15.Ke Li and Jitendra Malik. Learning to optimize. International Conference on Learning Representations (ICLR), 2017.
  16. 16.Tsendsuren Munkhdalai and Hong Yu. Meta networks. International Conference on Machine Learning (ICML), 2017.
  17. 17.Devang K Naik and RJ Mammone. Meta-neural networks that learn by learning. In Neural Networks, 1992. IJCNN., International Joint Conference on, volume 1, pp. 437–442. IEEE, 1992.
  18. 18.Ian Osband and Benjamin Van Roy. Why is posterior sampling better than optimism for reinforcement learning. International Conference on Machine Learning (ICML), 2017.
  19. 19.Sachin Ravi and Hugo Larochelle. Optimization as a model for few-shot learning. In International Conference on Learning Representations (ICLR), 2017.
  20. 20.Adam Santoro, Sergey Bartunov, Matthew Botvinick, Daan Wierstra, and Timothy Lillicrap. Meta-learning with memory-augmented neural networks. In International Conference on Machine Learning (ICML), 2016.
  21. 21.Jurgen Schmidhuber. Evolutionary principles in self-referential learning. On learning how to learn: The meta-meta-... hook.) Diploma thesis, Institut f. Informatik, Tech. Univ. Munich, 1987.
  22. 22.John Schulman, Sergey Levine, Philipp Moritz, Michael Jordan, and Pieter Abbeel. Trust region policy optimization. International Conference on Machine Learning (ICML), 2015.
  23. 23.John Schulman, Philipp Moritz, Sergey Levine, Michael Jordan, and Pieter Abbeel. High-dimensional continuous control using generalized advantage estimation. International Conference on Learning Representations (ICLR), 2016.
  24. 24.Jake Snell, Kevin Swersky, and Richard S Zemel. Prototypical networks for few-shot learning. arXiv preprint arXiv:1703.05175, 2017.
  25. 25.Malcolm Strens. A bayesian framework for reinforcement learning. In International Conference on Machine Learning (ICML), 2000.
  26. 26.Sebastian Thrun and Lorien Pratt. Learning to learn: Introduction and overview. In Learning to learn. Springer, 1998.
  27. 27.Aaron van den Oord, Sander Dieleman, Heig Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew W. Senior, and Koray Kavukcuoglu. Wavenet: A generative model for raw audio. CoRR, abs/1609.03499, 2016a.
  28. 28.Aaron van den Oord, Nal Kalchbrenner, Lasse Espeholt, Oriol Vinyals, Alex Graves, et al. Conditional image generation with pixelcnn decoders. In Advances in Neural Information Processing Systems (NIPS), 2016b.
  29. 29.Ashish Vaswani, Noah Shazeer, Jakob Uszkoreit, Llion Jones, Aidan Gomez N., Lukas Kaiser, and Illia Polosukhin. Attention is all you need. arXiv preprint arXiv:1706.03762, 2017a.
  30. 30.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. arXiv preprint arXiv:1706.03762, 2017b.
  31. 31.Oriol Vinyals, Charles Blundell, Tim Lillicrap, Daan Wierstra, et al. Matching networks for one shot learning. In Advances in Neural Information Processing Systems (NIPS), 2016.
  32. 32.Jane X Wang, Zeb Kurth-Nelson, Dhruva Tirumala, Hubert Soyer, Joel Z Leibo, Remi Munos, Charles Blundell, Dharshan Kumaran, and Matt Botvinick. Learning to reinforcement learn. arXiv preprint arXiv:1611.05763, 2016.

Citation

MLA
Mishra, N., et al. “A Simple Neural Attentive Meta-Learner”. arXiv, 2017, http://arxiv.org/abs/1707.03141v3.
APA
Mishra, N., Rohaninejad, M., Chen, X., & Abbeel, P. (2017). A Simple Neural Attentive Meta-Learner. arXiv. http://arxiv.org/abs/1707.03141v3
Chicago
Mishra, N., M. Rohaninejad, X. Chen, and P. Abbeel. 2017. “A Simple Neural Attentive Meta-Learner”. arXiv. http://arxiv.org/abs/1707.03141v3.
Harvard
Mishra, N. et al. (2017) “A Simple Neural Attentive Meta-Learner”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1707.03141v3.
Vancouver
1. Mishra N, Rohaninejad M, Chen X, Abbeel P (2017) A Simple Neural Attentive Meta-Learner. arXiv

BibTeX

@article{mishra2017simple,
  title = {A Simple Neural Attentive Meta-Learner},
  author = {Mishra, Nikhil and Rohaninejad, Mostafa and Chen, Xi and Abbeel, Pieter},
  year = {2017},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1707.03141v3},
  eprint = {1707.03141}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Published with permission