Few-Shot Learning with Graph Neural Networks

Victor GarciaJoan Bruna

article2017ICLR1,374 citations

Proposes a graph neural network architecture that unifies prior few-shot learning models through message passing on partially labeled graphs, achieving competitive classification performance while directly supporting semi-supervised and active learning extensions.

Listen

Modern machine learning models achieve remarkable accuracy when trained on vast amounts of annotated data, but they struggle in real-world scenarios where only a handful of labeled examples are available. Acquiring extensive labels is often costly, time-consuming, or impractical. While few-shot learning methods attempt to overcome this by transferring knowledge across related tasks, existing approaches frequently rely on rigid architectures, complex recurrent models, or separate image and label processing pipelines that require heavy parameter budgets.

The article demonstrates that few-shot learning can be reformulated as information propagation over a graph neural network, unifying standard few-shot classification, semi-supervised learning, and active learning under a single end-to-end architecture.

To evaluate this framework, the authors map collections of labeled and unlabeled images into a fully connected graph, where nodes represent image-label representations and edge weights correspond to a learned similarity metric. They conduct extensive image classification experiments on standard benchmarks, including the Omniglot character dataset (1,623 classes) and the more complex Mini-ImageNet dataset (100 classes with 84x84 color images), testing configurations across varying class counts, label ratios, and sample sizes.

Key findings show that the proposed approach delivers competitive accuracy while dramatically reducing model complexity. On Omniglot, the graph neural network achieves state-of-the-art accuracy of 99.2% on 5-way 1-shot tasks and 97.4% on 20-way 1-shot tasks, matching leading recurrent architectures while reducing parameter counts by roughly 94% (from approximately 5 million to around 300,000). On Mini-ImageNet, the method reaches 50.33% accuracy in 1-shot and 66.41% in 5-shot settings, reducing parameter size by roughly 96% compared to temporal convolution baselines (from about 11 million to approximately 400,000 parameters). In semi-supervised settings with only 20% labeled data, the model extracts sufficient structural information from unlabeled samples to match performance levels achieved with 40% supervised labels on Omniglot and yields an absolute accuracy boost of approximately 2% on Mini-ImageNet. When augmented with active learning, the architecture successfully queries the most informative unlabeled samples, improving Mini-ImageNet classification accuracy by roughly 3.4 percentage points over random querying.

These results demonstrate that graph-based message passing effectively captures relational dependencies between images and labels without needing bloated parameter sets. Organizations implementing this strategy can achieve substantial compute savings, lower deployment costs, and faster prototyping turnaround times for low-data computer vision systems. Furthermore, the framework's native support for active learning lowers labeling costs by strategically pinpointing high-value data points during annotation campaigns.

Organizations operating in data-constrained domains should pilot graph neural network architectures for visual recognition, especially where labeling capacity is limited and active querying can provide a high return on investment. Future technical efforts should explore graph coarsening and hierarchical scaling methods to support larger datasets with millions of nodes, as well as extending the active query mechanism into reinforcement learning environments.

Decision-makers should note that evaluations are currently confined to standard academic vision benchmarks using relatively small sample sizes per task. While confidence in the comparative efficiency and accuracy is high across these benchmarks, practitioners should conduct pilot validations before deploying the model on highly complex visual distributions, where richer feature extraction backbones may be necessary to maximize real-world performance.

Cover for Few-Shot Learning with Graph Neural Networks

Abstract

We propose to study the problem of few-shot learning with the prism of inference on a partially observed graphical model, constructed from a collection of input images whose label can be either observed or not. By assimilating generic message-passing inference algorithms with their neural-network counterparts, we define a graph neural network architecture that generalizes several of the recently proposed few-shot learning models. Besides providing improved numerical performance, our framework is easily extended to variants of few-shot learning, such as semi-supervised or active learning, demonstrating the ability of graph-based models to operate well on 'relational' tasks.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Problem Set-up
  • 4 Model
  • 4.1 Set and Graph Input Representations
  • 4.2 Graph Neural Networks
  • 4.3 Relationship with Existing Models
  • 5 Training
  • 5.1 Few-Shot and Semi-Supervised Learning
  • 5.2 Active Learning
  • 6 Experiments
  • 6.1 Datasets and Implementation
  • 6.1.1 Omniglot
  • 6.1.2 Mini-Imagenet
  • 6.2 Few-Shot
  • 6.3 Semi-Supervised
  • 6.4 Active Learning
  • 7 Conclusions
  • References

Knowls

  1. Knowl 1 — Graph Neural Network Architecture for Few-Shot Learning

    model/method

    The model formulates few-shot learning as supervised message passing on a fully connected graph G=(V,E)G = (V, E), where each node vi∈Vv_i \in V corresponds to an input image xix_i (either labeled or unlabeled).

    Initial Node Features: For an input collection T\mathcal{T} containing samples from KK classes, each node viv_i is initialized with: xi(0)=(ϕ(xi),h(li))x^{(0)}_i = (\phi(x_i), h(l_i)) where ϕ(xi)∈Rd\phi(x_i) \in \mathbb{R}^d is a feature vector extracted by a convolutional neural network (CNN), and h(li)∈RKh(l_i) \in \mathbb{R}^K encodes label information:

    • If image xix_i has observed label li∈{1,…,K}l_i \in \{1, \dots, K\}, h(li)h(l_i) is a one-hot indicator vector.
    • If xix_i has an unknown label (such as unlabeled auxiliary images or the query image to classify), h(li)=K−11Kh(l_i) = K^{-1} \mathbf{1}_K, representing a uniform distribution over the KK classes.

    Trainable Edge Feature / Adjacency Learning: At each GNN layer kk, pairwise edge weights A~i,j(k)\tilde{A}^{(k)}_{i,j} are computed from the current node representations xi(k)x^{(k)}_i and xj(k)x^{(k)}_j using a parametric symmetric metric: A~i,j(k)=MLPθ~(∣xi(k)−xj(k)∣)\tilde{A}^{(k)}_{i,j} = \text{MLP}_{\tilde{\theta}}\left(\left|x^{(k)}_i - x^{(k)}_j\right|\right) where MLPθ~\text{MLP}_{\tilde{\theta}} is a multi-layer perceptron acting on element-wise absolute differences, guaranteeing the distance symmetry property A~i,j(k)=A~j,i(k)\tilde{A}^{(k)}_{i,j} = \tilde{A}^{(k)}_{j,i}. The adjacency matrix is normalized along each row to a stochastic kernel: Ai,j(k)=exp⁡(A~i,j(k))∑j′exp⁡(A~i,j′(k))A^{(k)}_{i,j} = \frac{\exp(\tilde{A}^{(k)}_{i,j})}{\sum_{j'} \exp(\tilde{A}^{(k)}_{i,j'})}

    Graph Convolutional Layer: Using the generator family A={A(k),1}\mathcal{A} = \{A^{(k)}, \mathbf{1}\}, node features are updated via: xl(k+1)=ρ(∑B∈ABx(k)θB,l(k)),l=1,…,dk+1x^{(k+1)}_l = \rho\left(\sum_{B \in \mathcal{A}} B x^{(k)} \theta^{(k)}_{B,l}\right), \quad l = 1, \dots, d_{k+1} where θB(k)∈Rdk×dk+1\theta^{(k)}_B \in \mathbb{R}^{d_k \times d_{k+1}} are trainable weight matrices, ρ(⋅)\rho(\cdot) is a Leaky ReLU activation, and 1\mathbf{1} is the identity operator. Dense skip connections concatenate preceding representations [Gc(x(k)),x(k)][Gc(x^{(k)}), x^{(k)}].

    Classification Objective: For a target query image x∗x_* associated with graph node ∗*, the final node features are mapped via softmax over KK classes to yield P(Y∗=yk∣T)P(Y_* = y_k \mid \mathcal{T}). The network parameters Θ\Theta are trained end-to-end using the cross-entropy loss: L(Φ(T;Θ),Y)=−∑k=1Kyklog⁡P(Y∗=yk∣T)\mathcal{L}(\Phi(\mathcal{T}; \Theta), Y) = -\sum_{k=1}^K y_k \log P(Y_* = y_k \mid \mathcal{T})

  2. Knowl 2 — Graph-Based Formulation of Few-Shot, Semi-Supervised, and Active Learning

    definition

    Meta-learning tasks are cast as supervised interpolation problems over an input collection of images T\mathcal{T} and target labels YY drawn from an underlying distribution: T=({(x1,l1),…,(xs,ls)},{x~1,…,x~r},{xˉ1,…,xˉt}),Y=(y1,…,yt)∈{1,…,K}t\mathcal{T} = \left(\{ (x_1, l_1), \dots, (x_s, l_s) \}, \{ \tilde{x}_1, \dots, \tilde{x}_r \}, \{ \bar{x}_1, \dots, \bar{x}_t \} \right), \quad Y = (y_1, \dots, y_t) \in \{1, \dots, K\}^t where ss is the number of labeled samples (li∈{1,…,K}l_i \in \{1, \dots, K\}), rr is the number of unlabeled auxiliary samples, tt is the number of query samples to classify, and KK is the number of classes. Each image is sampled from a class-specific distribution Pl(RN)P_l(\mathbb{R}^N). The collection T\mathcal{T} defines a fully connected graph GT=(V,E)G_{\mathcal{T}} = (V, E) with ∣V∣=s+r+t|V| = s + r + t nodes.

    The framework covers three learning paradigms by varying these parameters:

    • qq-shot, KK-way Few-Shot Learning: r=0r = 0, t=1t = 1, and s=qKs = qK, with exactly qq labeled examples per class and one unlabeled query image to classify.
    • Semi-Supervised Few-Shot Learning: r>0r > 0 and t=1t = 1, where rr unlabeled images from the same underlying class distributions are included in the graph to aid label propagation.
    • Active Learning: r>0r > 0 and t=1t = 1, where the model can actively query the true labels of a subset of the unlabeled auxiliary images {x~1,…,x~r}\{\tilde{x}_1, \dots, \tilde{x}_r\} to maximally reduce query classification error.
  3. Knowl 3 — Generalization of Existing Few-Shot Learning Models via Graph Message Passing

    theoretical result

    The graph neural network (GNN) formulation for few-shot learning unifies and generalizes several existing meta-learning approaches as specific restricted cases of graph message passing:

    1. Siamese Networks: Corresponds to a single-layer message-passing iteration on initial node features xi(0)=(ϕ(xi),hi)x^{(0)}_i = (\phi(x_i), h_i) with a fixed non-trainable distance kernel: A~i,j(0)=softmax(−∥ϕ(xi)−ϕ(xj)∥)\tilde{A}^{(0)}_{i,j} = \text{softmax}(-\|\phi(x_i) - \phi(x_j)\|) and target label prediction Y^∗=∑jA~∗,j(0)⟨xj(0),u⟩\hat{Y}_* = \sum_j \tilde{A}^{(0)}_{*,j} \langle x^{(0)}_j, u \rangle, where uu extracts the label field.

    2. Prototypical Networks: Corresponds to a two-step GNN procedure. In the first step, an exact block-diagonal adjacency averages embeddings within each class cluster: A~i,j(0)={q−1if li=lj0otherwise  ⟹  xi(1)=∑jA~i,j(0)xj(0)\tilde{A}^{(0)}_{i,j} = \begin{cases} q^{-1} & \text{if } l_i = l_j \\ 0 & \text{otherwise} \end{cases} \implies x^{(1)}_i = \sum_j \tilde{A}^{(0)}_{i,j} x^{(0)}_j In the second step, prototype representations x(1)x^{(1)} interact with the query via a standard softmax kernel A~(1)=softmax(ϕ)\tilde{A}^{(1)} = \text{softmax}(\phi) to yield Y^∗=∑jA~∗,j(1)⟨xj(1),u⟩\hat{Y}_* = \sum_j \tilde{A}^{(1)}_{*,j} \langle x^{(1)}_j, u \rangle.

    3. Matching Networks: Uses edge attention A~∗,j(k)=ϕ(x∗(k),xj(T))\tilde{A}^{(k)}_{*, j} = \phi(x^{(k)}_*, x^{(T)}_j) where the support set encodings xj(T)x^{(T)}_j are computed independently of the query via bidirectional LSTMs, and labels are linearly combined in a final step. In contrast, the general GNN updates node and edge representations jointly across all nodes across multiple layers, capturing multi-hop dependencies between images and labels.

  4. Knowl 4 — Active Learning Acquisition via GNN Node Attention

    model/method

    In the active learning regime with rr unlabeled auxiliary samples {x~1,…,x~r}\{\tilde{x}_1, \dots, \tilde{x}_r\}, the GNN queries the label of the most informative unlabeled sample after the first layer of graph convolution:

    1. Attention Scoring: A two-layer neural network g:Rd1→Rg: \mathbb{R}^{d_1} \to \mathbb{R} computes an unnormalized scalar score for each unlabeled node representation xi(1)x^{(1)}_i, which is normalized across all rr unlabeled candidates using a softmax: Attention=Softmax(g(x{1,…,r}(1)))\text{Attention} = \text{Softmax}\left(g\left(x^{(1)}_{\{1, \dots, r\}}\right)\right)

    2. Sample Selection:

    • At training time, an index i∗i^* is sampled according to the multinomial distribution defined by Attention\text{Attention}.
    • At test time, the index with the maximum score is chosen deterministically: i∗=arg⁡max⁡iAttentionii^* = \arg\max_i \text{Attention}_i.
    1. Label Update: The true label vector h(li∗)h(l_{i^*}) of the chosen sample is queried, scaled by a hyperparameter w∈(0,1)w \in (0, 1), and added into the label portion of the node representation: xi∗(1)=[Gc(xi∗(0)),(ϕ(xi∗),h(li∗)+w⋅h(li∗))]x^{(1)}_{i^*} = \left[ G_c(x^{(0)}_{i^*}), \left(\phi(x_{i^*}), h(l_{i^*}) + w \cdot h(l_{i^*})\right) \right]

    2. Forward Propagation: The updated node representation is propagated through the subsequent GNN layers. The acquisition function gg is trained end-to-end via backpropagation from the final query classification cross-entropy loss.

  5. Knowl 5 — Omniglot Few-Shot Classification Performance

    data/table

    Evaluation of few-shot classification accuracy on the Omniglot benchmark across 5-way and 20-way settings under 1-shot and 5-shot conditions.

    Model 5-Way 20-Way
    1-shot 5-shot 1-shot 5-shot
    Pixels 41.7% 63.2% 26.7% 42.6%
    Siamese Net 97.3% 98.4% 88.2% 97.0%
    Matching Networks 98.1% 98.9% 93.8% 98.5%
    Neural Statistician 98.1% 99.5% 93.2% 98.1%
    Res. Pair-Wise – – 94.8% –
    Prototypical Networks 97.4% 99.3% 95.4% 98.8%
    ConvNet with Memory 98.4% 99.6% 95.0% 98.6%
    Agnostic Meta-learner (MAML) 98.7 ±\pm 0.4% 99.9 ±\pm 0.3% 95.8 ±\pm 0.3% 98.9 ±\pm 0.2%
    Meta Networks 98.9% – 97.0% –
    TCML 98.96 ±\pm 0.20% 99.75 ±\pm 0.11% 97.64 ±\pm 0.30% 99.36 ±\pm 0.18%
    Our GNN 99.2% 99.7% 97.4% 99.0%

    The GNN achieves competitive classification accuracy across all evaluation settings, matching or outperforming other meta-learning baselines. The 3-layer GNN model contains approximately 300K parameters, compared to approximately 5M parameters in the Temporal Convolutional Meta-Learner (TCML) architecture.

  6. Knowl 6 — Mini-ImageNet Few-Shot Classification Performance

    data/table

    Classification accuracy on 5-way 1-shot and 5-shot Mini-ImageNet tasks reported with 95% confidence intervals.

    Model 5-Way
    1-shot 5-shot
    Matching Networks 43.6% 55.3%
    Prototypical Networks 46.61 ±\pm 0.78% 65.77 ±\pm 0.70%
    Model Agnostic Meta-learner (MAML) 48.70 ±\pm 1.84% 63.1 ±\pm 0.92%
    Meta Networks 49.21 ±\pm 0.96% –
    Ravi Larochelle 43.4 ±\pm 0.77% 60.2 ±\pm 0.71%
    TCML 55.71 ±\pm 0.99% 68.88 ±\pm 0.92%
    Our metric learning + KNN 49.44 ±\pm 0.28% 64.02 ±\pm 0.51%
    Our GNN 50.33 ±\pm 0.36% 66.41 ±\pm 0.63%

    The baseline "Our metric learning + KNN" applies kk-nearest neighbors directly to the learned pairwise metric ϕθ(xi(0),xj(0))\phi_\theta(x^{(0)}_i, x^{(0)}_j) without graph message passing. Incorporating the full GNN architecture to aggregate information across nodes provides an accuracy improvement of 0.89% in the 1-shot setting and 2.39% (from 64.02% to 66.41%) in the 5-shot setting. The 3-layer GNN uses approximately 400K parameters, compared to ~11M parameters for TCML.

  7. Knowl 7 — Semi-Supervised Few-Shot Classification Performance

    data/table

    Performance of the GNN model in 5-way 5-shot semi-supervised classification on Omniglot and Mini-ImageNet, where either 20% (1 labeled example per class) or 40% (2 labeled examples per class) of the support set is labeled, with remaining instances provided as unlabeled nodes.

    Omniglot Accuracies (5-Way 5-Shot):

    Model 20%-labeled 40%-labeled 100%-labeled
    GNN - Trained only with labeled 99.18% 99.59% 99.71%
    GNN - Semi supervised 99.59% 99.63% 99.71%

    Mini-ImageNet Accuracies (5-Way 5-Shot, 95% Confidence Intervals):

    Model 20%-labeled 40%-labeled 100%-labeled
    GNN - Trained only with labeled 50.33 ±\pm 0.36% 56.91 ±\pm 0.42% 66.41 ±\pm 0.63%
    GNN - Semi supervised 52.45 ±\pm 0.88% 58.76 ±\pm 0.86% 66.41 ±\pm 0.63%

    On Omniglot, supplying unlabeled examples in the 20%-labeled setting (1 labeled sample per class + 4 unlabeled samples per class) achieves 99.59% accuracy, matching the accuracy of having twice as many labeled examples (40%-labeled supervised setting). On Mini-ImageNet, incorporating unlabeled data improves accuracy by approximately 2% absolute over the strictly supervised baseline in both the 20% and 40% labeled regimes.

  8. Knowl 8 — Active Learning Performance on Omniglot and Mini-ImageNet

    data/table

    Evaluation of active learning on 5-way 5-shot tasks with initially 20% labeled samples (1 labeled and 4 unlabeled samples per class), comparing learned query selection (GNN - AL) against random query selection (GNN - Random) when requesting the true label of one unlabeled instance.

    Method Omniglot (5-Way 5-shot 20%-labeled) Mini-ImageNet (5-Way 5-shot 20%-labeled)
    GNN - AL 99.62% 55.99 ±\pm 1.35%
    GNN - Random 99.59% 52.56 ±\pm 1.18%

    On Mini-ImageNet, random sample selection (GNN - Random) achieves 52.56%, which is virtually indistinguishable from the passive semi-supervised setting without label queries (52.45%). The learned active query mechanism (GNN - AL) achieves 55.99%, providing an absolute improvement of ~3.43% by identifying informative instances to label. On Omniglot, active selection increases accuracy from 99.59% to 99.62%.

  9. Knowl 9 — Embedding and Graph Convolutional Architectures

    experimental setup

    The experimental implementation uses specialized CNN feature extractor networks ϕ\phi and 3-block GNN architectures:

    Omniglot Embedding Network ϕ\phi: Consists of 4 convolutional blocks, each containing a 3×33 \times 3 convolution with 64 filters, batch normalization, 2×22 \times 2 max pooling, and Leaky ReLU activations, followed by a fully connected layer producing a 64-dimensional feature vector.

    Mini-ImageNet Embedding Network ϕ\phi: Consists of 4 convolutional layers and a fully connected layer producing a 128-dimensional feature vector:

    1. 3×33 \times 3 conv (64 filters), batch normalization, 2×22 \times 2 max pooling, Leaky ReLU.
    2. 3×33 \times 3 conv (96 filters), batch normalization, 2×22 \times 2 max pooling, Leaky ReLU.
    3. 3×33 \times 3 conv (128 filters), batch normalization, 2×22 \times 2 max pooling, Leaky ReLU, Dropout (0.5).
    4. 3×33 \times 3 conv (256 filters), batch normalization, 2×22 \times 2 max pooling, Leaky ReLU, Dropout (0.5).
    5. Fully connected layer (128 units), batch normalization.

    GNN Architecture: Composed of 3 stacked blocks, where each block kk contains:

    • Adjacency module: Computes absolute feature differences ∣xi(k)−xj(k)∣|x^{(k)}_i - x^{(k)}_j|, processed through two fully connected layers with 2×nf2 \times nf hidden units, two fully connected layers with nfnf hidden units (all with batch normalization and Leaky ReLU), and a linear layer to 1 scalar output, normalized across rows with softmax (nf=96nf = 96).
    • Graph convolution module: Applies a graph convolution with 48 output channels, batch normalization, Leaky ReLU, and concatenation with the input features x(k)x^{(k)} via dense connections.

Coverage note — None was omitted; all key contributions including problem setup, GNN architecture, active learning extension, theoretical generalizations, experimental setups, and empirical results across Omniglot and Mini-ImageNet are covered.

References

  1. 1.Peter Battaglia, Razvan Pascanu, Matthew Lai, Danilo Jimenez Rezende, et al. Interaction networks for learning about objects, relations and physics. In Advances in Neural Information Processing Systems, pp. 4502–4510, 2016.
  2. 2.Michael M Bronstein, Joan Bruna, Yann LeCun, Arthur Szlam, and Pierre Vandergheynst. Geometric deep learning: going beyond euclidean data. IEEE Signal Processing Magazine, 34(4):18–42, 2017.
  3. 3.Joan Bruna and Xiang Li. Community detection with graph neural networks. arXiv preprint arXiv:1705.08415, 2017.
  4. 4.Joan Bruna, Wojciech Zaremba, Arthur Szlam, and Yann LeCun. Spectral networks and locally connected networks on graphs. Proc. ICLR, 2013.
  5. 5.Michael B. Chang, Tomer Ullman, Antonio Torralba, and Joshua B. Tenenbaum. A compositional object-based approach to learning physical dynamics. ICLR, 2016.
  6. 6.Michaãl Defferrard, Xavier Bresson, and Pierre Vandergheynst. Convolutional neural networks on graphs with fast localized spectral filtering. In Advances in Neural Information Processing Systems, pp. 3837–3845, 2016.
  7. 7.David Duvenaud, Dougal Maclaurin, Jorge Aguilera-Iparraguirre, Rafael Gómez-Bombarelli, Timothy Hirzel, Alán Aspuru-Guzik, and Ryan P Adams. Convolutional networks on graphs for learning molecular fingerprints. In Neural Information Processing Systems, 2015.
  8. 8.Harrison Edwards and Amos Storkey. Towards a neural statistician. arXiv preprint arXiv:1606.02185, 2016.
  9. 9.Li Fei-Fei, Rob Fergus, and Pietro Perona. One-shot learning of object categories. IEEE transactions on pattern analysis and machine intelligence, 28(4):594–611, 2006.
  10. 10.Chelsea Finn, Pieter Abbeel, and Sergey Levine. Model-agnostic meta-learning for fast adaptation of deep networks. arXiv preprint arXiv:1703.03400, 2017.
  11. 11.Justin Gilmer, Samuel S Schoenholz, Patrick F Riley, Oriol Vinyals, and George E Dahl. Neural message passing for quantum chemistry. arXiv preprint arXiv:1704.01212, 2017.
  12. 12.M. Gori, G. Monfardini, and F. Scarselli. A new model for learning in graph domains. In Proc. IJCNN, 2005.
  13. 13.M. Henaff, J. Bruna, and Y. LeCun. Deep convolutional networks on graph-structured data. arXiv:1506.05163, 2015.
  14. 14.Łukasz Kaiser, Ofir Nachum, Aurko Roy, and Samy Bengio. Learning to remember rare events. arXiv preprint arXiv:1703.03129, 2017.
  15. 15.Steven Kearnes, Kevin McCloskey, Marc Berndl, Vijay Pande, and Patrick Riley. Molecular graph convolutions: moving beyond fingerprints. Journal of computer-aided molecular design, 30(8): 595–608, 2016.
  16. 16.Thomas N Kipf and Max Welling. Semi-supervised classification with graph convolutional networks. arXiv preprint arXiv:1609.02907, 2016.
  17. 17.Gregory Koch, Richard Zemel, and Ruslan Salakhutdinov. Siamese neural networks for one-shot image recognition. In ICML Deep Learning Workshop, volume 2, 2015.
  18. 18.Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. Imagenet classification with deep convolutional neural networks. In Advances in neural information processing systems, pp. 1097–1105, 2012.
  19. 19.Brenden M Lake, Ruslan Salakhutdinov, and Joshua B Tenenbaum. Human-level concept learning through probabilistic program induction. Science, 350(6266):1332–1338, 2015.
  20. 20.Yujia Li, Daniel Tarlow, Marc Brockschmidt, and Richard Zemel. Gated graph sequence neural networks. arXiv preprint arXiv:1511.05493, 2015.
  21. 21.Akshay Mehrotra and Ambedkar Dukkipati. Generative adversarial residual pairwise networks for one shot learning. arXiv preprint arXiv:1703.08033, 2017.
  22. 22.Nikhil Mishra, Mostafa Rohaninejad, Xi Chen, and Pieter Abbeel. Meta-learning with temporal convolutions. arXiv preprint arXiv:1707.03141, 2017.
  23. 23.Tsendsuren Munkhdalai and Hong Yu. Meta networks. arXiv preprint arXiv:1703.00837, 2017.
  24. 24.Sachin Ravi and Hugo Larochelle. Optimization as a model for few-shot learning. ICLR, 2016.
  25. 25.Anselm Rothe, Brenden Lake, and Todd Gureckis. Question asking as program generation. NIPS, 2017.
  26. 26.Adam Santoro, Sergey Bartunov, Matthew Botvinick, Daan Wierstra, and Timothy Lillicrap. Meta-learning with memory-augmented neural networks. In International conference on machine learning, pp. 1842–1850, 2016.
  27. 27.Franco Scarselli, Marco Gori, Ah Chung Tsoi, Markus Hagenbuchner, and Gabriele Monfardini. The graph neural network model. IEEE Transactions on Neural Networks, 20(1):61–80, 2009.
  28. 28.Jake Snell, Kevin Swersky, and Richard S Zemel. Prototypical networks for few-shot learning. arXiv preprint arXiv:1703.05175, 2017.
  29. 29.Sainbayar Sukhbaatar, Rob Fergus, et al. Learning multiagent communication with backpropagation. In Advances in Neural Information Processing Systems, pp. 2244–2252, 2016.
  30. 30.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. arXiv preprint arXiv:1706.03762, 2017.
  31. 31.Oriol Vinyals, Samy Bengio, and Manjunath Kudlur. Order matters: Sequence to sequence for sets. arXiv preprint arXiv:1511.06391, 2015.
  32. 32.Oriol Vinyals, Charles Blundell, Tim Lillicrap, Daan Wierstra, et al. Matching networks for one shot learning. In Advances in Neural Information Processing Systems, pp. 3630–3638, 2016.
  33. 33.Bing Xu, Naiyan Wang, Tianqi Chen, and Mu Li. Empirical evaluation of rectified activations in convolutional network. arXiv preprint arXiv:1505.00853, 2015.

Citation

MLA
Garcia, V., and J. Bruna. “Few-Shot Learning with Graph Neural Networks”. arXiv, 2017, http://arxiv.org/abs/1711.04043v3.
APA
Garcia, V., & Bruna, J. (2017). Few-Shot Learning with Graph Neural Networks. arXiv. http://arxiv.org/abs/1711.04043v3
Chicago
Garcia, V., and J. Bruna. 2017. “Few-Shot Learning with Graph Neural Networks”. arXiv. http://arxiv.org/abs/1711.04043v3.
Harvard
Garcia, V. and Bruna, J. (2017) “Few-Shot Learning with Graph Neural Networks”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1711.04043v3.
Vancouver
1. Garcia V, Bruna J (2017) Few-Shot Learning with Graph Neural Networks. arXiv

BibTeX

@article{garcia2017few,
  title = {Few-Shot Learning with Graph Neural Networks},
  author = {Garcia, Victor and Bruna, Joan},
  year = {2017},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1711.04043v3},
  eprint = {1711.04043}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors