Few-Shot Learning with Graph Neural Networks
Victor GarciaJoan Bruna
Proposes a graph neural network architecture that unifies prior few-shot learning models through message passing on partially labeled graphs, achieving competitive classification performance while directly supporting semi-supervised and active learning extensions.
Modern machine learning models achieve remarkable accuracy when trained on vast amounts of annotated data, but they struggle in real-world scenarios where only a handful of labeled examples are available. Acquiring extensive labels is often costly, time-consuming, or impractical. While few-shot learning methods attempt to overcome this by transferring knowledge across related tasks, existing approaches frequently rely on rigid architectures, complex recurrent models, or separate image and label processing pipelines that require heavy parameter budgets.
The article demonstrates that few-shot learning can be reformulated as information propagation over a graph neural network, unifying standard few-shot classification, semi-supervised learning, and active learning under a single end-to-end architecture.
To evaluate this framework, the authors map collections of labeled and unlabeled images into a fully connected graph, where nodes represent image-label representations and edge weights correspond to a learned similarity metric. They conduct extensive image classification experiments on standard benchmarks, including the Omniglot character dataset (1,623 classes) and the more complex Mini-ImageNet dataset (100 classes with 84x84 color images), testing configurations across varying class counts, label ratios, and sample sizes.
Key findings show that the proposed approach delivers competitive accuracy while dramatically reducing model complexity. On Omniglot, the graph neural network achieves state-of-the-art accuracy of 99.2% on 5-way 1-shot tasks and 97.4% on 20-way 1-shot tasks, matching leading recurrent architectures while reducing parameter counts by roughly 94% (from approximately 5 million to around 300,000). On Mini-ImageNet, the method reaches 50.33% accuracy in 1-shot and 66.41% in 5-shot settings, reducing parameter size by roughly 96% compared to temporal convolution baselines (from about 11 million to approximately 400,000 parameters). In semi-supervised settings with only 20% labeled data, the model extracts sufficient structural information from unlabeled samples to match performance levels achieved with 40% supervised labels on Omniglot and yields an absolute accuracy boost of approximately 2% on Mini-ImageNet. When augmented with active learning, the architecture successfully queries the most informative unlabeled samples, improving Mini-ImageNet classification accuracy by roughly 3.4 percentage points over random querying.
These results demonstrate that graph-based message passing effectively captures relational dependencies between images and labels without needing bloated parameter sets. Organizations implementing this strategy can achieve substantial compute savings, lower deployment costs, and faster prototyping turnaround times for low-data computer vision systems. Furthermore, the framework's native support for active learning lowers labeling costs by strategically pinpointing high-value data points during annotation campaigns.
Organizations operating in data-constrained domains should pilot graph neural network architectures for visual recognition, especially where labeling capacity is limited and active querying can provide a high return on investment. Future technical efforts should explore graph coarsening and hierarchical scaling methods to support larger datasets with millions of nodes, as well as extending the active query mechanism into reinforcement learning environments.
Decision-makers should note that evaluations are currently confined to standard academic vision benchmarks using relatively small sample sizes per task. While confidence in the comparative efficiency and accuracy is high across these benchmarks, practitioners should conduct pilot validations before deploying the model on highly complex visual distributions, where richer feature extraction backbones may be necessary to maximize real-world performance.
- Paper: Matching Networks for One Shot Learning, Oriol Vinyals et al. (2016). It introduces episodic meta-learning with memory and attention mechanisms for one-shot classification, which provides foundational motivation for graph-based few-shot formulations.
- Paper: Prototypical Networks for Few-shot Learning, Jake Snell et al. (2017). It establishes metric-based few-shot classification using learned class representations, a direct baseline generalized by graph message-passing models.
- Paper: The Graph Neural Network Model, Franco Scarselli et al. (2009). It provides the foundational graph neural network architecture and state-update mechanisms adapted to model relational inference over partially labeled datasets.
- Paper: Gated Graph Sequence Neural Networks, Yujia Li et al. (2015). It introduces modern gated message-passing dynamics on graphs that directly inform neural inference procedures over interconnected support and query sets.
- Paper: Combining active learning and semi-supervised learning using Gaussian fields and harmonic functions, Xiaojin Zhu et al. (2003). It outlines graph-based label propagation across partially labeled nodes for combined active and semi-supervised learning, framing the exact problem setting solved by neural message passing.
- Paper: Efficient Inference in Fully Connected CRFs with Gaussian Edge Potentials, Philipp Krähenbühl et al. (2011). It details fast mean-field message passing in dense graphical models, motivating the formulation of few-shot image classification as inference on fully connected graphs.
- Paper: Optimization as a Model for Few-Shot Learning, Sachin Ravi et al. (2017). It formalizes meta-learning as learning optimization and inference procedures over episodic few-shot classification benchmarks.
- Paper: Meta-Learning with Memory-Augmented Neural Networks, Adam Santoro et al. (2016). It demonstrates how neural networks can rapidly bind novel inputs to labels in few-shot regimes using episodic memory mechanisms.
- Paper: Meta-Learning for Semi-Supervised Few-Shot Classification, Mengye Ren et al. (2018). It builds on the semi-supervised few-shot learning paradigm by refining prototypes with unlabeled examples and handling distractor classes.
- Paper: Learning to Compare: Relation Network for Few-Shot Learning, Flood Sung et al. (2017). It presents a closely related metric-relational approach by learning deep non-linear relation modules for pairwise and episodic few-shot comparisons.
- Paper: TADAM: Task dependent adaptive metric for improved few-shot learning, Boris N. Oreshkin et al. (2018). It extends few-shot metric learning architectures by incorporating task-dependent feature conditioning and adaptive metric scaling.
- Paper: Meta-Learning With Differentiable Convex Optimization, Kwonjoon Lee et al. (2019). It advances beyond simple message passing by embedding convex optimization and discriminative SVM base learners directly into the meta-learning loop.
- Paper: Meta-Learning with Latent Embedding Optimization, Andrei A. Rusu et al. (2018). It extends meta-learning adaptation into compact latent spaces to overcome optimization instabilities in few-shot regimes.
- Paper: A Closer Look at Few-shot Classification, Wei-Yu Chen et al. (2019). It offers an empirical critique and standardized benchmarking across few-shot architectures, assessing the relative value of meta-learning versus standard representations.
- Paper: Graph Neural Networks: A Review of Methods and Applications, Jie Zhou et al. (2018). It synthesizes subsequent progress in graph neural network formulations, including applications where graphs are inferred dynamically for relational tasks.
