EASE: Unsupervised Discriminant Subspace Learning for Transductive Few-Shot Learning
Hao ZhuPiotr Koniusz
Proposes an unsupervised subspace projection and a constrained Wasserstein clustering method to efficiently separate novel classes at test time without backbone fine-tuning, achieving strong performance gains across multiple transductive few-shot benchmarks.
Modern computer vision models typically require large volumes of labeled data to achieve high accuracy, making deployment costly and difficult in specialized domains where human annotation requires rare expertise. Few-shot learning addresses this bottleneck by enabling models to recognize new visual categories from only a few labeled examples. The article introduces an efficient inference approach designed to boost classification accuracy by learning a compact, discriminant feature space directly during test time without requiring expensive retraining or complex meta-learning pipelines.
The article develops and evaluates two modular techniques: Unsupervised Discriminant Subspace Learning (EASE) and Constrained Wasserstein Mean Shift Clustering (SIAMESE). EASE captures the underlying structure of image data by generating similarity and dissimilarity relationships across labeled and unlabeled samples, solving for an optimal linear projection via closed-form singular value decomposition. SIAMESE then refines the estimated class centers and final label assignments using optimal transport principles that account for class balance constraints and known support labels. The framework was evaluated across five standard benchmark image datasets (mini-ImageNet, tiered-ImageNet, CIFAR-FS, CUB, and OpenMIC) under both transductive and semi-supervised conditions using standard deep neural network backbones.
The empirical findings demonstrate significant improvements over existing approaches. When paired together, the proposed methods outperformed previous state-of-the-art methods across all tested benchmarks, achieving up to an 84.54% accuracy in 1-shot tasks on tiered-ImageNet with a ResNet-12 backbone. In semi-supervised settings, the method provided consistent accuracy gains ranging between 3% and 6% over leading baselines on mini-ImageNet. The approach proved robust across different network architectures, including standard pre-trained networks trained without episodic meta-learning, and operated roughly ten times faster than competing transductive algorithms, requiring only 6 to 9 milliseconds per classification task.
These results show that high-accuracy few-shot classification can be achieved using fast, plug-and-play mathematical operations at inference time rather than costly architectural modifications or specialized training regimes. This significantly lowers computational costs and latency for deploying vision models to edge or enterprise systems that must adapt to new classes on the fly. Organizations seeking to deploy image recognition under severe data constraints should consider integrating test-time subspace projection and optimal transport clustering into their existing computer vision backbones.
Decision-makers should note that the primary performance gains rely on the transductive setting, which assumes access to a batch of unlabeled query samples rather than classifying a single isolated image at a time. Performance also scales with query set size and may degrade if unlabeled data contains entirely unrelated or out-of-distribution classes. Before enterprise deployment, teams should conduct pilot testing on domain-specific data to ensure query batch sizes and data distributions align with operating conditions.
- Paper: Meta-Learning for Semi-Supervised Few-Shot Classification, Mengye Ren et al. (2018). It introduces the semi-supervised episodic few-shot classification setup and prototype clustering on unlabeled data that EASE directly builds upon.
- Paper: Prototypical Networks for Few-shot Learning, Jake Snell et al. (2017). It establishes metric-based prototypical classification, which provides the foundational class-prototype formulation extended by SIAMESE.
- Paper: A Closer Look at Few-shot Classification, Wei-Yu Chen et al. (2019). It provides the standardized benchmarking protocols and pre-trained feature backbone baselines upon which transductive few-shot subspace methods operate.
- Paper: Few-Shot Learning with Graph Neural Networks, Victor Garcia et al. (2017). It formalizes transductive few-shot learning by modeling support and query samples jointly via similarity structures.
- Paper: Meta-Learning With Differentiable Convex Optimization, Kwonjoon Lee et al. (2019). It demonstrates learning convex, discriminative classifiers directly on top of deep embeddings for few-shot adaptation.
- Paper: TADAM: Task dependent adaptive metric for improved few-shot learning, Boris N. Oreshkin et al. (2018). It examines task-dependent metric adaptation and distance scaling for few-shot image classification.
- Paper: Unsupervised Visual Domain Adaptation Using Subspace Alignment, Basura Fernando et al. (2013). It introduces foundational linear subspace alignment methods via SVD for unsupervised domain adaptation.
- Paper: A survey on semi-supervised learning, Jesper E. van Engelen et al. (2019). It surveys the core assumptions and graph-based label propagation principles underlying transductive and semi-supervised learning.
- Paper: Revisiting Prototypical Network for Cross Domain Few-Shot Learning, Fei Zhou et al. (2023). It extends prototypical few-shot classification to challenging cross-domain distributions where test-time feature spaces diverge significantly.
- Paper: Rethinking the Correlation in Few-Shot Segmentation: A Buoys View, Yuan Wang et al. (2023). It applies optimal transport and structural correlation alignment mechanisms between support and query representations to few-shot dense prediction tasks.
- Paper: Few-Shot Object Detection with Foundation Models, Guangxing Han et al. (2024). It builds on modern foundation model representations and metric matching principles to scale few-shot learning to object detection.
