EASE: Unsupervised Discriminant Subspace Learning for Transductive Few-Shot Learning

Hao ZhuPiotr Koniusz

article2022CVPR63 citations

Proposes an unsupervised subspace projection and a constrained Wasserstein clustering method to efficiently separate novel classes at test time without backbone fine-tuning, achieving strong performance gains across multiple transductive few-shot benchmarks.

Listen

Modern computer vision models typically require large volumes of labeled data to achieve high accuracy, making deployment costly and difficult in specialized domains where human annotation requires rare expertise. Few-shot learning addresses this bottleneck by enabling models to recognize new visual categories from only a few labeled examples. The article introduces an efficient inference approach designed to boost classification accuracy by learning a compact, discriminant feature space directly during test time without requiring expensive retraining or complex meta-learning pipelines.

The article develops and evaluates two modular techniques: Unsupervised Discriminant Subspace Learning (EASE) and Constrained Wasserstein Mean Shift Clustering (SIAMESE). EASE captures the underlying structure of image data by generating similarity and dissimilarity relationships across labeled and unlabeled samples, solving for an optimal linear projection via closed-form singular value decomposition. SIAMESE then refines the estimated class centers and final label assignments using optimal transport principles that account for class balance constraints and known support labels. The framework was evaluated across five standard benchmark image datasets (mini-ImageNet, tiered-ImageNet, CIFAR-FS, CUB, and OpenMIC) under both transductive and semi-supervised conditions using standard deep neural network backbones.

The empirical findings demonstrate significant improvements over existing approaches. When paired together, the proposed methods outperformed previous state-of-the-art methods across all tested benchmarks, achieving up to an 84.54% accuracy in 1-shot tasks on tiered-ImageNet with a ResNet-12 backbone. In semi-supervised settings, the method provided consistent accuracy gains ranging between 3% and 6% over leading baselines on mini-ImageNet. The approach proved robust across different network architectures, including standard pre-trained networks trained without episodic meta-learning, and operated roughly ten times faster than competing transductive algorithms, requiring only 6 to 9 milliseconds per classification task.

These results show that high-accuracy few-shot classification can be achieved using fast, plug-and-play mathematical operations at inference time rather than costly architectural modifications or specialized training regimes. This significantly lowers computational costs and latency for deploying vision models to edge or enterprise systems that must adapt to new classes on the fly. Organizations seeking to deploy image recognition under severe data constraints should consider integrating test-time subspace projection and optimal transport clustering into their existing computer vision backbones.

Decision-makers should note that the primary performance gains rely on the transductive setting, which assumes access to a batch of unlabeled query samples rather than classifying a single isolated image at a time. Performance also scales with query set size and may degrade if unlabeled data contains entirely unrelated or out-of-distribution classes. Before enterprise deployment, teams should conduct pilot testing on domain-specific data to ensure query batch sizes and data distributions align with operating conditions.

Cover for EASE: Unsupervised Discriminant Subspace Learning for Transductive Few-Shot Learning

Abstract

Few-shot learning (FSL) has received a lot of attention due to its remarkable ability to adapt to novel classes. Although many techniques have been proposed for FSL, they mostly focus on improving FSL backbones. Some works also focus on learning on top of the features generated by these backbones to adapt them to novel classes. We present an unsuPervised discriminAnt Subspace lEarning (EASE) that improves transductive few-shot learning performance by learning a linear projection onto a subspace built from features of the support set and the unlabeled query set in the test time. Specifically, based on the support set and the unlabeled query set, we generate the similarity matrix and the dissimilarity matrix based on the structure prior for the proposed EASE method, which is efficiently solved with SVD. We also introduce conStraIned wAsserstein MEan Shift clustEring (SIAMESE) which extends Sinkhorn K-means by incorporating labeled support samples. SIAMESE works on the features obtained from EASE to estimate class centers and query predictions. On the mini-ImageNet, tiered-ImageNet, CIFAR-FS, CUB and OpenMIC benchmarks, both steps significantly boost the performance in transductive FSL and semi-supervised FSL.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Methodology
  • 3.1. Model Description
  • 3.2. Unsupervised Discriminant Subspace Learning
  • 3.3. Constrained Wasserstein Mean Shift Clustering (SIAMESE)
  • 4. Experiments
  • 4.1. FSL benchmarks used in our experiments
  • 4.2. Ablations
  • 5. Conclusions
  • References

Knowls

  1. Knowl 1 — Unsupervised Discriminant Subspace Learning (EASE) Objective and Solution

    model/method

    Unsupervised Discriminant Subspace Learning (EASE) optimizes an orthonormal linear projection matrix W∈RO′×OW \in \mathbb{R}^{O' \times O} (with O′≪OO' \ll O) at test time to project feature representations fθ(X)∈R(L+U)×Of_\theta(X) \in \mathbb{R}^{(L+U) \times O} of a support set SS (L=K×NL = K \times N samples) and query set QQ (U=K×BU = K \times B samples) into a lower-dimensional, discriminative subspace without using label supervision.

    Let X=[x1,…,xL+U]⊤X = [x_1, \dots, x_{L+U}]^\top contain all support and query image samples, and let A~sim=Dsim−1/2AsimDsim−1/2\tilde{A}_{sim} = D_{sim}^{-1/2} A_{sim} D_{sim}^{-1/2} and A~dis=Ddis−1/2AdisDdis−1/2\tilde{A}_{dis} = D_{dis}^{-1/2} A_{dis} D_{dis}^{-1/2} denote the normalized affinity matrices for sample similarity and dissimilarity, with Laplacian matrices Lsim=I−A~simL_{sim} = I - \tilde{A}_{sim} and Ldis=I−A~disL_{dis} = I - \tilde{A}_{dis}. The optimization problem balances preserving similarity and maximizing dissimilarity using a balancing hyperparameter α>0\alpha > 0:

    W∗=arg⁡min⁡WW⊤=IWfθ(X)⊤(αA~dis−A~sim)fθ(X)W⊤W^* = \arg\min_{W W^\top = I} W f_\theta(X)^\top (\alpha \tilde{A}_{dis} - \tilde{A}_{sim}) f_\theta(X) W^\top

    The optimal projection matrix W∗W^* is obtained in closed form by solving a generalized eigenvalue problem: selecting the O′O' leading eigenvectors corresponding to the top-O′O' largest eigenvalues of the matrix:

    MEASE=fθ(X)⊤(A~sim−αA~dis)fθ(X)M_{EASE} = f_\theta(X)^\top (\tilde{A}_{sim} - \alpha \tilde{A}_{dis}) f_\theta(X)

  2. Knowl 2 — Low-Rank Representation (LRR) Similarity Matrix Construction

    model/method

    To construct the sample similarity matrix AsimA_{sim} under the prior that feature vectors in a KK-way task lie in a union of KK distinct subspaces (yielding a KK-block-diagonal affinity structure), EASE computes a self-representation coefficient matrix Z∈R(L+U)×(L+U)Z \in \mathbb{R}^{(L+U) \times (L+U)} via low-rank representation (LRR):

    arg⁡min⁡Z∥fθ(X)−fθ(X)Z∥F2s.t.rank(Z)=K\arg\min_Z \|f_\theta(X) - f_\theta(X)Z\|_F^2 \quad \text{s.t.} \quad \text{rank}(Z) = K

    where fθ(X)∈R(L+U)×Of_\theta(X) \in \mathbb{R}^{(L+U) \times O} is the concatenated feature matrix of LL support points and UU query points.

    This rank-constrained problem is solved via the thin/skinny Singular Value Decomposition (SVD) of fθ(X)=UΣV⊤f_\theta(X) = U \Sigma V^\top:

    1. For each row of VV, only the components corresponding to the top-KK largest singular values in Σ\Sigma are retained.
    2. The representation matrix is calculated as Z=V⊤VZ = V^\top V.
    3. The non-negative, zero-diagonal similarity affinity matrix is formed as:

    Wsim=∣Z∣−diag(∣Z∣)W_{sim} = |Z| - \text{diag}(|Z|)

    1. The normalized similarity matrix is A~sim=Dsim−1/2WsimDsim−1/2\tilde{A}_{sim} = D_{sim}^{-1/2} W_{sim} D_{sim}^{-1/2}, where Dsim=diag(d1,…,dL+U)D_{sim} = \text{diag}(d_1, \dots, d_{L+U}) with di=∑j(Wsim)ijd_i = \sum_j (W_{sim})_{ij}.
  3. Knowl 3 — Dense Dissimilarity Matrix Construction and Equivalence to Variance Maximization

    model/method

    In a KK-way few-shot task with NN support and BB query examples per class (total L+U=(N+B)KL+U = (N+B)K samples), off-diagonal pairs can be assumed to represent differing entities in the absence of class labels. The unnormalized dissimilarity matrix Adis∈R(N+B)K×(N+B)KA_{dis} \in \mathbb{R}^{(N+B)K \times (N+B)K} is constructed as the adjacency matrix of a densely connected graph without self-loops:

    Adis=1(N+B)Kee⊤−IA_{dis} = \frac{1}{(N+B)K} e e^\top - I

    where e=[1,1,…,1]⊤∈R(N+B)Ke = [1, 1, \dots, 1]^\top \in \mathbb{R}^{(N+B)K} is the all-ones vector and II is the identity matrix.

    Because each row and column sum of AdisA_{dis} equals (N+B)K(N+B)K, symmetric degree normalization yields A~dis=1(N+B)KAdis\tilde{A}_{dis} = \frac{1}{(N+B)K} A_{dis}.

    Minimizing Wfθ(X)⊤A~disfθ(X)W⊤W f_\theta(X)^\top \tilde{A}_{dis} f_\theta(X) W^\top is equivalent to maximizing the total sample variance in the projected subspace (analogous to Principal Component Analysis, PCA), since:

    fθ(X)⊤(I−1(N+B)Kee⊤)fθ(X)f_\theta(X)^\top \left( I - \frac{1}{(N+B)K} e e^\top \right) f_\theta(X)

    represents the empirical sample covariance matrix of the task features.

  4. Knowl 4 — Constrained Wasserstein Mean Shift Clustering (SIAMESE) Algorithm

    algorithm

    Constrained Wasserstein Mean Shift Clustering (SIAMESE) estimates class prototypes C~=[c~1,…,c~K]⊤∈RK×O′\tilde{C} = [\tilde{c}_1, \dots, \tilde{c}_K]^\top \in \mathbb{R}^{K \times O'} and soft assignment probabilities P∈R(L+U)×KP \in \mathbb{R}^{(L+U) \times K} for projected features H=fθ(X)W⊤∈R(L+U)×O′H = f_\theta(X)W^\top \in \mathbb{R}^{(L+U) \times O'}. It extends Sinkhorn K-means by enforcing hard label clamping P1:L=Y1:LP_{1:L} = Y_{1:L} on the LL labeled support samples, ensuring class distribution balance via optimal transport marginals r=1L+Ur = \mathbf{1}_{L+U} and c=(N+B)1Kc = (N+B)\mathbf{1}_K.

    Input: Projected sample matrix H=[h1,…,hL+U]⊤∈R(L+U)×O′H = [h_1, \dots, h_{L+U}]^\top \in \mathbb{R}^{(L+U) \times O'}, one-hot support labels Y1:L∈{0,1}L×KY_{1:L} \in \{0, 1\}^{L \times K}, support size per class NN, queries per class BB, regularization λ\lambda, learning rate α\alpha, step limit nstepn_{\text{step}}, convergence threshold ϵ\epsilon
    Output: Predicted class labels yi∈{1,…,K}y_i \in \{1, \dots, K\} for query samples i∈{L+1,…,L+U}i \in \{L+1, \dots, L+U\}
    Initialize prototypes c~k=1∣Sk∣∑(hi,yi)∈Skhi\tilde{c}_k = \frac{1}{|S_k|} \sum_{(h_i, y_i) \in S_k} h_i for k=1,…,Kk = 1, \dots, K
    Initialize marginal constraints r=1L+Ur = \mathbf{1}_{L+U}, c=(N+B)1Kc = (N + B) \mathbf{1}_K
    Initialize iteration counter t=0t = 0
    while t<nstept < n_{\text{step}} do
        Compute cost matrix Mi,j=∥hi−c~j∥22M_{i,j} = \|h_i - \tilde{c}_j\|_2^2 for all i∈{1,…,L+U},j∈{1,…,K}i \in \{1, \dots, L+U\}, j \in \{1, \dots, K\}
        Compute initial transport weights P=exp⁡(−λM)P = \exp(-\lambda M)
        Normalize P=P/∑i,jPi,jP = P / \sum_{i,j} P_{i,j}
        Initialize vector u=0L+Uu = \mathbf{0}_{L+U}
        
        while max⁡i∣ui−∑jPi,j∣>ϵ\max_i |u_i - \sum_j P_{i,j}| > \epsilon do
            u=∑jP⋅,ju = \sum_j P_{\cdot, j}
            P=diag(r/u)PP = \text{diag}(r / u) P
            P=Pdiag(c/∑iPi,⋅)P = P \text{diag}(c / \sum_i P_{i, \cdot})
            Clamp support labels P1:L=Y1:LP_{1:L} = Y_{1:L}
        end while
        
        Compute target prototype centers Ω=P⊤H/(N+B)\Omega = P^\top H / (N + B)
        Update prototypes C~←C~+α(Ω−C~)\tilde{C} \leftarrow \tilde{C} + \alpha (\Omega - \tilde{C})
        t←t+1t \leftarrow t + 1
    end while
    return yi=arg⁡max⁡jPi,jy_i = \arg\max_j P_{i,j} for all i∈{L+1,…,L+U}i \in \{L+1, \dots, L+U\}
  5. Knowl 5 — Transductive Few-Shot Classification Performance on mini-ImageNet and tiered-ImageNet

    data/table

    The combination of EASE subspace projection and SIAMESE clustering was evaluated on standard 5-way 1-shot and 5-shot transductive few-shot image classification tasks over 10,000 episodes on mini-ImageNet and tiered-ImageNet benchmarks using ResNet-12 and WRN-28-10 backbones.

    mini-ImageNet tiered-ImageNet
    Method (Backbone) 1-shot (%) 5-shot (%) 1-shot (%) 5-shot (%)
    ResNet-12 Backbone:
    TPN 55.51±0.8655.51 \pm 0.86 69.86±0.6569.86 \pm 0.65 59.91±0.9459.91 \pm 0.94 73.30±0.7573.30 \pm 0.75
    Transductive Tuning 62.35±0.6662.35 \pm 0.66 74.53±0.5474.53 \pm 0.54 – –
    MetaOptNet 62.64±0.6162.64 \pm 0.61 78.63±0.4678.63 \pm 0.46 65.99±0.7265.99 \pm 0.72 81.56±0.5381.56 \pm 0.53
    DSN-MR 64.60±0.7264.60 \pm 0.72 79.51±0.5079.51 \pm 0.50 67.39±0.8267.39 \pm 0.82 82.85±0.5682.85 \pm 0.56
    CAN-T 67.19±0.5567.19 \pm 0.55 80.64±0.3580.64 \pm 0.35 73.21±0.5873.21 \pm 0.58 84.93±0.3884.93 \pm 0.38
    EASE + Soft K-means (ours) 57.00±0.2657.00 \pm 0.26 75.07±0.2175.07 \pm 0.21 69.74±0.3169.74 \pm 0.31 85.17±0.2185.17 \pm 0.21
    QR + SIAMESE (ours) 68.66±0.3768.66 \pm 0.37 79.36±0.2279.36 \pm 0.22 75.87±0.2975.87 \pm 0.29 87.80±0.2187.80 \pm 0.21
    EASE + SIAMESE (ours) 70.47±0.30\mathbf{70.47 \pm 0.30} 80.73±0.16\mathbf{80.73 \pm 0.16} 84.54±0.27\mathbf{84.54 \pm 0.27} 89.63±0.15\mathbf{89.63 \pm 0.15}
    WRN-28-10 Backbone:
    BD-CSPN 70.31±0.9370.31 \pm 0.93 81.89±0.6081.89 \pm 0.60 78.74±0.9578.74 \pm 0.95 86.92±0.6386.92 \pm 0.63
    TIM 77.877.8 87.487.4 82.182.1 89.889.8
    EPNet 70.74±0.8570.74 \pm 0.85 84.34±0.5384.34 \pm 0.53 78.50±0.9178.50 \pm 0.91 88.36±0.5788.36 \pm 0.57
    LaplacianShot 74.86±0.1974.86 \pm 0.19 84.13±0.1484.13 \pm 0.14 80.18±0.2180.18 \pm 0.21 87.56±0.1587.56 \pm 0.15
    iLCT 83.05±0.7983.05 \pm 0.79 88.82±0.4288.82 \pm 0.42 88.50±0.7588.50 \pm 0.75 92.46±0.4292.46 \pm 0.42
    Oblique Manifold 80.64±0.3480.64 \pm 0.34 89.39±0.3989.39 \pm 0.39 85.22±0.3485.22 \pm 0.34 91.35±0.4091.35 \pm 0.40
    EASE + Soft K-means (ours) 67.42±0.2767.42 \pm 0.27 84.45±0.1884.45 \pm 0.18 75.87±0.2975.87 \pm 0.29 85.17±0.2185.17 \pm 0.21
    QR + SIAMESE (ours) 79.90±0.3479.90 \pm 0.34 86.88±0.1986.88 \pm 0.19 84.31±0.3084.31 \pm 0.30 90.55±0.1990.55 \pm 0.19
    EASE + SIAMESE (ours) 83.00±0.21\mathbf{83.00 \pm 0.21} 88.92±0.13\mathbf{88.92 \pm 0.13} 88.96±0.23\mathbf{88.96 \pm 0.23} 92.63±0.13\mathbf{92.63 \pm 0.13}

    The results demonstrate that EASE+SIAMESE consistently outperforms previous transductive state-of-the-art methods across both backbones and datasets, achieving up to 84.54% 1-shot accuracy on tiered-ImageNet with ResNet-12 and 88.96% with WRN-28-10.

  6. Knowl 6 — Transductive Classification Performance on CUB, CIFAR-FS, and OpenMIC

    empirical result

    On fine-grained and cross-domain few-shot benchmarks under the 5-way 1-shot and 5-shot transductive settings:

    1. CUB Benchmark:

      • ResNet-12: EASE+SIAMESE achieves 90.11±0.21%90.11 \pm 0.21\% (1-shot) and 93.13±0.11%93.13 \pm 0.11\% (5-shot), outperforming iLPC (89.00%89.00\% and 92.74%92.74\%) and LR+ICI (86.53%86.53\% and 92.11%92.11\%).
      • WRN-28-10: EASE+SIAMESE achieves 91.68±0.19%91.68 \pm 0.19\% (1-shot) and 94.12±0.09%94.12 \pm 0.09\% (5-shot), outperforming PT+MAP (91.37%91.37\% and 93.93%93.93\%) and iLPC (91.03%91.03\% and 94.11%94.11\%).
    2. CIFAR-FS Benchmark:

      • ResNet-12: EASE+SIAMESE reaches 78.41±0.29%78.41 \pm 0.29\% (1-shot) and 85.67±0.11%85.67 \pm 0.11\% (5-shot), surpassing iLPC (77.14%77.14\%) and DSN-MR (75.60%75.60\%).
      • WRN-28-10: EASE+SIAMESE attains 87.60±0.23%87.60 \pm 0.23\% (1-shot) and 90.60±0.16%90.60 \pm 0.16\% (5-shot), outperforming PT+MAP (86.91%86.91\% and 90.50%90.50\%) and SIB (80.00%80.00\% and 85.30%85.30\%).
    3. OpenMIC Benchmark (5-way 1-shot cross-domain transfers):

      • Across four domain transfer splits (p1→p2p1 \to p2: 81.60%81.60\%, p2→p3p2 \to p3: 68.68%68.68\%, p3→p4p3 \to p4: 86.38%86.38\%, p4→p1p4 \to p1: 66.53%66.53\%), EASE+SIAMESE achieves an average accuracy of 75.80%75.80\%, substantially exceeding inductive baselines including SoSN (67.85%67.85\%) and DSN (69.59%69.59\%).
  7. Knowl 7 — Semi-Supervised Few-Shot Classification with Distractor Classes

    data/table

    In the semi-supervised few-shot learning setting, test episodes are augmented with an additional pool of unlabeled images that contain samples from both target classes and distractor categories. Experiments evaluate 5-way 1-shot and 5-shot tasks with a 30/5030/50 unlabeled pool setting across mini-ImageNet, tiered-ImageNet, CIFAR-FS, and CUB.

    mini-ImageNet tiered-ImageNet CIFAR-FS CUB
    Method (Backbone) 1-shot 5-shot 1-shot 5-shot 1-shot 5-shot 1-shot 5-shot
    ResNet-12:
    LR+ICI 67.5767.57 79.0779.07 83.3283.32 89.0689.06 75.9975.99 84.0184.01 88.5088.50 –
    iLPC 70.9970.99 81.0681.06 85.0485.04 89.6389.63 78.5778.57 85.8485.84 90.1190.11 –
    EASE+SIAMESE 73.90\mathbf{73.90} 81.68\mathbf{81.68} 85.86\mathbf{85.86} 89.64\mathbf{89.64} 80.51\mathbf{80.51} 85.96\mathbf{85.96} 90.61\mathbf{90.61} –
    WRN-28-10:
    LR+ICI 81.3181.31 88.5388.53 88.4888.48 92.0392.03 86.0386.03 89.5789.57 90.8290.82 –
    PT+MAP 83.1483.14 88.9588.95 89.1689.16 92.3092.30 87.0587.05 89.9889.98 91.5291.52 –
    iLPC 83.5883.58 89.6889.68 89.3589.35 92.6192.61 87.0387.03 90.3490.34 91.6991.69 –
    EASE+SIAMESE 84.89\mathbf{84.89} 89.47\mathbf{89.47} 90.08\mathbf{90.08} 92.67\mathbf{92.67} 87.89\mathbf{87.89} 90.18\mathbf{90.18} 92.11\mathbf{92.11} –

    EASE+SIAMESE achieves performance gains of 3%3\% to 6%6\% over competing approaches on 1-shot mini-ImageNet (ResNet-12), demonstrating robustness against unrelated distractor samples in the unlabeled set.

  8. Knowl 8 — Comparison of EASE with PCA and ICA Dimensionality Reduction on DenseNet Backbone

    data/table

    When evaluated on a standard DenseNet backbone trained directly with cross-entropy classification on base classes (without episodic meta-learning), replacing PCA or ICA in the task-adaptive feature subspace learning pipeline (TAFSSL) with EASE leads to superior accuracy across mini-ImageNet and tiered-ImageNet.

    mini-ImageNet tiered-ImageNet
    Method 1-shot (%) 5-shot (%) 1-shot (%) 5-shot (%)
    SimpleShot (Inductive) 65.77±0.1965.77 \pm 0.19 82.23±0.1382.23 \pm 0.13 71.20±0.2271.20 \pm 0.22 86.33±0.1586.33 \pm 0.15
    LaplacianShot 75.57±0.1975.57 \pm 0.19 84.72±0.1384.72 \pm 0.13 80.30±0.2080.30 \pm 0.20 87.93±0.1587.93 \pm 0.15
    RAP-LaplacianShot 75.58±0.2075.58 \pm 0.20 85.63±0.1385.63 \pm 0.13 – –
    TAFSSL (PCA) 70.53±0.2570.53 \pm 0.25 80.71±0.1680.71 \pm 0.16 80.07±0.2580.07 \pm 0.25 86.42±0.1786.42 \pm 0.17
    TAFSSL (ICA) 72.10±0.2572.10 \pm 0.25 81.85±0.1681.85 \pm 0.16 80.82±0.2580.82 \pm 0.25 86.97±0.1786.97 \pm 0.17
    EASE + Soft K-means 74.30±0.2674.30 \pm 0.26 82.08±0.1782.08 \pm 0.17 82.67±0.2582.67 \pm 0.25 87.60±0.1787.60 \pm 0.17
    QR + SIAMESE 75.75±0.3275.75 \pm 0.32 85.10±0.1885.10 \pm 0.18 82.87±0.3382.87 \pm 0.33 89.11±0.2089.11 \pm 0.20
    EASE + SIAMESE (ours) 79.42±0.27\mathbf{79.42 \pm 0.27} 86.76±0.14\mathbf{86.76 \pm 0.14} 86.17±0.25\mathbf{86.17 \pm 0.25} 90.54±0.15\mathbf{90.54 \pm 0.15}

    Compared to SimpleShot, EASE+SIAMESE improves performance by 13.65%13.65\% on 1-shot mini-ImageNet and 14.97%14.97\% on 1-shot tiered-ImageNet. It also surpasses TAFSSL(ICA) by 7.32%7.32\% and 5.35%5.35\% in 1-shot settings, indicating that EASE constructs a significantly more discriminative subspace than classical unsupervised projection techniques.

  9. Knowl 9 — Ablation on Query Set Size, Subspace Dimension, and Inference Latency

    empirical result

    Ablation studies on mini-ImageNet reveal three key operational properties of EASE and SIAMESE:

    1. Query Set Scaling: Varying query count per class from 2 to 50 shows that EASE maintains effective subspace learning starting from 5 queries. While pure SVD-based adaptation plateaus quickly as queries increase, EASE accuracy increases approximately linearly with query volume.
    2. Subspace Dimension Robustness: Classification performance rises steeply when the projection dimension O′O' increases from 5 to 20, after which performance gains stabilize slowly. Setting O′=40O' = 40 provides an optimal balance between accuracy and computational cost across all benchmarks. Unlike SVD, whose performance degrades beyond 20 dimensions, EASE maintains stable performance as dimension scales toward sample count.
    3. Inference Latency: On mini-ImageNet (measured on an AMD 2700 CPU), EASE+SIAMESE requires 6.0 ms6.0\text{ ms} (1-shot) and 7.4 ms7.4\text{ ms} (5-shot) per episode with ResNet-12, and 8.0 ms8.0\text{ ms} (1-shot) and 8.7 ms8.7\text{ ms} (5-shot) with WRN-28-10. This is approximately 10×10\times faster than iterative transductive alternatives (iLPC at 45–70 ms45\text{--}70\text{ ms}, ICI at 34–52 ms34\text{--}52\text{ ms}).

Coverage note — None was omitted; all primary methodological contributions (EASE subspace learning, LRR similarity, dense dissimilarity, SIAMESE clustering) and empirical evaluations (transductive FSL, semi-supervised FSL, ablations, latency) are covered.

References

  1. 1.Malik Boudiaf, Imtiaz Ziko, Jertome Rony, Jose Dolz, Pablo Piantanida, and Ismail Ben Ayed. Information maximization for few-shot learning. Advances in Neural Information Processing Systems, 33, 2020.
  2. 2.Wei-Yu Chen, Yen-Cheng Liu, Zsolt Kira, Yu-Chiang Frank Wang, and Jia-Bin Huang. A closer look at few-shot classification. arXiv preprint arXiv:1904.04232, 2019.
  3. 3.Marco Cuturi. Sinkhorn distances: Lightspeed computation of optimal transport. Advances in neural information processing systems, 26:2292–2300, 2013.
  4. 4.Guneet S Dhillon, Pratik Chaudhari, Avinash Ravichandran, and Stefano Soatto. A baseline for few-shot image classification. arXiv preprint arXiv:1909.02729, 2019.
  5. 5.Li Fei-Fei, Rob Fergus, and Pietro Perona. One-shot learning of object categories. IEEE transactions on pattern analysis and machine intelligence, 28(4):594–611, 2006.
  6. 6.Chelsea Finn, Pieter Abbeel, and Sergey Levine. Model-agnostic meta-learning for fast adaptation of deep networks. In International Conference on Machine Learning, pages 1126–1135. PMLR, 2017.
  7. 7.Spyros Gidaris and Nikos Komodakis. Generating classification weights with gnn denoising autoencoders for few-shot learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 21–30, 2019.
  8. 8.Jie Hong, Pengfei Fang, Weihao Li, Tong Zhang, Christian Simon, Mehrtash Harandi, and Lars Petersson. Reinforced attention for few-shot learning and beyond. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 913–923, 2021.
  9. 9.Ruibing Hou, Hong Chang, Bingpeng Ma, Shiguang Shan, and Xilin Chen. Cross attention network for few-shot classification. arXiv preprint arXiv:1910.07677, 2019.
  10. 10.Shell Xu Hu, Pablo G Moreno, Yang Xiao, Xi Shen, Guillaume Obozinski, Neil D Lawrence, and Andreas Damianou. Empirical bayes transductive meta-learning with synthetic gradients. arXiv preprint arXiv:2004.12696, 2020.
  11. 11.Yuqing Hu, Vincent Gripon, and Stephane Pateux. Leveraging the feature distribution in transfer-based few-shot learning. arXiv preprint arXiv:2006.03806, 2020.
  12. 12.Gabriel Huang, Hugo Larochelle, and Simon Lacoste-Julien. Are few-shot learning benchmarks too simple? solving them without task supervision at test-time. arXiv preprint arXiv:1902.08605, 2019.
  13. 13.Gao Huang, Zhuang Liu, Laurens Van Der Maaten, and Kilian Q Weinberger. Densely connected convolutional networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 4700–4708, 2017.
  14. 14.Huaxi Huang, Junjie Zhang, Jian Zhang, Qiang Wu, and Chang Xu. Ptn: A poisson transfer network for semi-supervised few-shot learning. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, pages 1602–1609, 2021.
  15. 15.Jongmin Kim, Taesup Kim, Sungwoong Kim, and Chang D Yoo. Edge-labeling graph neural network for few-shot learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 11–20, 2019.
  16. 16.Piotr Koniusz, Yusuf Tas, Hongguang Zhang, Mehrtash Harandi, Fatih Porikli, and Rui Zhang. Museum exhibit identification challenge for the supervised domain adaptation and beyond. In The European Conference on Computer Vision (ECCV), September 2018.
  17. 17.Piotr Koniusz and Hongguang Zhang. Power normalizations in fine-grained image, few-shot image and graph classification. In IEEE Transactions on Pattern Analysis and Machine Intelligence, 2020.
  18. 18.Alex Krizhevsky, Geoffrey Hinton, et al. Learning multiple layers of features from tiny images. 2009.
  19. 19.Seong Min Kye, Hae Beom Lee, Hoirin Kim, and Sung Ju Hwang. Meta-learned confidence for few-shot learning. arXiv preprint arXiv:2002.12017, 2020.
  20. 20.Michalis Lazarou, Tania Stathaki, and Yannis Avrithis. Iterative label cleaning for transductive and semi-supervised few-shot learning. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 8751–8760, 2021.
  21. 21.Kwonjoon Lee, Subhransu Maji, Avinash Ravichandran, and Stefano Soatto. Meta-learning with differentiable convex optimization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10657–10665, 2019.
  22. 22.Xinzhe Li, Qianru Sun, Yaoyao Liu, Shibao Zheng, Qin Zhou, Tat-Seng Chua, and Bernt Schiele. Learning to self-train for semi-supervised few-shot classification (6 2019). In Advances in Neural Information Processing Systems: 33rd Conference on Neural Information Processing Systems (NeurIPS 2019), Vancouver, Canada, December, volume 8, pages 1–11, 1906.
  23. 23.Moshe Lichtenstein, Prasanna Sattigeri, Rogerio Feris, Raja Giryes, and Leonid Karlinsky. Tafssl: Task-adaptive feature sub-space learning for few-shot classification. In European Conference on Computer Vision, pages 522–539. Springer, 2020.
  24. 24.Guangcan Liu, Zhouchen Lin, and Yong Yu. Robust subspace segmentation by low-rank representation. In Proceedings of the 27th International Conference on International Conference on Machine Learning, pages 663–670, 2010.
  25. 25.Jinlu Liu, Liang Song, and Yongqiang Qin. Prototype rectification for few-shot learning. arXiv preprint arXiv:1911.10713, 2019.
  26. 26.Yanbin Liu, Juho Lee, Minseop Park, Saehoon Kim, Eunho Yang, Sung Ju Hwang, and Yi Yang. Learning to propagate labels: Transductive propagation network for few-shot learning. arXiv preprint arXiv:1805.10002, 2018.
  27. 27.Puneet Mangla, Nupur Kumari, Abhishek Sinha, Mayank Singh, Balaji Krishnamurthy, and Vineeth N Balasubramanian. Charting the right manifold: Manifold mixup for few-shot learning. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 2218–2227, 2020.
  28. 28.Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S Corrado, and Jeff Dean. Distributed representations of words and phrases and their compositionality. In Advances in neural information processing systems, pages 3111–3119, 2013.
  29. 29.Boris N Oreshkin, Pau Rodriguez, and Alexandre Lacoste. Tadam: Task dependent adaptive metric for improved few-shot learning. arXiv preprint arXiv:1805.10123, 2018.
  30. 30.Guodong Qi, Huimin Yu, Zhaohui Lu, and Shuzhao Li. Transductive few-shot classification on the oblique manifold. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 8412–8422, 2021.
  31. 31.Limeng Qiao, Yemin Shi, Jia Li, Yaowei Wang, Tiejun Huang, and Yonghong Tian. Transductive episodic-wise adaptive metric for few-shot learning. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 3603–3612, 2019.
  32. 32.Sachin Ravi and Hugo Larochelle. Optimization as a model for few-shot learning. 2017.
  33. 33.Mengye Ren, Eleni Triantafillou, Sachin Ravi, Jake Snell, Kevin Swersky, Joshua B Tenenbaum, Hugo Larochelle, and Richard S Zemel. Meta-learning for semi-supervised few-shot classification. arXiv preprint arXiv:1803.00676, 2018.
  34. 34.Pau Rodríguez, Issam Laradji, Alexandre Drouin, and Alexandre Lacoste. Embedding propagation: Smoother manifold for few-shot classification. In European Conference on Computer Vision, pages 121–138. Springer, 2020.
  35. 35.Andrei A Rusu, Dushyant Rao, Jakub Sygnowski, Oriol Vinyals, Razvan Pascanu, Simon Osindero, and Raia Hadsell. Meta-learning with latent embedding optimization. arXiv preprint arXiv:1807.05960, 2018.
  36. 36.Kuniaki Saito, Donghyun Kim, Stan Sclaroff, Trevor Darrell, and Kate Saenko. Semi-supervised domain adaptation via minimax entropy. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 8050–8058, 2019.
  37. 37.Victor Garcia Satorras and Joan Bruna Estrach. Few-shot learning with graph neural networks. In International Conference on Learning Representations, 2018.
  38. 38.Christian Simon, Piotr Koniusz, and Mehrtash Harandi. Meta-learning for multi-label few-shot classification. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), pages 3951–3960, 2022.
  39. 39.Christian Simon, Piotr Koniusz, Richard Nock, and Mehrtash Harandi. Adaptive subspaces for few-shot learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4136–4145, 2020.
  40. 40.Jake Snell, Kevin Swersky, and Richard S Zemel. Prototypical networks for few-shot learning. arXiv preprint arXiv:1703.05175, 2017.
  41. 41.Flood Sung, Yongxin Yang, Li Zhang, Tao Xiang, Philip HS Torr, and Timothy M Hospedales. Learning to compare: Relation network for few-shot learning. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 1199–1208, 2018.
  42. 42.Lloyd N Trefethen and David Bau III. Numerical linear algebra, volume 50. Siam, 1997.
  43. 43.Oriol Vinyals, Charles Blundell, Timothy Lillicrap, Koray Kavukcuoglu, and Daan Wierstra. Matching networks for one shot learning. arXiv preprint arXiv:1606.04080, 2016.
  44. 44.Yan Wang, Wei-Lun Chao, Kilian Q Weinberger, and Laurens van der Maaten. Simpleshot: Revisiting nearest-neighbor classification for few-shot learning. arXiv preprint arXiv:1911.04623, 2019.
  45. 45.Yikai Wang, Chengming Xu, Chen Liu, Li Zhang, and Yanwei Fu. Instance credibility inference for few-shot learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12836–12845, 2020.
  46. 46.Peter Welinder, Steve Branson, Takeshi Mita, Catherine Wah, Florian Schroff, Serge Belongie, and Pietro Perona. Caltech-ucsd birds 200. 2010.
  47. 47.Yaochen Xie, Zhao Xu, Zhengyang Wang, and Shuiwang Ji. Self-supervised learning of graph neural networks: A unified review. CoRR, abs/2102.10757, 2021.
  48. 48.Weijian Xu, yifan xu, Huaijin Wang, and Zhuowen Tu. Attentional constellation nets for few-shot learning. In International Conference on Learning Representations, 2021.
  49. 49.Han-Jia Ye, Hexiang Hu, De-Chuan Zhan, and Fei Sha. Few-shot learning via embedding adaptation with set-to-set functions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8808–8817, 2020.
  50. 50.Chi Zhang, Yujun Cai, Guosheng Lin, and Chunhua Shen. Deepemd: Few-shot image classification with differentiable earth mover’s distance and structured classifiers. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 12203–12213, 2020.
  51. 51.Hongguang Zhang and Piotr Koniusz. Power normalizing second-order similarity network for few-shot learning. In Winter Conference on Applications of Computer Vision (WACV), January 2019.
  52. 52.Hongguang Zhang, Piotr Koniusz, Songlei Jian, Hongdong Li, and Philip H. S. Torr. Rethinking class relations: Absolute-relative supervised and unsupervised few-shot learning. In IEEE Conference on Computer Vision and Pattern Recognition, pages 9432–9441, 2021.
  53. 53.Hongguang Zhang, Hongdong Li, and Piotr Koniusz. Multi-level second-order few-shot learning. IEEE Transactions on Multimedia, 2022.
  54. 54.Hongguang Zhang, Jing Zhang, and Piotr Koniusz. Few-shot learning via saliency-guided hallucination of samples. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 2770–2779, 2019.
  55. 55.Shan Zhang, Dawei Luo, Lei Wang, and Piotr Koniusz. Few-shot object detection by second-order pooling. In Asian Conference on Computer Vision, 2020.
  56. 56.Hao Zhu and Piotr Koniusz. Refine: Random range finder for network embedding. In Proceedings of the 30th ACM International Conference on Information & Knowledge Management, pages 3682–3686, 2021.
  57. 57.Hao Zhu, Ke Sun, and Piotr Koniusz. Contrastive laplacian eigenmaps. Advances in Neural Information Processing Systems, 34, 2021.
  58. 58.Imtiaz Ziko, Jose Dolz, Eric Granger, and Ismail Ben Ayed. Laplacian regularized few-shot learning. In International Conference on Machine Learning, pages 11660–11670. PMLR, 2020.

Citation

MLA
Zhu, H., and P. Koniusz. “EASE: Unsupervised Discriminant Subspace Learning for Transductive Few-Shot Learning”. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022, pp. 9068–78, https://doi.org/10.1109/CVPR52688.2022.00887.
APA
Zhu, H., & Koniusz, P. (2022). EASE: Unsupervised Discriminant Subspace Learning for Transductive Few-Shot Learning. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 9068–9078. https://doi.org/10.1109/CVPR52688.2022.00887
Chicago
Zhu, H., and P. Koniusz. 2022. “EASE: Unsupervised Discriminant Subspace Learning for Transductive Few-Shot Learning”. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 9068–78. https://doi.org/10.1109/CVPR52688.2022.00887.
Harvard
Zhu, H. and Koniusz, P. (2022) “EASE: Unsupervised Discriminant Subspace Learning for Transductive Few-Shot Learning”, 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp. 9068–9078. Available at: https://doi.org/10.1109/CVPR52688.2022.00887.
Vancouver
1. Zhu H, Koniusz P (2022) EASE: Unsupervised Discriminant Subspace Learning for Transductive Few-Shot Learning. In: 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp 9068–9078

BibTeX

@inproceedings{Zhu_2022, title={EASE: Unsupervised Discriminant Subspace Learning for Transductive Few-Shot Learning}, url={http://dx.doi.org/10.1109/CVPR52688.2022.00887}, DOI={10.1109/cvpr52688.2022.00887}, booktitle={2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)}, publisher={IEEE}, author={Zhu, Hao and Koniusz, Piotr}, year={2022}, month=June, pages={9068–9078} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE