GKEAL: Gaussian Kernel Embedded Analytic Learning for Few-Shot Class Incremental Task

Huiping ZhuangZhenyu WengRun HeZhiping LinZiqian Zeng

article2023CVPR118 citations

Proposes a closed-form analytic learning framework for few-shot class-incremental learning that combines Gaussian kernel embeddings with recursive least-squares updates to guarantee weight-invariant classifier solutions and prevent catastrophic forgetting.

Listen

Deploying computer vision systems in dynamic real-world environments requires machine learning models to continually absorb new categories over time without losing previously acquired knowledge. In practice, new classes frequently arrive with very limited training examples—a challenging setup termed few-shot class-incremental learning. Standard deep learning models suffer severely from catastrophic forgetting, rapidly overwriting historical base knowledge, or heavily over-fitting to scarce new data.

The article demonstrates and evaluates a novel framework called Gaussian Kernel Embedded Analytic Learning (GKEAL), designed to achieve high classification accuracy across incremental learning stages while mathematically preventing catastrophic forgetting.

The evaluated approach converts neural network classifier updates into direct, closed-form linear algebra operations rather than relying on iterative gradient updates. After initial feature extractor training on base classes, the backbone network is frozen, and the final classification layer is replaced by a Kernel Analytic Module. This module uses Gaussian kernel embeddings to make features more discriminative and calculates parameters recursively via least-squares solutions. To correct for the extreme sample size disparity between well-represented base classes and data-scarce new classes, the framework incorporates an Augmented Feature Concatenation module that mathematically scales up the impact of new classes in a single-shot calculation.

Empirical evaluations on standard benchmark image datasets demonstrate significant performance gains. First, GKEAL achieved the highest final-phase classification accuracy across all tested benchmarks, reaching 51.31% on mini-ImageNet, 51.40% on CIFAR-100, and 58.67% on CUB200-2011, consistently outperforming existing methods. Second, it demonstrated the lowest performance drop rates from the initial base phase to the final incremental phase (dropping only 20.21% to 22.61%), confirming effective mitigation of catastrophic forgetting. Third, ablation studies proved that both core components are critical: removing the Gaussian kernel embedding caused severe mathematical breakdown, dropping final accuracy on mini-ImageNet from 51.21% to 7.22%, while adding feature concatenation improved accuracy by 6.42 percentage points by preventing base-class bias.

These results establish that combining recursive closed-form analytic learning with kernel transformations provides a reliable alternative to traditional iterative retraining for incremental tasks. By eliminating iterative back-propagation during incremental phases, the approach reduces the computational instability, tuning complexity, and performance degradation typically caused by few-shot data streams.

For practical implementation, organizations deploying incremental image classification systems under data constraints should consider adopting recursive analytic classifier architectures. Operating teams must carefully tune the kernel count, kernel width, and data augmentation scaling parameters based on the specific ratio of base to incremental data. Future work should focus on developing more compact kernel structures to reduce memory overhead and testing the approach across larger-scale real-world continuous learning deployments.

The primary limitation of this method is the requirement for a large number of kernel centers (ranging from 5,000 to 12,000 in the tests) to prevent under-fitting, which increases parameter storage requirements. Additionally, the approach relies on the assumption that a frozen backbone trained on base classes provides sufficient feature representations for future unseen classes. Confidence in the reported results is high, supported by multi-run evaluations across diverse standard benchmarks.

Cover for GKEAL: Gaussian Kernel Embedded Analytic Learning for Few-Shot Class Incremental Task

Abstract

Few-shot class incremental learning (FSCIL) aims to address catastrophic forgetting during class incremental learning in a few-shot learning setting. In this paper, we approach the FSCIL by adopting analytic learning, a technique that converts network training into linear problems. This is inspired by the fact that the recursive implementation (batch-by-batch learning) of analytic learning gives identical weights to that produced by training on the entire dataset at once. The recursive implementation and the weight-identical property highly resemble the FSCIL setting (phase-by-phase learning) and its goal of avoiding catastrophic forgetting. By bridging the FSCIL with the analytic learning, we propose a Gaussian kernel embedded analytic learning (GKEAL) for FSCIL. The key components of GKEAL include the kernel analytic module which allows the GKEAL to conduct FSCIL in a recursive manner, and the augmented feature concatenation module that balances the preference between old and new tasks especially effectively under the few-shot setting. Our experiments show that the GKEAL gives state-of-the-art performance on several benchmark datasets.

Table of Contents

  • 1. Introduction
  • 2. Related Works
  • 3. The Proposed Method
  • 3.1. Base Training
  • 3.2. Few-shot Class Incremental Learning
  • 3.3. Augmented Feature Concatenation For New Tasks
  • 4. Experiments
  • 4.1. Experimental Setup
  • 4.2. Comparison with State-of-the-arts
  • 4.3. Ablation Study and Parameter Analysis
  • 5. Conclusion
  • References

Knowls

  1. Knowl 1 — Recursive Weight Update and Weight-Invariant Property in GKEAL

    theoretical result

    In Gaussian Kernel Embedded Analytic Learning (GKEAL) for few-shot class incremental learning (FSCIL), the ridge regression classification weight matrix can be updated recursively at each incremental phase kk without storing or re-accessing training data from prior phases {0,…,k−1}\{0, \dots, k-1\}, while producing the exact weight matrix that would result from joint ridge regression on all accumulated training data D0:ktrain\mathcal{D}_{0:k}^{\text{train}}.

    Let X0:k−1(ke)∈R(∑j=0k−1Nj)×IX_{0:k-1}^{(\text{ke})} \in \mathbb{R}^{(\sum_{j=0}^{k-1} N_j) \times I} denote the kernelized feature matrix for all training samples up to phase k−1k-1, and let Y0:k−1trainY_{0:k-1}^{\text{train}} denote the corresponding block-diagonal one-hot label matrix. The ridge regression problem over phases 0,…,k−10, \dots, k-1 with regularization parameter γ>0\gamma > 0 is: W^FCN(k−1)=arg⁡min⁡W∥Y0:k−1train−X0:k−1(ke)W∥F2+γ∥W∥F2=(X0:k−1(ke)⊤X0:k−1(ke)+γI)−1X0:k−1(ke)⊤Y0:k−1train\hat{W}_{\text{FCN}}^{(k-1)} = \arg\min_{W} \|Y_{0:k-1}^{\text{train}} - X_{0:k-1}^{(\text{ke})} W\|_F^2 + \gamma \|W\|_F^2 = (X_{0:k-1}^{(\text{ke})\top} X_{0:k-1}^{(\text{ke})} + \gamma I)^{-1} X_{0:k-1}^{(\text{ke})\top} Y_{0:k-1}^{\text{train}}

    Define the recursive inverse matrix Rk−1=(X0:k−1(ke)⊤X0:k−1(ke)+γI)−1∈RI×IR_{k-1} = (X_{0:k-1}^{(\text{ke})\top} X_{0:k-1}^{(\text{ke})} + \gamma I)^{-1} \in \mathbb{R}^{I \times I}. Given new phase-kk training samples with kernelized embeddings Xk(ke)∈RNk×IX_k^{(\text{ke})} \in \mathbb{R}^{N_k \times I} and one-hot labels Yktrain∈RNk×dykY_k^{\text{train}} \in \mathbb{R}^{N_k \times d_{y_k}}, the updated correlation inverse RkR_k and updated classifier weight matrix W^FCN(k)\hat{W}_{\text{FCN}}^{(k)} are computed recursively by: Rk=Rk−1−Rk−1Xk(ke)⊤(INk+Xk(ke)Rk−1Xk(ke)⊤)−1Xk(ke)Rk−1R_k = R_{k-1} - R_{k-1} X_k^{(\text{ke})\top} (I_{N_k} + X_k^{(\text{ke})} R_{k-1} X_k^{(\text{ke})\top})^{-1} X_k^{(\text{ke})} R_{k-1} W^FCN(k)=[W^FCN(k−1)−RkXk(ke)⊤Xk(ke)W^FCN(k−1)RkXk(ke)⊤Yktrain]\hat{W}_{\text{FCN}}^{(k)} = \begin{bmatrix} \hat{W}_{\text{FCN}}^{(k-1)} - R_k X_k^{(\text{ke})\top} X_k^{(\text{ke})} \hat{W}_{\text{FCN}}^{(k-1)} & R_k X_k^{(\text{ke})\top} Y_k^{\text{train}} \end{bmatrix} where the left sub-matrix corresponds to updating the classifier output channels for old classes and the right sub-matrix represents the newly created output channels for the phase-kk classes.

  2. Knowl 2 — Kernel Analytic Module and Analytic Initialization

    model/method

    The Kernel Analytic Module (KAM) is a two-layer non-iterative classifier attached to a frozen convolutional neural network (CNN) backbone fCNN(⋅,WCNN)f_{\text{CNN}}(\cdot, W_{\text{CNN}}) that was pre-trained via back-propagation on the base dataset D0train={X0train,Y0train}\mathcal{D}_0^{\text{train}} = \{X_0^{\text{train}}, Y_0^{\text{train}}\}. The CNN backbone parameters WCNNW_{\text{CNN}} are frozen throughout all subsequent incremental learning phases.

    For base dataset input X0train∈RN0×w×h×3X_0^{\text{train}} \in \mathbb{R}^{N_0 \times w \times h \times 3}, base CNN feature representations X0(cnn)=fflat(fCNN(X0train,WCNN))∈RN0×dcnnX_0^{(\text{cnn})} = f_{\text{flat}}(f_{\text{CNN}}(X_0^{\text{train}}, W_{\text{CNN}})) \in \mathbb{R}^{N_0 \times d_{\text{cnn}}} are extracted, where fflatf_{\text{flat}} flattens tensors into 1-D vectors. The first layer of KAM performs Gaussian Kernel Embedding (GKE) g{c1,…,cI}:RN0×dcnn→RN0×Ig_{\{c_1, \dots, c_I\}}: \mathbb{R}^{N_0 \times d_{\text{cnn}}} \to \mathbb{R}^{N_0 \times I}, where {c1,…,cI}⊂Rdcnn\{c_1, \dots, c_I\} \subset \mathbb{R}^{d_{\text{cnn}}} is a set of II center vectors randomly sampled without replacement from the rows of X0(cnn)X_0^{(\text{cnn})}. The jj-th row of the kernelized embedding matrix X0(ke)∈RN0×IX_0^{(\text{ke})} \in \mathbb{R}^{N_0 \times I} is defined as: X0(ke)[j,:]=[e−β∥X0(cnn)[j,:]−c1∥22e−β∥X0(cnn)[j,:]−c2∥22…e−β∥X0(cnn)[j,:]−cI∥22]X_0^{(\text{ke})}[j, :] = \begin{bmatrix} e^{-\beta \|X_0^{(\text{cnn})}[j,:] - c_1\|_2^2} & e^{-\beta \|X_0^{(\text{cnn})}[j,:] - c_2\|_2^2} & \dots & e^{-\beta \|X_0^{(\text{cnn})}[j,:] - c_I\|_2^2} \end{bmatrix} where β>0\beta > 0 is a kernel width hyperparameter.

    The second layer of KAM is a linear classifier solved analytically via regularized least squares on base task labels Y0train∈RN0×dy0Y_0^{\text{train}} \in \mathbb{R}^{N_0 \times d_{y_0}} (Analytic Initialization, AInit): W^FCN(0)=(X0(ke)⊤X0(ke)+γII)−1X0(ke)⊤Y0train\hat{W}_{\text{FCN}}^{(0)} = (X_0^{(\text{ke})\top} X_0^{(\text{ke})} + \gamma I_I)^{-1} X_0^{(\text{ke})\top} Y_0^{\text{train}} R0=(X0(ke)⊤X0(ke)+γII)−1R_0 = (X_0^{(\text{ke})\top} X_0^{(\text{ke})} + \gamma I_I)^{-1} where γ>0\gamma > 0 is the regularization parameter and III_I is the I×II \times I identity matrix.

  3. Knowl 3 — Augmented Feature Concatenation for Base-New Class Balancing

    model/method

    In few-shot class incremental learning, base classes contain abundant samples while incremental phases contain very few samples per class (e.g., 5 samples per class), causing models to bias heavily toward base classes. The Augmented Feature Concatenation (AFC) module balances the preference between old and new tasks by augmenting and vertically stacking new-task features and labels in single-shot matrix form.

    For phase-kk training data Xktrain∈RNk×w×h×3X_k^{\text{train}} \in \mathbb{R}^{N_k \times w \times h \times 3}, C∈Z+C \in \mathbb{Z}^+ stochastic data augmentations A1,A2,…,ACA_1, A_2, \dots, A_C (such as random horizontal flips, random cropping, and normalization) are applied. The augmented kernelized feature matrix Xˉk(fe)∈RCNk×I\bar{X}_k^{(\text{fe})} \in \mathbb{R}^{C N_k \times I} and expanded label matrix Yˉktrain∈RCNk×dyk\bar{Y}_k^{\text{train}} \in \mathbb{R}^{C N_k \times d_{y_k}} are formed by vertical concatenation: Xˉk(fe)=[g{c1,…,cI}(fCNN(A1(Xktrain),WCNN))g{c1,…,cI}(fCNN(A2(Xktrain),WCNN))⋮g{c1,…,cI}(fCNN(AC(Xktrain),WCNN))],Yˉktrain=[YktrainYktrain⋮Yktrain]\bar{X}_k^{(\text{fe})} = \begin{bmatrix} g_{\{c_1,\dots,c_I\}}(f_{\text{CNN}}(A_1(X_k^{\text{train}}), W_{\text{CNN}})) \\ g_{\{c_1,\dots,c_I\}}(f_{\text{CNN}}(A_2(X_k^{\text{train}}), W_{\text{CNN}})) \\ \vdots \\ g_{\{c_1,\dots,c_I\}}(f_{\text{CNN}}(A_C(X_k^{\text{train}}), W_{\text{CNN}})) \end{bmatrix}, \quad \bar{Y}_k^{\text{train}} = \begin{bmatrix} Y_k^{\text{train}} \\ Y_k^{\text{train}} \\ \vdots \\ Y_k^{\text{train}} \end{bmatrix}

    Substituting Xˉk(fe)\bar{X}_k^{(\text{fe})} and Yˉktrain\bar{Y}_k^{\text{train}} into the recursive update formula yields: W^FCN(k)=[W^FCN(k−1)−RkXˉk(fe)⊤Xˉk(fe)W^FCN(k−1)RkXˉk(fe)⊤Yˉktrain]\hat{W}_{\text{FCN}}^{(k)} = \begin{bmatrix} \hat{W}_{\text{FCN}}^{(k-1)} - R_k \bar{X}_k^{(\text{fe})\top} \bar{X}_k^{(\text{fe})} \hat{W}_{\text{FCN}}^{(k-1)} & R_k \bar{X}_k^{(\text{fe})\top} \bar{Y}_k^{\text{train}} \end{bmatrix}

    Because Xˉk(fe)⊤Xˉk(fe)≈CXk(ke)⊤Xk(ke)\bar{X}_k^{(\text{fe})\top} \bar{X}_k^{(\text{fe})} \approx C X_k^{(\text{ke})\top} X_k^{(\text{ke})} and Xˉk(fe)⊤Yˉktrain≈CXk(ke)⊤Yktrain\bar{X}_k^{(\text{fe})\top} \bar{Y}_k^{\text{train}} \approx C X_k^{(\text{ke})\top} Y_k^{\text{train}}, this operation approximates: W^FCN(k)≈[W^FCN(k−1)−CRkXk(ke)⊤Xk(ke)W^FCN(k−1)CRkXk(ke)⊤Yktrain]\hat{W}_{\text{FCN}}^{(k)} \approx \begin{bmatrix} \hat{W}_{\text{FCN}}^{(k-1)} - C R_k X_k^{(\text{ke})\top} X_k^{(\text{ke})} \hat{W}_{\text{FCN}}^{(k-1)} & C R_k X_k^{(\text{ke})\top} Y_k^{\text{train}} \end{bmatrix} Increasing the augmentation count CC directly scales up the output channel gains for new tasks by approximately a factor of CC while attenuating the gains of old tasks, directly controlling the stability-plasticity trade-off.

  4. Knowl 4 — GKEAL Algorithm for Few-Shot Class Incremental Learning

    algorithm

    The Gaussian Kernel Embedded Analytic Learning (GKEAL) algorithm executes base training via backpropagation followed by non-iterative, recursive incremental analytic learning steps across KK few-shot phases.

    Input: Base dataset D0train\mathcal{D}_0^{\text{train}}, incremental datasets D1:Ktrain\mathcal{D}_{1:K}^{\text{train}}, number of kernel center vectors II, regularization parameter γ\gamma, kernel width parameter β\beta, augmentation count CC
    Output: Sequence of classifier weights W^FCN(0),W^FCN(1),…,W^FCN(K)\hat{W}_{\text{FCN}}^{(0)}, \hat{W}_{\text{FCN}}^{(1)}, \dots, \hat{W}_{\text{FCN}}^{(K)}
    1. Train CNN backbone fCNN(⋅,WCNN)f_{\text{CNN}}(\cdot, W_{\text{CNN}}) and initial classifier using backpropagation on base dataset D0train\mathcal{D}_0^{\text{train}}
    2. Freeze CNN backbone weights WCNNW_{\text{CNN}}
    3. Extract base CNN embeddings X0(cnn)=fflat(fCNN(X0train,WCNN))X_0^{(\text{cnn})} = f_{\text{flat}}(f_{\text{CNN}}(X_0^{\text{train}}, W_{\text{CNN}}))
    4. Randomly sample II center vectors {c1,…,cI}\{c_1, \dots, c_I\} from the rows of X0(cnn)X_0^{(\text{cnn})}
    5. Compute Gaussian kernel base embeddings X0(ke)=g{c1,…,cI}(X0(cnn))X_0^{(\text{ke})} = g_{\{c_1, \dots, c_I\}}(X_0^{(\text{cnn})}) with width β\beta
    6. Compute base inverse correlation matrix R0=(X0(ke)⊤X0(ke)+γII)−1R_0 = (X_0^{(\text{ke})\top} X_0^{(\text{ke})} + \gamma I_I)^{-1}
    7. Compute initial classifier weights W^FCN(0)=R0X0(ke)⊤Y0train\hat{W}_{\text{FCN}}^{(0)} = R_0 X_0^{(\text{ke})\top} Y_0^{\text{train}}
    8. for k=1k = 1 to KK do
    9. Generate CC augmented views A1(Xktrain),…,AC(Xktrain)A_1(X_k^{\text{train}}), \dots, A_C(X_k^{\text{train}})
    10. Extract CNN features and apply kernel embedding to form stacked matrix Xˉk(fe)\bar{X}_k^{(\text{fe})} and stacked label matrix Yˉktrain\bar{Y}_k^{\text{train}}
    11. Update inverse matrix Rk=Rk−1−Rk−1Xˉk(fe)⊤(ICNk+Xˉk(fe)Rk−1Xˉk(fe)⊤)−1Xˉk(fe)Rk−1R_k = R_{k-1} - R_{k-1} \bar{X}_k^{(\text{fe})\top} (I_{C N_k} + \bar{X}_k^{(\text{fe})} R_{k-1} \bar{X}_k^{(\text{fe})\top})^{-1} \bar{X}_k^{(\text{fe})} R_{k-1}
    12. Update classifier weights W^FCN(k)=[W^FCN(k−1)−RkXˉk(fe)⊤Xˉk(fe)W^FCN(k−1),  RkXˉk(fe)⊤Yˉktrain]\hat{W}_{\text{FCN}}^{(k)} = [\hat{W}_{\text{FCN}}^{(k-1)} - R_k \bar{X}_k^{(\text{fe})\top} \bar{X}_k^{(\text{fe})} \hat{W}_{\text{FCN}}^{(k-1)}, \; R_k \bar{X}_k^{(\text{fe})\top} \bar{Y}_k^{\text{train}}]
    13. end for
  5. Knowl 5 — Experimental Setup and Evaluation Protocols for FSCIL Benchmarks

    experimental setup

    GKEAL is evaluated on three benchmark datasets with standard FSCIL split configurations:

    • CIFAR-100: 100 classes with 32×3232 \times 32 images. 60 base classes (phase 0) and 8 incremental phases of 5-way 5-shot tasks (5 classes, 5 samples per class per phase). Backbone is ResNet-20 trained on base data for 300 epochs with initial learning rate 0.1, divided by 10 at epochs 150, 225, and 275.
    • mini-ImageNet: 100 classes with 84×8484 \times 84 images. 60 base classes (phase 0) and 8 incremental phases of 5-way 5-shot tasks. Backbone is ResNet-18 trained on base data for 300 epochs with identical learning rate schedule to CIFAR-100.
    • CUB200-2011: 200 bird categories with 224×224224 \times 224 images. 100 base classes (phase 0) and 10 incremental phases of 10-way 5-shot tasks. Backbone is ResNet-18 pre-trained on ImageNet, fine-tuned on base data for 30 epochs with initial learning rate 0.01 divided by 10 at epoch 15.

    Evaluation is measured by top-1 classification accuracy AkA_k on all seen classes up to phase kk (evaluated on test set D0:ktest\mathcal{D}_{0:k}^{\text{test}}) and the Performance Drop rate PD=A0−AKPD = A_0 - A_K, where lower PDPD indicates less catastrophic forgetting. The regularization parameter is fixed to γ=1\gamma = 1. Hyperparameters selected via validation grid search are:

    • CIFAR-100: I=5000I = 5000, β=10\beta = 10, C=200C = 200.
    • mini-ImageNet: I=10000I = 10000, β=10\beta = 10, C=200C = 200.
    • CUB200-2011: I=10000I = 10000, β=15\beta = 15, C=10C = 10.
  6. Knowl 6 — Benchmark Accuracy and Performance Drop on FSCIL Datasets

    data/table

    GKEAL achieves state-of-the-art performance across all phases on mini-ImageNet, CIFAR-100, and CUB200-2011, exhibiting both the highest final-phase accuracy (AKA_K) and the lowest performance drop (PD=A0−AKPD = A_0 - A_K) among compared methods.

    mini-ImageNet Phase 0 Phase 1 Phase 2 Phase 3 Phase 4 Phase 5 Phase 6 Phase 7 Phase 8 PD↓PD\downarrow
    iCaRL 61.31 46.32 42.94 37.63 30.49 24.00 20.89 18.80 17.21 44.10
    EEIL 61.31 46.58 44.00 37.29 33.14 27.12 24.10 21.57 19.58 41.73
    LUCIR 61.31 47.80 39.31 31.91 25.68 21.35 18.67 17.24 14.17 47.14
    TOPIC 61.31 50.09 45.17 41.16 37.48 35.52 32.19 29.46 24.42 36.89
    CEC 72.00 66.83 62.97 59.43 56.70 53.73 51.19 49.24 47.63 24.37
    F2M 72.05 67.47 63.16 59.70 56.71 53.77 51.11 49.21 47.84 24.21
    MetaFSCIL 72.04 67.94 63.77 60.29 57.58 55.16 52.90 50.79 49.19 22.85
    Entropy-reg 71.84 67.12 63.21 59.77 57.01 53.95 51.55 49.52 48.21 23.63
    GKEAL 73.59 68.90 65.33 62.29 59.39 56.70 54.20 52.59 51.31 22.28
    CIFAR-100 Phase 0 Phase 1 Phase 2 Phase 3 Phase 4 Phase 5 Phase 6 Phase 7 Phase 8 PD↓PD\downarrow
    iCaRL 64.10 53.28 41.69 34.13 27.93 25.06 20.41 15.48 13.73 50.37
    EEIL 64.10 53.11 43.71 35.15 28.96 24.98 21.01 17.26 15.85 48.25
    LUCIR 64.10 53.05 43.96 36.97 31.61 26.73 21.23 16.78 13.54 50.56
    TOPIC 64.10 55.88 47.07 45.16 40.11 36.38 33.96 31.55 29.37 34.73
    CEC 73.07 68.88 65.26 61.19 58.09 55.57 53.22 51.34 49.14 23.93
    F2M 71.45 68.10 64.43 60.80 57.76 55.26 53.53 51.57 49.35 22.06
    MetaFSCIL 74.50 70.10 66.84 62.77 59.48 56.52 54.36 52.56 49.97 24.53
    Entropy-reg 74.40 70.20 66.54 62.51 59.71 56.58 54.52 52.39 50.14 24.26
    GKEAL 74.01 70.45 67.01 63.08 60.01 57.30 55.50 53.39 51.40 22.61
    CUB200-2011 Phase 0 Phase 2 Phase 4 Phase 6 Phase 8 Phase 10 PD↓PD\downarrow
    iCaRL 68.68 48.61 36.62 27.83 24.01 21.16 47.52
    EEIL 68.68 47.91 36.30 25.93 23.95 22.11 46.57
    LUCIR 68.68 44.21 26.71 24.62 20.12 19.87 48.81
    TOPIC 68.68 54.81 45.25 38.35 32.22 26.26 42.40
    CEC 75.85 68.50 62.43 57.73 54.83 52.28 23.57
    F2M 77.13 70.27 64.34 60.52 57.15 55.89 21.24
    MetaFSCIL 75.90 68.78 62.96 58.30 54.78 52.64 23.26
    Entropy-reg 75.90 68.64 62.58 57.82 54.92 52.39 23.51
    GKEAL 78.88 72.32 67.23 62.98 60.20 58.67 20.21

    On mini-ImageNet, GKEAL achieves a final accuracy of 51.31%51.31\%, exceeding the previous best method (MetaFSCIL at 49.19%49.19\%) by 2.12%2.12\%, and records the lowest PDPD of 22.28%22.28\%. On CIFAR-100, GKEAL reaches a final accuracy of 51.40%51.40\%. On CUB200-2011, GKEAL achieves 58.67%58.67\% final accuracy with a lowest PDPD of 20.21%20.21\%.

  7. Knowl 7 — Ablation Study of GKE and AFC Modules

    data/table

    An ablation study conducted on CIFAR-100 shows the individual and joint necessity of the Gaussian Kernel Embedding (GKE) and Augmented Feature Concatenation (AFC) modules in GKEAL.

    GKE AFC Phase 0 Phase 1 Phase 2 Phase 3 Phase 4 Phase 5 Phase 6 Phase 7 Phase 8
    ×\times ×\times 13.20 11.99 11.29 10.01 9.39 9.22 8.81 8.10 7.99
    ×\times ✓ 12.56 10.80 10.29 9.81 9.36 8.60 8.00 7.89 7.22
    ✓ ×\times 74.80 68.98 64.11 59.35 55.78 52.28 49.08 47.02 43.79
    ✓ ✓ 74.35 70.32 66.21 62.37 60.01 56.98 55.12 53.39 51.21
    • Without GKE: Applying analytic least-squares regression directly to raw CNN features (X0(cnn)X_0^{(\text{cnn})}) leads to severe breakdown due to ill-conditioned correlation matrices, resulting in catastrophic failure (accuracy dropping from 13.20%13.20\% at Phase 0 to 7.99%7.99\% at Phase 8 without AFC, and 7.22%7.22\% with AFC).
    • With GKE alone: Embedding the CNN features into a higher-dimensional kernel space restores performance, achieving 74.80%74.80\% at Phase 0 and 43.79%43.79\% at Phase 8.
    • With GKE and AFC combined: AFC improves final phase accuracy from 43.79%43.79\% to 51.21%51.21\% (a +7.42%+7.42\% gain), mitigating base-new class imbalance.
  8. Knowl 8 — Hyperparameter Sensitivity in GKEAL

    empirical result

    Empirical evaluation of the hyperparameters in GKEAL reveals the following behavior across CIFAR-100, mini-ImageNet, and CUB200-2011:

    1. Number of kernel centers II: Increasing II monotonicially boosts final-phase accuracy AKA_K, peaking around I∈[8k,12k]I \in [8\text{k}, 12\text{k}] (e.g., I=5kI=5\text{k} on CIFAR-100, I=10kI=10\text{k} on mini-ImageNet and CUB200-2011). Higher dimensionality compensates for under-fitting tendencies in linear least-squares estimators.
    2. Kernel width β\beta: Accuracy is sensitive to β\beta, with optimal performance in the range β∈[5,15]\beta \in [5, 15] for CIFAR-100 and mini-ImageNet, and peaking at β=15\beta = 15 for CUB200-2011. Values that are too small or too large collapse accuracy.
    3. Augmentation count CC: Optimal performance is achieved around C=200C = 200 for CIFAR-100 and mini-ImageNet, and C=10C = 10 for CUB200-2011. This difference aligns with the ratio of augmentation count to base class sample size: C/Nbase_per_class=200/500=0.40C / N_{\text{base\_per\_class}} = 200 / 500 = 0.40 on CIFAR-100/mini-ImageNet versus 10/30=0.3310 / 30 = 0.33 on CUB200-2011. Evaluating test accuracy separately on base classes (D0test\mathcal{D}_0^{\text{test}}) and new classes (D1:Ktest\mathcal{D}_{1:K}^{\text{test}}) confirms that larger CC monotonically increases new class accuracy while moderately reducing base class accuracy.
  9. Knowl 9 — High Parameter Footprint from Large Kernel Dimension in GKEAL

    limitation

    A key limitation of GKEAL is its reliance on a large number of kernel center vectors II (typically I∈[5000,12000]I \in [5000, 12000]) in the Kernel Analytic Module to overcome the under-fitting tendency of linear least-squares regression. This large kernel embedding dimension requires storing and updating an I×II \times I correlation inverse matrix RkR_k and an I×(∑j=0kdyj)I \times (\sum_{j=0}^k d_{y_j}) weight matrix, substantially increasing the parameter count and memory footprint compared to low-dimensional linear classifiers.

Coverage note — None. All primary contributions—including the GKEAL framework, recursive weight-invariant mathematical derivation, Kernel Analytic Module, Augmented Feature Concatenation module, full pseudocode algorithm, main benchmark comparisons across CIFAR-100/mini-ImageNet/CUB200-2011, ablation experiments, hyperparameter analyses, and stated limitations—are fully covered.

References

  1. 1.Francisco M. Castro, Manuel J. Marin-Jimenez, Nicolas Guil, Cordelia Schmid, and Karteek Alahari. End-to-end incremental learning. In Proceedings of the European Conference on Computer Vision (ECCV), September 2018. 1, 2, 5, 6, 7
  2. 2.Ali Cheraghian, Shafin Rahman, Pengfei Fang, Soumava Kumar Roy, Lars Petersson, and Mehrtash Harandi. Semantic-aware knowledge distillation for few-shot class-incremental learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 2534–2543, June 2021. 2
  3. 3.Zhixiang Chi, Li Gu, Huan Liu, Yang Wang, Yuanhao Yu, and Jin Tang. Metafscil: A meta-learning approach for few-shot class incremental learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 14166–14175, June 2022. 5, 6, 7
  4. 4.Arthur Douillard, Matthieu Cord, Charles Ollion, Thomas Robert, and Eduardo Valle. Podnet: Pooled outputs distillation for small-tasks incremental learning. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XX 16, pages 86–102. Springer, 2020. 2
  5. 5.Chelsea Finn, Pieter Abbeel, and Sergey Levine. Model-agnostic meta-learning for fast adaptation of deep networks. In Doina Precup and Yee Whye Teh, editors, Proceedings of the 34th International Conference on Machine Learning, volume 70 of Proceedings of Machine Learning Research, pages 1126–1135. PMLR, 06–11 Aug 2017. 2
  6. 6.Spyros Gidaris and Nikos Komodakis. Dynamic few-shot visual learning without forgetting. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), June 2018. 2
  7. 7.Ping Guo, Michael R Lyu, and NE Mastorakis. Pseudoinverse learning algorithm for feedforward neural networks. Advances in Neural Networks and Applications, pages 321–326, 2001. 2
  8. 8.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), June 2016. 5
  9. 9.Saihui Hou, Xinyu Pan, Chen Change Loy, Zilei Wang, and Dahua Lin. Learning a unified classifier incrementally via rebalancing. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2019. 1, 2, 5, 6, 7
  10. 10.Muhammad Abdullah Jamal and Guo-Jun Qi. Task agnostic meta-learning for few-shot learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2019. 2
  11. 11.James Kirkpatrick, Razvan Pascanu, Neil Rabinowitz, Joel Veness, Guillaume Desjardins, Andrei A Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, et al. Overcoming catastrophic forgetting in neural networks. Proceedings of the national academy of sciences, 114(13):3521–3526, 2017. 1, 2
  12. 12.Alex Krizhevsky, Geoffrey Hinton, et al. Learning multiple layers of features from tiny images. 2009. 5
  13. 13.Zhizhong Li and Derek Hoiem. Learning without forgetting. IEEE Transactions on Pattern Analysis and Machine Intelligence, 40(12):2935–2947, 2018. 1, 2
  14. 14.Huan Liu, Li Gu, Zhixiang Chi, Yang Wang, Yuanhao Yu, Jun Chen, and Jin Tang. Few-shot class-incremental learning via entropy-regularized data-free replay. In Computer Vision–ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part XXIV, pages 146–162. Springer, 2022. 5, 6, 7
  15. 15.Xialei Liu, Marc Masana, Luis Herranz, Joost Van de Weijer, Antonio M. Lopez, and Andrew D. Bagdanov. Rotate your networks: Better weight consolidation and less catastrophic forgetting. In 2018 24th International Conference on Pattern Recognition (ICPR), pages 2262–2268, 2018. 2
  16. 16.Yaoyao Liu, Bernt Schiele, and Qianru Sun. An ensemble of epoch-wise empirical bayes for few-shot learning. In Andrea Vedaldi, Horst Bischof, Thomas Brox, and Jan-Michael Frahm, editors, Computer Vision – ECCV 2020. Springer International Publishing, 2020. 2
  17. 17.Yaoyao Liu, Bernt Schiele, and Qianru Sun. Adaptive aggregation networks for class-incremental learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 2544–2553, June 2021. 1, 2
  18. 18.Yaoyao Liu, Bernt Schiele, and Qianru Sun. Rmm: Reinforced memory management for class-incremental learning. Advances in Neural Information Processing Systems, 34, 2021. 2
  19. 19.J. Park and I. W. Sandberg. Universal approximation using radial-basis-function networks. Neural Computation, 3(2):246–257, 1991. 2
  20. 20.Sylvestre-Alvise Rebuffi, Alexander Kolesnikov, Georg Sperl, and Christoph H. Lampert. icarl: Incremental classifier and representation learning. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), July 2017. 1, 2, 5, 6, 7
  21. 21.Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, et al. Imagenet large scale visual recognition challenge. International journal of computer vision, 115(3):211–252, 2015. 5
  22. 22.Guangyuan Shi, Jiaxin Chen, Wenlong Zhang, Li-Ming Zhan, and Xiao-Ming Wu. Overcoming catastrophic forgetting in incremental few-shot learning by finding flat minima. In M. Ranzato, A. Beygelzimer, Y. Dauphin, P.S. Liang, and J. Wortman Vaughan, editors, Advances in Neural Information Processing Systems, volume 34, pages 6747–6761. Curran Associates, Inc., 2021. 1, 2, 5, 6, 7
  23. 23.Xiaoyu Tao, Xiaopeng Hong, Xinyuan Chang, Songlin Dong, Xing Wei, and Yihong Gong. Few-shot class-incremental learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2020. 1, 2, 5, 6, 7
  24. 24.Kar-Ann Toh. Learning from the kernel and the range space. In the Proceedings of the 17th 2018 IEEE Conference on Computer and Information Science, pages 417–422. IEEE, June 2018. 2
  25. 25.C. Wah, S. Branson, P. Welinder, P. Perona, and S. Belongie. The caltech-ucsd birds-200-2011 dataset. 2011. 5
  26. 26.X. Wang, T. Zhang, and R. Wang. Noniterative deep learning: Incorporating restricted boltzmann machine into multilayer random weight neural networks. IEEE Transactions on Systems, Man, and Cybernetics: Systems, 49(7):1299–1308, 2019. 2
  27. 27.Yue Wu, Yinpeng Chen, Lijuan Wang, Yuancheng Ye, Zicheng Liu, Yandong Guo, and Yun Fu. Large scale incremental learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2019. 2
  28. 28.Han-Jia Ye, Hexiang Hu, and De-Chuan Zhan. Learning adaptive classifiers synthesis for generalized few-shot learning. International Journal of Computer Vision, 129(6):1930–1953, 2021. 2
  29. 29.Han-Jia Ye, Hexiang Hu, De-Chuan Zhan, and Fei Sha. Few-shot learning via embedding adaptation with set-to-set functions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2020. 2
  30. 30.Chi Zhang, Yujun Cai, Guosheng Lin, and Chunhua Shen. Deepemd: Few-shot image classification with differentiable earth mover's distance and structured classifiers. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2020. 2
  31. 31.Chi Zhang, Nan Song, Guosheng Lin, Yun Zheng, Pan Pan, and Yinghui Xu. Few-shot incremental learning with continually evolved classifiers. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 12455–12464, June 2021. 1, 2, 5, 6, 7
  32. 32.Hanbin Zhao, Yongjian Fu, Mintong Kang, Qi Tian, Fei Wu, and Xi Li. Mgsvf: Multi-grained slow vs. fast framework for few-shot class-incremental learning. IEEE Transactions on Pattern Analysis and Machine Intelligence, pages 1–1, 2021. 2
  33. 33.Huiping Zhuang, Zhiping Lin, and Kar-Ann Toh. Blockwise recursive Moore-Penrose inverse for network learning. IEEE Transactions on Systems, Man, and Cybernetics: Systems, pages 1–14, 2021. 1, 2, 4
  34. 34.Huiping Zhuang, Zhiping Lin, and Kar-Ann Toh. Correlation projection for analytic learning of a classification network. Neural Processing Letters, pages 1–22, 2021. 2, 4
  35. 35.Huiping Zhuang, Zhiping Lin, and Kar-Ann Toh. Training multilayer neural networks analytically using kernel projection. In 2021 IEEE International Symposium on Circuits and Systems (ISCAS), pages 1–5, 2021. 4
  36. 36.Huiping Zhuang, Zhenyu Weng, Hongxin Wei, Renchunzi Xie, Kar-Ann Toh, and Zhiping Lin. ACIL: Analytic class-incremental learning with absolute memorization and privacy protection. In Advances in Neural Information Processing Systems, 2022. 3

Citation

MLA
Zhuang, H., et al. “GKEAL: Gaussian Kernel Embedded Analytic Learning for Few-Shot Class Incremental Task”. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023, pp. 7746–55, https://doi.org/10.1109/CVPR52729.2023.00748.
APA
Zhuang, H., Weng, Z., He, R., Lin, Z., & Zeng, Z. (2023). GKEAL: Gaussian Kernel Embedded Analytic Learning for Few-Shot Class Incremental Task. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 7746–7755. https://doi.org/10.1109/CVPR52729.2023.00748
Chicago
Zhuang, H., Z. Weng, R. He, Z. Lin, and Z. Zeng. 2023. “GKEAL: Gaussian Kernel Embedded Analytic Learning for Few-Shot Class Incremental Task”. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 7746–55. https://doi.org/10.1109/CVPR52729.2023.00748.
Harvard
Zhuang, H. et al. (2023) “GKEAL: Gaussian Kernel Embedded Analytic Learning for Few-Shot Class Incremental Task”, 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp. 7746–7755. Available at: https://doi.org/10.1109/CVPR52729.2023.00748.
Vancouver
1. Zhuang H, Weng Z, He R, Lin Z, Zeng Z (2023) GKEAL: Gaussian Kernel Embedded Analytic Learning for Few-Shot Class Incremental Task. In: 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp 7746–7755

BibTeX

@inproceedings{Zhuang_2023, title={GKEAL: Gaussian Kernel Embedded Analytic Learning for Few-Shot Class Incremental Task}, url={http://dx.doi.org/10.1109/CVPR52729.2023.00748}, DOI={10.1109/cvpr52729.2023.00748}, booktitle={2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)}, publisher={IEEE}, author={Zhuang, Huiping and Weng, Zhenyu and He, Run and Lin, Zhiping and Zeng, Ziqian}, year={2023}, month=June, pages={7746–7755} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE