ACIL: Analytic Class-Incremental Learning with Absolute Memorization and Privacy Protection

Huiping ZhuangZhenyu WengHongxin WeiRenchunzi XieKar-Ann TohZhiping Lin

article2022NeurIPS79 citations

Proposes an analytic class-incremental learning framework that mathematically matches the performance of joint training without storing historical exemplar data, eliminating catastrophic forgetting while protecting data privacy across multi-phase learning tasks.

Listen

Modern computer vision systems frequently need to learn new visual categories over time without retraining from scratch. Traditional incremental learning systems suffer from catastrophic forgetting, where the model forgets older knowledge when exposed to new tasks. While storing and replaying a subset of historical user images mitigates this problem, it introduces severe data privacy risks and increases memory costs. Existing privacy-preserving alternatives that avoid storing past data typically fail to match state-of-the-art accuracy, especially as the number of incremental phases grows.

The article demonstrates an incremental learning framework called Analytic Class-Incremental Learning, which mathematically guarantees exact retention of past knowledge without retaining any historical training images. The authors evaluate this approach against leading benchmarks across various sequential learning setups.

The evaluated method first trains a core neural network feature extractor on initial base data. It then applies an expanded linear classification layer solved via a direct mathematical formula rather than iterative optimization. For all subsequent tasks, the system updates its classification weights recursively using only new training samples and a compact, summary correlation matrix that compresses prior knowledge without revealing underlying raw data. The authors validated this mechanism across standard benchmark image datasets—including CIFAR-100, ImageNet-Subset, and full ImageNet—over multi-phase schedules ranging from 5 to 50 phases.

The analysis yields four key findings. First, the recursive incremental formula yields identical mathematical results to training on all historical and new data together in a single batch, ensuring complete memorization. Second, because knowledge retention does not degrade over time, performance remains stable regardless of phase count; for instance, accuracy on CIFAR-100 stayed around 66% whether divided into 5 or 50 phases. Third, the method outperforms top replay-based competitors in long learning sequences, leading 25-phase CIFAR-100 benchmarks with 65.95% accuracy compared to 64.12% for the closest alternative. Finally, the framework significantly reduces forgetting of initial base classes, showing an initial-class accuracy drop of only 2.75% on full ImageNet compared to 13.63% in top-performing baseline techniques.

These results demonstrate that organizations do not need to compromise privacy compliance to maintain high-accuracy continuous learning models. By relying on a fixed-size summary matrix instead of stored image exemplars, the framework mitigates regulatory privacy exposure, prevents reverse-engineering of user data, and offers substantial memory savings on high-resolution image workloads.

Organizations deploying continuous learning across privacy-sensitive domains should consider adopting analytic classification layers to streamline model updates. Decision-makers should note that the system relies on a frozen feature extractor after initial training; therefore, foundational base training must encompass a diverse and representative data sample. Future work and pilot evaluations should focus on testing hardware scaling limits for larger expansion sizes and optimizing initial feature extraction to further close performance gaps in early-phase learning.

arXiv: 2205.14922
Cover for ACIL: Analytic Class-Incremental Learning with Absolute Memorization and Privacy Protection

Abstract

Class-incremental learning (CIL) learns a classification model with training data of different classes arising progressively. Existing CIL either suffers from serious accuracy loss due to catastrophic forgetting, or invades data privacy by revisiting used exemplars. Inspired by linear learning formulations, we propose an analytic class-incremental learning (ACIL) with absolute memorization of past knowledge while avoiding breaching of data privacy (i.e., without storing historical data). The absolute memorization is demonstrated in the sense that class-incremental learning using ACIL given present data would give identical results to that from its joint-learning counterpart which consumes both present and historical samples. This equality is theoretically validated. Data privacy is ensured since no historical data are involved during the learning process. Empirical validations demonstrate ACIL's competitive accuracy performance with near-identical results for various incremental task settings (e.g., 5-50 phases). This also allows ACIL to outperform the state-of-the-art methods for large-phase scenarios (e.g., 25 and 50 phases).

Table of Contents

  • 1 Introduction
  • 2 Related Works
  • 2.1 Class-Incremental Learning
  • 2.2 Analytic Learning
  • 3 The Proposed Method
  • 3.1 The Base Training Agenda
  • 3.2 The Class-Incremental Learning Agenda
  • 4 Experiments
  • 4.1 Datasets and Implementation Details
  • 4.2 Evaluation Metric
  • 4.3 Result Comparison
  • 4.4 Ablation Study
  • 4.5 Potential Positive and Negative Societal Impacts
  • 5 Conclusion
  • 6 Acknowledgment
  • References
  • Checklist

Knowls

  1. Knowl 1 — Recursive Update of Analytic Classifier and Regularized Feature Autocorrelation Matrix in ACIL

    theoretical result

    In Analytic Class-Incremental Learning (ACIL), the classifier weight update across incremental learning phases is formulated to achieve absolute memorization of historical knowledge without storing any past training samples.

    Let Dktrain={Xktrain,Yktrain}\mathcal{D}_k^{\text{train}} = \{X_k^{\text{train}}, Y_k^{\text{train}}\} denote the training dataset at incremental phase k∈{1,2,…,K}k \in \{1, 2, \dots, K\}, where Xktrain∈RNk×w×h×cX_k^{\text{train}} \in \mathbb{R}^{N_k \times w \times h \times c} is the batch of NkN_k input images and Yktrain∈RNk×dykY_k^{\text{train}} \in \mathbb{R}^{N_k \times d_{y_k}} is the one-hot label matrix for dykd_{y_k} newly introduced classes. Let Xk(fe)∈RNk×dfeX_k^{(\text{fe})} \in \mathbb{R}^{N_k \times d_{\text{fe}}} denote the expanded feature matrix extracted from XktrainX_k^{\text{train}}, and let γ>0\gamma > 0 be a ridge regularization coefficient. The regularized feature autocorrelation matrix (RFAuM) at phase k−1k-1 is defined as:

    Rk−1=(∑i=0k−1(Xi(fe))TXi(fe)+γI)−1∈Rdfe×dfeR_{k-1} = \left( \sum_{i=0}^{k-1} (X_i^{(\text{fe})})^T X_i^{(\text{fe})} + \gamma I \right)^{-1} \in \mathbb{R}^{d_{\text{fe}} \times d_{\text{fe}}}

    where I∈Rdfe×dfeI \in \mathbb{R}^{d_{\text{fe}} \times d_{\text{fe}}} is the identity matrix.

    At phase kk, the updated RFAuM RkR_k is computed recursively using the Woodbury matrix identity without accessing historical datasets D0:k−1train\mathcal{D}_{0:k-1}^{\text{train}}:

    Rk=Rk−1−Rk−1(Xk(fe))T(I+Xk(fe)Rk−1(Xk(fe))T)−1Xk(fe)Rk−1R_k = R_{k-1} - R_{k-1} (X_k^{(\text{fe})})^T \left( I + X_k^{(\text{fe})} R_{k-1} (X_k^{(\text{fe})})^T \right)^{-1} X_k^{(\text{fe})} R_{k-1}

    Using RkR_k, the expanded classifier weight matrix W^FCN(k)\hat{W}_{\text{FCN}}^{(k)} is recursively updated from the previous weight W^FCN(k−1)\hat{W}_{\text{FCN}}^{(k-1)} by:

    W^FCN(k)=[W^FCN(k−1)−Rk(Xk(fe))TXk(fe)W^FCN(k−1)|Rk(Xk(fe))TYktrain]\hat{W}_{\text{FCN}}^{(k)} = \left[ \hat{W}_{\text{FCN}}^{(k-1)} - R_k (X_k^{(\text{fe})})^T X_k^{(\text{fe})} \hat{W}_{\text{FCN}}^{(k-1)} \quad\middle|\quad R_k (X_k^{(\text{fe})})^T Y_k^{\text{train}} \right]

    This recursive solution is mathematically identical to the closed-form regularized least-squares estimator obtained by joint training on the full cumulative dataset D0:ktrain\mathcal{D}_{0:k}^{\text{train}}:

    W^FCN(k)=(∑i=0k(Xi(fe))TXi(fe)+γI)−1[(X0(fe))TY0train(X1(fe))TY1train…(Xk(fe))TYktrain]\hat{W}_{\text{FCN}}^{(k)} = \left( \sum_{i=0}^k (X_i^{(\text{fe})})^T X_i^{(\text{fe})} + \gamma I \right)^{-1} \left[ (X_0^{(\text{fe})})^T Y_0^{\text{train}} \quad (X_1^{(\text{fe})})^T Y_1^{\text{train}} \quad \dots \quad (X_k^{(\text{fe})})^T Y_k^{\text{train}} \right]

  2. Knowl 2 — Analytic Class-Incremental Learning Algorithm

    algorithm

    Analytic Class-Incremental Learning (ACIL) executes class-incremental learning in two successive agendas: a Base Training Agenda on base dataset D0train\mathcal{D}_0^{\text{train}} and a recursive Class-Incremental Learning (CIL) Agenda on sequential task datasets D1train,…,DKtrain\mathcal{D}_1^{\text{train}}, \dots, \mathcal{D}_K^{\text{train}}.

    Input: Base dataset D0train={X0train,Y0train}\mathcal{D}_0^{\text{train}} = \{X_0^{\text{train}}, Y_0^{\text{train}}\}, sequential incremental datasets Dktrain={Xktrain,Yktrain}\mathcal{D}_k^{\text{train}} = \{X_k^{\text{train}}, Y_k^{\text{train}}\} for k=1,…,Kk = 1, \dots, K, expansion dimension dfed_{\text{fe}}, regularization parameter γ\gamma
    Output: Sequence of classifier weights W^FCN(0),W^FCN(1),…,W^FCN(K)\hat{W}_{\text{FCN}}^{(0)}, \hat{W}_{\text{FCN}}^{(1)}, \dots, \hat{W}_{\text{FCN}}^{(K)}
    Base Training Agenda:
      1. Train CNN backbone fCNN(⋅,WCNN)f_{\text{CNN}}(\cdot, W_{\text{CNN}}) and initial classifier using backpropagation for MM epochs on D0train\mathcal{D}_0^{\text{train}}
      2. Freeze CNN backbone parameters WCNNW_{\text{CNN}} for all subsequent operations
      3. Initialize random feature expansion matrix Wfe∈Rdcnn×dfeW_{\text{fe}} \in \mathbb{R}^{d_{\text{cnn}} \times d_{\text{fe}}} with elements sampled i.i.d. from N(0,1)\mathcal{N}(0, 1)
      4. Extract base features: X0(cnn)=fflat(fCNN(X0train,WCNN))X_0^{(\text{cnn})} = f_{\text{flat}}(f_{\text{CNN}}(X_0^{\text{train}}, W_{\text{CNN}}))
      5. Compute expanded base features: X0(fe)=ReLU(X0(cnn)Wfe)X_0^{(\text{fe})} = \text{ReLU}(X_0^{(\text{cnn})} W_{\text{fe}})
      6. Compute initial RFAuM: R0=((X0(fe))TX0(fe)+γI)−1R_0 = ((X_0^{(\text{fe})})^T X_0^{(\text{fe})} + \gamma I)^{-1}
      7. Compute initial analytic classifier: W^FCN(0)=R0(X0(fe))TY0train\hat{W}_{\text{FCN}}^{(0)} = R_0 (X_0^{(\text{fe})})^T Y_0^{\text{train}}
    Class-Incremental Learning Agenda:
      for k=1k = 1 to KK do
        1. Extract features for phase kk: Xk(cnn)=fflat(fCNN(Xktrain,WCNN))X_k^{(\text{cnn})} = f_{\text{flat}}(f_{\text{CNN}}(X_k^{\text{train}}, W_{\text{CNN}}))
        2. Compute expanded features: Xk(fe)=ReLU(Xk(cnn)Wfe)X_k^{(\text{fe})} = \text{ReLU}(X_k^{(\text{cnn})} W_{\text{fe}})
        3. Update RFAuM:
           Rk=Rk−1−Rk−1(Xk(fe))T(I+Xk(fe)Rk−1(Xk(fe))T)−1Xk(fe)Rk−1R_k = R_{k-1} - R_{k-1} (X_k^{(\text{fe})})^T (I + X_k^{(\text{fe})} R_{k-1} (X_k^{(\text{fe})})^T)^{-1} X_k^{(\text{fe})} R_{k-1}
        4. Update classifier weights:
           W^FCN(k)=[W^FCN(k−1)−Rk(Xk(fe))TXk(fe)W^FCN(k−1)∣Rk(Xk(fe))TYktrain]\hat{W}_{\text{FCN}}^{(k)} = [ \hat{W}_{\text{FCN}}^{(k-1)} - R_k (X_k^{(\text{fe})})^T X_k^{(\text{fe})} \hat{W}_{\text{FCN}}^{(k-1)} \mid R_k (X_k^{(\text{fe})})^T Y_k^{\text{train}} ]
        5. Discard Dktrain\mathcal{D}_k^{\text{train}} to preserve data privacy
      end for

    Base training utilizes conventional backpropagation (e.g., SGD with momentum for 160 epochs on ResNet-32 or 90 epochs on ResNet-18). All incremental learning phases operate in a single epoch per task batch using purely analytic matrix operations.

  3. Knowl 3 — Random Feature Expansion in Analytic Re-alignment Base Training

    model/method

    Because analytic linear regression on fixed feature representations is prone to underfitting, Analytic Re-alignment Base Training (ARaBT) expands the convolutional feature representations into a higher-dimensional space prior to the analytic classifier solution.

    Let fCNN(⋅,WCNN)f_{\text{CNN}}(\cdot, W_{\text{CNN}}) be the feature extractor trained on the base dataset and frozen thereafter, and let fflatf_{\text{flat}} denote the flattening operator reshaping the output of the convolutional backbone into a 1-D vector. For an input batch Xitrain∈RNi×w×h×cX_i^{\text{train}} \in \mathbb{R}^{N_i \times w \times h \times c}, the extracted backbone feature matrix is:

    Xi(cnn)=fflat(fCNN(Xitrain,WCNN))∈RNi×dcnnX_i^{(\text{cnn})} = f_{\text{flat}}(f_{\text{CNN}}(X_i^{\text{train}}, W_{\text{CNN}})) \in \mathbb{R}^{N_i \times d_{\text{cnn}}}

    To increase representation capacity, an additional linear layer parameterized by a fixed expansion matrix Wfe∈Rdcnn×dfeW_{\text{fe}} \in \mathbb{R}^{d_{\text{cnn}} \times d_{\text{fe}}} is applied, where dfe≥dcnnd_{\text{fe}} \ge d_{\text{cnn}} is the expansion dimension. Each element of WfeW_{\text{fe}} is drawn independently from a standard normal distribution N(0,1)\mathcal{N}(0, 1) and fixed. The expanded feature matrix Xi(fe)X_i^{(\text{fe})} is produced by passing the linear projection through an element-wise activation function factf_{\text{act}} (specifically ReLU\text{ReLU}):

    Xi(fe)=fact(Xi(cnn)Wfe)=ReLU(Xi(cnn)Wfe)∈RNi×dfeX_i^{(\text{fe})} = f_{\text{act}}(X_i^{(\text{cnn})} W_{\text{fe}}) = \text{ReLU}(X_i^{(\text{cnn})} W_{\text{fe}}) \in \mathbb{R}^{N_i \times d_{\text{fe}}}

    The expanded features are directly mapped onto classification targets via regularized linear regression.

  4. Knowl 4 — Average Incremental Accuracy and Forgetting Rate Metrics in Class-Incremental Learning

    definition

    In class-incremental learning protocols where a network learns KK sequential tasks after an initial base task (phase 0), performance across phases is evaluated using Average Incremental Accuracy Aˉ\bar{\mathcal{A}} and Forgetting Rate F\mathcal{F}.

    1. Average Incremental Accuracy (Aˉ\bar{\mathcal{A}}) evaluates the overall classification performance across all phases:

    Aˉ=1K+1∑k=0KAk\bar{\mathcal{A}} = \frac{1}{K+1} \sum_{k=0}^K \mathcal{A}_k

    where Ak\mathcal{A}_k is the classification test accuracy of the model after completing training at phase kk, evaluated on the cumulative test dataset D0:ktest\mathcal{D}_{0:k}^{\text{test}} containing all classes encountered from phase 00 through phase kk.

    1. Forgetting Rate (F\mathcal{F}) measures the degradation in performance on the base classes (phase 0) after learning subsequent incremental classes:

    F=A0Z−AKZ\mathcal{F} = \mathcal{A}_0^Z - \mathcal{A}_K^Z

    where AkZ\mathcal{A}_k^Z denotes the test accuracy of the model evaluated exclusively on the base class test set D0test\mathcal{D}_0^{\text{test}} at learning phase kk. A lower F\mathcal{F} indicates better retention of initial knowledge.

  5. Knowl 5 — Class-Incremental Benchmark Evaluation on CIFAR-100 and ImageNet

    data/table

    ACIL was benchmarked against exemplar-free privacy-preserving methods (LwF, EWC, SDC) and replay-based methods storing 20 exemplars per class (BiC, iCaRL, LUCIR, PODNet, LUCIR+Mnemonics, POD+AANets, POD+AANets+RMM) using ResNet-32 on CIFAR-100 and ResNet-18 on ImageNet-Subset and ImageNet-Full. Base phase 0 contained 50% of the classes, and the remaining 50% were partitioned evenly across K∈{5,10,25,50}K \in \{5, 10, 25, 50\} phases.

    Method CIFAR-100 (Aˉ\bar{\mathcal{A}} %) ImageNet-Subset (Aˉ\bar{\mathcal{A}} %) ImageNet-Full (Aˉ\bar{\mathcal{A}} %)
    K=5K=5 K=10K=10 K=25K=25 K=50K=50 K=5K=5 K=10K=10 K=25K=25 K=50K=50 K=5K=5 K=10K=10 K=25K=25 K=50K=50
    LwF 49.59 46.98 45.51 - 53.62 47.64 44.32 - 51.50 46.89 43.14 -
    EWC 34.01 32.33 - - 42.35 26.76 - - - - - -
    SDC 55.96 56.56 - - - 62.97 - - - - - -
    BiC 59.36 54.20 50.00 - 70.07 64.96 57.73 - 62.65 58.72 53.47 -
    iCaRL 57.12 52.66 48.22 - 65.44 59.88 52.97 - 51.50 46.89 43.14 -
    LUCIR 63.17 60.14 57.54 - 70.84 68.32 61.44 - 64.45 61.57 56.56 -
    PODNet 64.83 63.19 60.72 57.98 75.54 74.33 68.31 62.48 66.95 64.13 59.17 -
    LUCIR+Mnemonics 64.95 63.25 63.70 - 73.30 72.17 71.50 - 66.15 63.12 63.08 -
    POD+AANets 66.31 64.31 62.31 - 76.96 75.58 71.78 - 67.73 64.85 61.78 -
    POD+AANets+RMM 68.36 66.67 64.12 - 79.50 78.11 75.01 - - - - -
    ACIL 66.30 66.07 65.95 66.01 74.81 74.76 74.59 74.13 65.34 64.84 64.63 64.35
    Method CIFAR-100 (F\mathcal{F} %) ImageNet-Subset (F\mathcal{F} %) ImageNet-Full (F\mathcal{F} %)
    K=5K=5 K=10K=10 K=25K=25 K=50K=50 K=5K=5 K=10K=10 K=25K=25 K=50K=50 K=5K=5 K=10K=10 K=25K=25 K=50K=50
    LwF 43.36 43.58 41.66 - 55.32 57.00 55.12 - 48.70 47.94 49.84 -
    iCaRL 57.12 34.10 36.48 - 43.40 45.84 47.60 - 26.03 33.76 38.80 -
    BiC 31.42 32.50 34.60 - 27.04 31.04 37.88 - 25.06 28.34 33.17 -
    LUCIR 18.70 21.34 26.46 - 31.88 33.48 35.40 - 24.08 27.29 30.30 -
    LUCIR+Mnemonics 11.64 10.90 9.96 - 10.20 9.88 11.76 - 13.63 13.45 14.40 -
    ACIL 9.00 9.72 9.28 9.32 3.91 3.40 3.20 3.43 2.75 3.45 3.31 3.40

    While baseline methods exhibit substantial accuracy degradation as the number of incremental phases KK increases (e.g., PODNet accuracy drops from 64.83%64.83\% at K=5K=5 to 57.98%57.98\% at K=50K=50 on CIFAR-100), ACIL maintains stable accuracy across all task lengths (66.30%66.30\% at K=5K=5 vs 66.01%66.01\% at K=50K=50). On ImageNet-Full, ACIL outperforms all competitors for K≥10K \ge 10, achieving 64.63%64.63\% at K=25K=25 compared to 63.08%63.08\% for LUCIR+Mnemonics. ACIL also achieves the lowest forgetting rates across all datasets (e.g., 2.75%2.75\%--3.45%3.45\% on ImageNet-Full vs 13.45%13.45\%--49.84%49.84\% for baselines).

  6. Knowl 6 — Ablation Analysis of Feature Expansion Size and Ridge Regularization in ACIL

    empirical result

    The effectiveness of ACIL depends on both the random feature expansion (FE) dimension dfed_{\text{fe}} and the ridge regularization parameter γ\gamma.

    A 5-phase incremental learning ablation using ResNet-32 on CIFAR-100 demonstrates the following performance shifts:

    FE Process (dfe=8kd_{\text{fe}} = 8\text{k}) Regularization γ\gamma Aˉ\bar{\mathcal{A}} (%)
    10−110^{-1} 10−210^{-2} 10−310^{-3}
    No Yes No No 52.99%
    Yes Yes No No 66.30%
    Yes No Yes No 66.25%
    Yes No No Yes 66.23%
    Yes No No No 51.12%

    Key findings include:

    1. Necessity of Feature Expansion: Omitting feature expansion (mapping CNN features directly to classes) causes a 13.31%13.31\% absolute decrease in average incremental accuracy (66.30%→52.99%66.30\% \to 52.99\%).
    2. Necessity of Regularization: Removing regularization completely (γ=0\gamma = 0) results in a 15.18%15.18\% accuracy drop (66.30%→51.12%66.30\% \to 51.12\%).
    3. Regularization Robustness: Performance remains stable across a wide range of values γ∈[10−3,10−1]\gamma \in [10^{-3}, 10^{-1}], varying by less than 0.07%0.07\%.
    4. Expansion Dimension Scaling: On ImageNet-Subset and ImageNet-Full, average accuracy monotonically increases as dfed_{\text{fe}} expands from 1k1\text{k} to 15k15\text{k}. On CIFAR-100, accuracy peaks around dfe=8kd_{\text{fe}} = 8\text{k}--10k10\text{k} and slightly declines at 15k15\text{k} because the expansion ratio (15k/6415\text{k} / 64) becomes excessively large relative to the backbone dimension.
  7. Knowl 7 — Exemplar-Free Privacy Preservation and Storage Complexity in ACIL

    model/method

    ACIL preserves data privacy across incremental phases by discarding all training samples from previous tasks D0:k−1train\mathcal{D}_{0:k-1}^{\text{train}} and caching only the regularized feature autocorrelation matrix Rk∈Rdfe×dfeR_k \in \mathbb{R}^{d_{\text{fe}} \times d_{\text{fe}}}.

    1. Privacy Preservation: Unlike replay-based methods that store exemplar images, ACIL never revisits historical data. The original input samples cannot be reconstructed or reverse-engineered from RkR_k because RkR_k is a compressed summation of outer products of expanded latent features with non-linear activation.
    2. Constant Storage Footprint: The storage requirement for historical information in ACIL is strictly fixed at dfe×dfed_{\text{fe}} \times d_{\text{fe}} float elements regardless of dataset sample size or number of incremental phases KK.
      • For dfe=8kd_{\text{fe}} = 8\text{k}, ACIL stores 8,000×8,000=6.4×1078,000 \times 8,000 = 6.4 \times 10^7 (64M) tensor elements.
      • In contrast, storing 20 exemplars per class for standard replay methods on ImageNet (1,0001,000 classes with 224×224×3224 \times 224 \times 3 resolution) requires 20×1,000×224×224×3≈3,010.6M20 \times 1,000 \times 224 \times 224 \times 3 \approx 3,010.6\text{M} tensor elements, which is roughly 47 times larger than ACIL's storage footprint.
  8. Knowl 8 — Frozen Feature Extractor Limitation in ACIL

    limitation

    ACIL relies on freezing the convolutional feature extractor WCNNW_{\text{CNN}} after the base training phase (phase 0). All subsequent class-incremental learning phases rely entirely on transfer learning through the frozen CNN backbone and random feature expansion.

    This architectural constraint introduces two limitations:

    1. Representation Inflexibility: When incremental classes possess visual patterns or domain distributions substantially different from base classes, the frozen CNN feature extractor cannot adapt its convolutional filters to extract domain-specific representations. This can cause an accuracy gap compared to joint backpropagation models trained on the entire dataset simultaneously.
    2. Memory Scaling of Feature Expansion: Maximizing ACIL's classification capacity requires increasing the feature expansion dimension dfed_{\text{fe}}. Because RFAuM storage scales quadratically as O(dfe2)\mathcal{O}(d_{\text{fe}}^2) and the Woodbury matrix inversion step scales with batch size and feature dimension, expanding beyond dfe=15kd_{\text{fe}} = 15\text{k} is constrained by available GPU RAM (e.g., out-of-memory errors on an 11GB GPU).

Coverage note — None was omitted. All contributed algorithmic components, theoretical formulation (Theorem 3.1), benchmark experimental results across CIFAR-100 and ImageNet, ablation studies, and stated limitations were fully captured.

References

  1. 1.Eden Belouadah, Adrian Popescu, and Ioannis Kanellos. A comprehensive study of class incremental learning algorithms for visual tasks. Neural Networks, 135:38–54, 2021.
  2. 2.Francisco M. Castro, Manuel J. Marin-Jimenez, Nicolas Guil, Cordelia Schmid, and Karteek Alahari. End-to-end incremental learning. In Proceedings of the European Conference on Computer Vision (ECCV), September 2018.
  3. 3.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In 2009 IEEE Conference on Computer Vision and Pattern Recognition, pages 248–255, 2009.
  4. 4.Arthur Douillard, Matthieu Cord, Charles Ollion, Thomas Robert, and Eduardo Valle. Podnet: Pooled outputs distillation for small-tasks incremental learning. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XX 16, pages 86–102. Springer, 2020.
  5. 5.Ping Guo, Michael R Lyu, and NE Mastorakis. Pseudoinverse learning algorithm for feedforward neural networks. Advances in Neural Networks and Applications, pages 321–326, 2001.
  6. 6.Saihui Hou, Xinyu Pan, Chen Change Loy, Zilei Wang, and Dahua Lin. Learning a unified classifier incrementally via rebalancing. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2019.
  7. 7.Heechul Jung, Jeongwoo Ju, Minju Jung, and Junmo Kim. Less-forgetting learning in deep neural networks. arXiv preprint arXiv:1607.00122, 2016.
  8. 8.James Kirkpatrick, Razvan Pascanu, Neil Rabinowitz, Joel Veness, Guillaume Desjardins, Andrei A Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, et al. Overcoming catastrophic forgetting in neural networks. Proceedings of the national academy of sciences, 114(13):3521–3526, 2017.
  9. 9.Zhizhong Li and Derek Hoiem. Learning without forgetting. IEEE Transactions on Pattern Analysis and Machine Intelligence, 40(12):2935–2947, 2018.
  10. 10.Xialei Liu, Marc Masana, Luis Herranz, Joost Van de Weijer, Antonio M. López, and Andrew D. Bagdanov. Rotate your networks: Better weight consolidation and less catastrophic forgetting. In 2018 24th International Conference on Pattern Recognition (ICPR), pages 2262–2268, 2018.
  11. 11.Yaoyao Liu, Bernt Schiele, and Qianru Sun. Adaptive aggregation networks for class-incremental learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 2544–2553, June 2021.
  12. 12.Yaoyao Liu, Bernt Schiele, and Qianru Sun. Rmm: Reinforced memory management for class-incremental learning. Advances in Neural Information Processing Systems, 34, 2021.
  13. 13.Yaoyao Liu, Yuting Su, An-An Liu, Bernt Schiele, and Qianru Sun. Mnemonics training: Multi-class incremental learning without forgetting. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2020.
  14. 14.Cheng-Yaw Low, Jaewoo Park, and Andrew Beng-Jin Teoh. Stacking-based deep neural network: Deep analytic network for pattern classification. IEEE Transactions on Cybernetics, 50(12):5021–5034, 2020.
  15. 15.J. Park and I. W. Sandberg. Universal approximation using radial-basis-function networks. Neural Computation, 3(2):246–257, 1991.
  16. 16.Sylvestre-Alvise Rebuffi, Alexander Kolesnikov, Georg Sperl, and Christoph H. Lampert. icarl: Incremental classifier and representation learning. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), July 2017.
  17. 17.Hanul Shin, Jung Kwon Lee, Jaehong Kim, and Jiwon Kim. Continual learning with deep generative replay. In Proceedings of the 31st International Conference on Neural Information Processing Systems, NIPS’17, page 2994–3003, 2017.
  18. 18.Kar-Ann Toh. Learning from the kernel and the range space. In the Proceedings of the 17th 2018 IEEE Conference on Computer and Information Science, pages 417–422. IEEE, June 2018.
  19. 19.Jue Wang, Ping Guo, and Yanjun Li. Densepilae: a feature reuse pseudoinverse learning algorithm for deep stacked autoencoder. Complex & Intelligent Systems, pages 1–11, 2021.
  20. 20.X. Wang, T. Zhang, and R. Wang. Noniterative deep learning: Incorporating restricted boltzmann machine into multilayer random weight neural networks. IEEE Transactions on Systems, Man, and Cybernetics: Systems, 49(7):1299–1308, 2019.
  21. 21.Chenshe Wu, Luis Herranz, Xialei Liu, Yaxing Wang, Joost van de Weijer, and Bogdan Raducanu. Memory replay gans: learning to generate images from new categories without forgetting. In Conference on Neural Information Processing Systems (NIPS), 2018.
  22. 22.Chenshen Wu, Luis Herranz, Xialei Liu, Yaxing Wang, Joost van de Weijer, and Bogdan Raducanu. Memory replay gans: Learning to generate images from new categories without forgetting. In Proceedings of the 32nd International Conference on Neural Information Processing Systems, NIPS’18, page 5966–5976, 2018.
  23. 23.Yue Wu, Yinpeng Chen, Lijuan Wang, Yuancheng Ye, Zicheng Liu, Yandong Guo, and Yun Fu. Large scale incremental learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2019.
  24. 24.Shipeng Yan, Jiangwei Xie, and Xuming He. DER: Dynamically expandable representation for class incremental learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 3014–3023, June 2021.
  25. 25.Lu Yu, Bartlomiej Twardowski, Xialei Liu, Luis Herranz, Kai Wang, Yongmei Cheng, Shangling Jui, and Joost van de Weijer. Semantic drift compensation for class-incremental learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2020.
  26. 26.Junting Zhang, Jie Zhang, Shalini Ghosh, Dawei Li, Serafettin Tasci, Larry Heck, Heming Zhang, and C.-C. Jay Kuo. Class-incremental learning via deep model consolidation. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), March 2020.
  27. 27.Bolei Zhou, Agata Lapedriza, Aditya Khosla, Aude Oliva, and Antonio Torralba. Places: A 10 million image database for scene recognition. IEEE Transactions on Pattern Analysis and Machine Intelligence, 40(6):1452–1464, 2018.
  28. 28.Huiping Zhuang, Zhiping Lin, and Kar-Ann Toh. Blockwise recursive Moore-Penrose inverse for network learning. IEEE Transactions on Systems, Man, and Cybernetics: Systems, pages 1–14, 2021.
  29. 29.Huiping Zhuang, Zhiping Lin, and Kar-Ann Toh. Correlation projection for analytic learning of a classification network. Neural Processing Letters, pages 1–22, 2021.

Citation

MLA
ZHUANG, H., et al. “ACIL: Analytic Class-Incremental Learning with Absolute Memorization and Privacy Protection”. Advances in Neural Information Processing Systems, vol. 35, 2022, pp. 11602–14, https://proceedings.neurips.cc/paper_files/paper/2022/file/4b74a42fc81fc7ee252f6bcb6e26c8be-Paper-Conference.pdf.
APA
ZHUANG, H., Weng, Z., Wei, H., XIE, R., Toh, K.-A., & Lin, Z. (2022). ACIL: Analytic Class-Incremental Learning with Absolute Memorization and Privacy Protection. Advances in Neural Information Processing Systems, 35, 11602–11614. https://proceedings.neurips.cc/paper_files/paper/2022/file/4b74a42fc81fc7ee252f6bcb6e26c8be-Paper-Conference.pdf
Chicago
ZHUANG, H., Z. Weng, H. Wei, R. XIE, K.-A. Toh, and Z. Lin. 2022. “ACIL: Analytic Class-Incremental Learning with Absolute Memorization and Privacy Protection”. Advances in Neural Information Processing Systems 35: 11602–14. https://proceedings.neurips.cc/paper_files/paper/2022/file/4b74a42fc81fc7ee252f6bcb6e26c8be-Paper-Conference.pdf.
Harvard
ZHUANG, H. et al. (2022) “ACIL: Analytic Class-Incremental Learning with Absolute Memorization and Privacy Protection”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 11602–11614. Available at: https://proceedings.neurips.cc/paper_files/paper/2022/file/4b74a42fc81fc7ee252f6bcb6e26c8be-Paper-Conference.pdf.
Vancouver
1. ZHUANG H, Weng Z, Wei H, XIE R, Toh K-A, Lin Z (2022) ACIL: Analytic Class-Incremental Learning with Absolute Memorization and Privacy Protection. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 11602–11614

BibTeX

@inproceedings{zhuang2022acil,
  title = {ACIL: Analytic Class-Incremental Learning with Absolute Memorization and Privacy Protection},
  author = {ZHUANG, HUIPING and Weng, Zhenyu and Wei, Hongxin and XIE, RENCHUNZI and Toh, Kar-Ann and Lin, Zhiping},
  year = {2022},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {35},
  pages = {11602-11614},
  url = {https://proceedings.neurips.cc/paper_files/paper/2022/file/4b74a42fc81fc7ee252f6bcb6e26c8be-Paper-Conference.pdf}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors