FeCAM: Exploiting the Heterogeneity of Class Distributions in Exemplar-Free Continual Learning

Dipam GoswamiYuyang LiuBartlomiej TwardowskiJoost van de Weijer

article2023NeurIPS101 citations

Proposes a training-free Bayes classifier based on Mahalanobis distance and covariance modeling that effectively captures heterogeneous class distributions to achieve state-of-the-art results in exemplar-free continual learning without updating the backbone network.

Listen

Artificial intelligence systems deployed in real-world environments frequently encounter new classes of data over time. When updated on new tasks without retraining on previous ones, standard deep learning models suffer from catastrophic forgetting, rapidly losing previously acquired knowledge. While storing past samples helps preserve performance, this approach introduces significant data privacy risks and high storage costs. Consequently, exemplar-free continual learning has emerged as a crucial area of research, aiming to incorporate new capabilities sequentially without storing past data or retraining the entire core model.

To address this challenge, the article introduces and evaluates the Feature Covariance-Aware Metric (FeCAM) framework. The primary objective is to demonstrate that modeling the directional spread and feature covariance of data distributions enables highly accurate classification of new classes when the main feature extractor network is frozen after an initial training phase.

To evaluate this approach, the authors conducted extensive empirical experiments across standard machine learning benchmarks, including CIFAR-100, TinyImageNet, ImageNet-Subset, miniImageNet, and CUB-200, spanning both many-shot and few-shot learning scenarios as well as domain-shift benchmarks. The methodology leverages a Bayesian classifier that measures anisotropic Mahalanobis distance rather than standard isotropic Euclidean distance. To ensure numerical stability and comparable distance metrics across classes, the pipeline incorporates feature normalizations (Tukey's transformation), covariance shrinkage to handle cases with limited training examples, and correlation normalization.

The investigation yielded several critical findings. First, the article reveals that while standard Euclidean distance works well for jointly trained data, newly introduced classes exhibit highly heterogeneous, non-spherical feature distributions when processed by a frozen network, making Euclidean metrics suboptimal. Second, FeCAM substantially outperformed existing exemplar-free methods across all benchmarks, improving final accuracy by 4 to 7 percentage points over previous top-performing techniques on many-shot tasks. Third, FeCAM surpassed most exemplar-based methods that store thousands of past images, trailing only complex models that multiply parameter counts by nearly six times. Fourth, the method proved exceptionally efficient, reducing incremental task completion time from 44 minutes down to 6 minutes on benchmark hardware while requiring no iterative classifier retraining.

These findings indicate that organizations can build highly adaptable, sequential machine learning pipelines at a fraction of standard computational costs and without violating privacy regulations. By eliminating the need to store sensitive historical data or continuously retrain deep neural backbones, FeCAM provides a practical path for rapid deployment in resource-constrained and privacy-regulated environments.

Organizations developing streaming or continually updated AI applications should consider adopting covariance-aware Mahalanobis classification in place of standard linear classifiers or nearest-mean Euclidean approaches. Because FeCAM requires no parameter retraining during incremental tasks, it can serve as a drop-in classifier layer for existing frozen backbones or pretrained vision transformer models.

The primary limitation of this approach is its reliance on a high-quality initial feature representation; it performs best when initialized on a substantial base dataset or a strong pretrained model, and performance drops if the initial task contains very few classes. Stakeholders can have high confidence in these results across structured image benchmarks, though future work should assess extending covariance modeling to dynamic scenarios where the underlying feature extractor itself must be continually updated.

Cover for FeCAM: Exploiting the Heterogeneity of Class Distributions in Exemplar-Free Continual Learning

Abstract

Exemplar-free class-incremental learning (CIL) poses several challenges since it prohibits the rehearsal of data from previous tasks and thus suffers from catastrophic forgetting. Recent approaches to incrementally learning the classifier by freezing the feature extractor after the first task have gained much attention. In this paper, we explore prototypical networks for CIL, which generate new class prototypes using the frozen feature extractor and classify the features based on the Euclidean distance to the prototypes. In an analysis of the feature distributions of classes, we show that classification based on Euclidean metrics is successful for jointly trained features. However, when learning from non-stationary data, we observe that the Euclidean metric is suboptimal and that feature distributions are heterogeneous. To address this challenge, we revisit the anisotropic Mahalanobis distance for CIL. In addition, we empirically show that modeling the feature covariance relations is better than previous attempts at sampling features from normal distributions and training a linear classifier. Unlike existing methods, our approach generalizes to both many- and few-shot CIL settings, as well as to domain-incremental settings. Interestingly, without updating the backbone network, our method obtains state-of-the-art results on several standard continual learning benchmarks. Code is available at https://github.com/dipamgoswami/FeCAM.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Proposed Approach
  • 3.1 Motivation
  • 3.2 Bayesian FeCAM Classifier
  • 3.3 On the Suboptimality of Learning Linear Classifier
  • 4 Experiments
  • 4.1 Experimental Setup
  • 4.2 Experimental Results
  • 4.2.1 Many-shot CIL Results
  • 4.2.2 Experiments with pre-trained models
  • 4.2.3 Few-shot CIL Results
  • 4.2.4 Ablation Studies
  • 5 Conclusions
  • References

Knowls

  1. Knowl 1 — FeCAM Bayesian Classifier for Exemplar-Free Class-Incremental Learning

    model/method

    In exemplar-free class-incremental learning (CIL) where the feature extractor ϕ:X→RD\phi: \mathcal{X} \to \mathbb{R}^D is frozen after the initial task, features belonging to new incremental classes exhibit anisotropic and heterogeneous distributions rather than isotropic spherical distributions. Feature Covariance-Aware Metric (FeCAM) classifies samples without training or rehearsal by assuming that feature vectors for each class y∈{1,…,Y}y \in \{1, \dots, Y\} follow a multivariate normal distribution N(μy,Σy)\mathcal{N}(\mu_y, \Sigma_y).

    Under this generative assumption with equal class priors P(Y=y)P(Y=y), the Bayes optimal decision rule assigns an input sample xx to the class minimizing the squared Mahalanobis distance between its Tukey-transformed feature representation ϕ~(x)\tilde{\phi}(x) and the class prototype μ~y\tilde{\mu}_y:

    y∗=arg⁡min⁡y=1,…,YDM(ϕ(x),μy)=arg⁡min⁡y=1,…,Y(ϕ~(x)−μ~y)T(Σy)^s−1(ϕ~(x)−μ~y)y^* = \arg\min_{y=1,\dots,Y} D_M(\phi(x), \mu_y) = \arg\min_{y=1,\dots,Y} (\tilde{\phi}(x) - \tilde{\mu}_y)^T \hat{(\Sigma_y)}_s^{-1} (\tilde{\phi}(x) - \tilde{\mu}_y)

    where ϕ~(x)=ϕ(x)λ\tilde{\phi}(x) = \phi(x)^\lambda (with λ=0.5\lambda = 0.5) is the Tukey-transformed feature vector, μ~y=1∣Xy∣∑x∈Xyϕ~(x)\tilde{\mu}_y = \frac{1}{|X_y|} \sum_{x \in X_y} \tilde{\phi}(x) is the class prototype computed from the set of training examples XyX_y of class yy, and (Σy)^s−1\hat{(\Sigma_y)}_s^{-1} is the inverse of the class covariance matrix after covariance shrinkage and correlation normalization.

  2. Knowl 2 — Correlation Matrix Normalization for Heterogeneous Class Covariances

    equation

    Because representations of old classes (used to train the feature extractor) and new classes (unseen during feature extractor training) have substantial differences in variance magnitudes, Mahalanobis distances computed with raw class covariance matrices Σy\Sigma_y have disparate scaling factors and are not directly comparable across classes. FeCAM eliminates scale discrepancies by normalizing each class covariance matrix into a correlation matrix with unit diagonal entries:

    Σ^y(i,j)=Σy(i,j)σy(i)σy(j)\hat{\Sigma}_y(i, j) = \frac{\Sigma_y(i, j)}{\sigma_y(i)\sigma_y(j)}

    σy(i)=Σy(i,i),σy(j)=Σy(j,j)\sigma_y(i) = \sqrt{\Sigma_y(i, i)}, \quad \sigma_y(j) = \sqrt{\Sigma_y(j, j)}

    where Σy(i,j)\Sigma_y(i, j) is the sample covariance between feature dimensions ii and jj (i,j∈{1,…,D}i, j \in \{1, \dots, D\}) for class yy, and σy(i)\sigma_y(i) is the standard deviation along dimension ii.

  3. Knowl 3 — Covariance Shrinkage for Rank-Deficient Feature Distributions

    equation

    When the number of training samples ∣Xy∣|X_y| available for class yy is fewer than the feature dimension DD, the sample covariance matrix Σ∈RD×D\Sigma \in \mathbb{R}^{D \times D} is rank-deficient and singular. FeCAM applies linear shrinkage to estimate a full-rank, invertible covariance matrix Σs\Sigma_s:

    Σs=Σ+γ1V1I+γ2V2(1−I)\Sigma_s = \Sigma + \gamma_1 V_1 I + \gamma_2 V_2 (1 - I)

    where I∈RD×DI \in \mathbb{R}^{D \times D} is the identity matrix, (1−I)(1 - I) is a matrix with zeros on the diagonal and ones off the diagonal, V1=1DTr⁡(Σ)V_1 = \frac{1}{D}\operatorname{Tr}(\Sigma) is the mean diagonal variance of Σ\Sigma, and V2=1D(D−1)∑i≠jΣ(i,j)V_2 = \frac{1}{D(D-1)} \sum_{i \ne j} \Sigma(i, j) is the mean off-diagonal covariance of Σ\Sigma. The regularization coefficients are set to γ1=1,γ2=1\gamma_1 = 1, \gamma_2 = 1 in many-shot settings, and γ1=100,γ2=100\gamma_1 = 100, \gamma_2 = 100 in few-shot settings.

  4. Knowl 4 — Tukey's Ladder of Powers Feature Gaussianization

    equation

    To reduce distribution skewness and better satisfy the multivariate normality assumption of the Bayesian classifier, extracted deep feature representations ϕ(x)∈RD\phi(x) \in \mathbb{R}^D are normalized using Tukey's Ladder of Powers transformation:

    ϕ~(x)={ϕ(x)λif λ≠0log⁡(ϕ(x))if λ=0\tilde{\phi}(x) = \begin{cases} \phi(x)^\lambda & \text{if } \lambda \ne 0 \\ \log(\phi(x)) & \text{if } \lambda = 0 \end{cases}

    where λ∈R\lambda \in \mathbb{R} is a transformation hyperparameter applied element-wise. In FeCAM, λ=0.5\lambda = 0.5 (square-root transformation) is used across all benchmarks prior to computing prototypes and covariance matrices.

  5. Knowl 5 — Incremental Common Covariance Matrix Update

    equation

    When storage constraints prevent maintaining a separate covariance matrix for each individual class, a single common covariance matrix Σ1:t∈RD×D\Sigma^{1:t} \in \mathbb{R}^{D \times D} representing all classes seen up to task tt can be updated incrementally via running average:

    Σ1:t=Σ1:t−1⋅∣Y1:t−1∣∣Y1:t∣+Σt⋅∣Y1:t∣−∣Y1:t−1∣∣Y1:t∣\Sigma^{1:t} = \Sigma^{1:t-1} \cdot \frac{|Y^{1:t-1}|}{|Y^{1:t}|} + \Sigma^t \cdot \frac{|Y^{1:t}| - |Y^{1:t-1}|}{|Y^{1:t}|}

    where Σt\Sigma^t is the empirical covariance matrix calculated across all sample features in task tt, ∣Y1:t−1∣|Y^{1:t-1}| is the total count of classes seen up to task t−1t-1, and ∣Y1:t∣|Y^{1:t}| is the total count of classes seen up to task tt.

  6. Knowl 6 — Suboptimality of Linear Classifiers on Gaussian-Sampled Features

    theoretical result

    Training a linear classifier (such as logistic regression) on synthetic features sampled from class Gaussian distributions N(μy,Σy)\mathcal{N}(\mu_y, \Sigma_y) assumes that the optimal decision boundaries between classes are hyperplanes. However, linear decision boundaries are Bayes-optimal only when all class covariance matrices are equal (homoscedasticity). When class covariance matrices Σy\Sigma_y are heterogeneous and anisotropic, the true Bayes-optimal decision boundary between classes ii and jj is quadratic:

    {x∈RD∣(x−μi)TΣi−1(x−μi)+log⁡∣Σi∣=(x−μj)TΣj−1(x−μj)+log⁡∣Σj∣}\{x \in \mathbb{R}^D \mid (x - \mu_i)^T \Sigma_i^{-1} (x - \mu_i) + \log|\Sigma_i| = (x - \mu_j)^T \Sigma_j^{-1} (x - \mu_j) + \log|\Sigma_j|\}

    Because a linear classifier cannot model quadratic surfaces, directly evaluating the Bayes Mahalanobis distance outperforms linear classifier training on synthetic samples (e.g., yielding 70.9% average accuracy versus 67.9% when sampling 20,000 features per class on CIFAR-100 T=5T=5).

  7. Knowl 7 — Exemplar-Free Many-Shot CIL Benchmark Results

    data/table

    The performance of FeCAM in exemplar-free many-shot class-incremental learning was evaluated on CIFAR-100, TinyImageNet, and ImageNet-Subset using a ResNet-18 backbone trained on the initial task and frozen thereafter. Top-1 average incremental accuracy (%) across 5, 10, and 20 incremental tasks (TT) demonstrates that FeCAM with per-class covariance (Σy\Sigma_y) or common covariance (Σ1:t\Sigma^{1:t}) outperforms prior exemplar-free methods and Euclidean Nearest Class Mean (Eucl-NCM).

    Method CIFAR-100 TinyImageNet ImageNet-Subset
    T=5T=5 T=10T=10 T=20T=20 T=5T=5 T=10T=10 T=20T=20 T=5T=5 T=10T=10 T=20T=20
    EWC 24.5 21.2 15.9 18.8 15.8 12.4 - 20.4 -
    LwF-MC 45.9 27.4 20.1 29.1 23.1 17.4 - 31.2 -
    DeeSIL 60.0 50.6 38.1 49.8 43.9 34.1 67.9 60.1 50.5
    MUC 49.4 30.2 21.3 32.6 26.6 21.9 - 35.1 -
    SDC 56.8 57.0 58.9 - - - - 61.2 -
    PASS 63.5 61.8 58.1 49.6 47.3 42.1 64.4 61.8 51.3
    IL2A 66.0 60.3 57.9 47.3 44.7 40.0 - - -
    SSRE 65.9 65.0 61.7 50.4 48.9 48.2 - 67.7 -
    FeTrIL 67.6 66.6 63.5 55.4 54.3 53.0 73.1 71.9 69.1
    Eucl-NCM 64.8 64.6 61.5 54.1 53.8 53.6 72.2 72.0 68.4
    FeCAM (ours) - Σ1:t\Sigma^{1:t} 68.8 68.6 67.4 56.0 55.7 55.5 75.8 75.6 73.5
    FeCAM (ours) - Σy\Sigma_y 70.9 70.8 69.4 59.6 59.4 59.3 78.3 78.2 75.1
    Upper Bound (Joint) 79.2 79.2 79.2 66.1 66.1 66.1 84.7 84.7 84.7
  8. Knowl 8 — Continual Learning Performance with Pretrained Vision Transformers

    data/table

    Using frozen feature representations extracted from a Vision Transformer (ViT-B/16) pretrained on ImageNet-21k, FeCAM was evaluated on class-incremental benchmarks (Split-CIFAR100 and Split-ImageNet-R) and the domain-incremental benchmark CoRe50. In the domain-incremental setup, a single covariance matrix per class is maintained across domains and updated in each new domain by averaging the previous domain's matrix with the current domain's matrix.

    Method Split-CIFAR100 Split-ImageNet-R CoRe50
    Avg Acc (%) Avg Acc (%) Test Acc (%)
    FT-frozen 17.7 39.5 -
    FT 33.6 28.9 -
    EWC 47.0 35.0 74.8
    LwF 60.7 38.5 75.5
    L2P 83.8 61.6 78.3
    NCM 83.7 55.7 85.4
    FeCAM (ours) 85.7 63.7 89.9
    Joint 90.9 79.1 -
  9. Knowl 9 — Ablation Study of FeCAM Classifier Pipeline

    data/table

    An ablation study on 5-task many-shot CIL settings for CIFAR-100 and ImageNet-Subset shows the contribution of each module in FeCAM: Tukey's transformation, covariance shrinkage, covariance structure (full vs. diagonal), and correlation matrix normalization.

    Distance Cov. Matrix Tukey Shrinkage Norm. CIFAR-100 (T=5T=5) ImageNet-Subset (T=5T=5)
    Last Acc Avg Acc Last Acc Avg Acc
    Euclidean - - - 51.6 64.8 60.0 72.2
    Euclidean - ✓ - - 54.4 66.6 66.2 73.6
    Mahalanobis Full 14.6 29.7 33.5 45.1
    Mahalanobis Full ✓ 20.6 36.2 54.0 65.6
    Mahalanobis Full ✓ 44.6 59.3 39.9 56.9
    Mahalanobis Full ✓ ✓ 52.1 62.8 56.5 67.3
    Mahalanobis Diagonal ✓ ✓ 55.2 66.9 64.0 74.1
    Mahalanobis Full ✓ ✓ 55.4 65.9 58.1 68.5
    Mahalanobis Full ✓ ✓ ✓ 62.1 70.9 70.9 78.3
  10. Knowl 10 — Limitations in Small Initial Task Scenarios

    limitation

    FeCAM relies on a frozen feature extractor and does not adapt feature representations after task 1. Consequently, its efficacy depends heavily on learning a rich representation from a sufficiently large initial task (e.g., 50% of classes) or from a strong pretrained backbone. When initialized with a small first task (e.g., starting with only 20 classes on CIFAR-100 and adding 20 classes per task), average incremental accuracy decreases (62.3% on CIFAR-100 and 66.4% on ImageNet-Subset). Applying the method to settings with continual feature representation updates would require modeling representation drift for both class means and class covariance matrices.

Coverage note — None was omitted; all key theoretical observations, mathematical formulations, algorithmic components, experimental results (MSCIL, pretrained ViT, FSCIL), ablations, and stated limitations are included.

References

  1. 1.Afra Feyza Akyürek, Ekin Akyürek, Derry Wijaya, and Jacob Andreas. Subspace regularizers for few-shot class incremental learning. In International Conference on Learning Representations (ICLR), 2022.
  2. 2.Nader Asadi, MohammadReza Davari, Sudhir Mudur, Rahaf Aljundi, and Eugene Belilovsky. Prototype-sample relation distillation: towards replay-free continual learning. In International Conference on Machine Learning (ICML), 2023.
  3. 3.Eden Belouadah and Adrian Popescu. Deesil: Deep-shallow incremental learning. In European Conference on Computer Vision (ECCV) Workshops, 2018.
  4. 4.Luca Bertinetto, Joao F. Henriques, Philip Torr, and Andrea Vedaldi. Meta-learning with differentiable closed-form solvers. In International Conference on Learning Representations (ICLR), 2019.
  5. 5.Prashant Shivaram Bhat, Bahram Zonooz, and Elahe Arani. Consistency is the key to further mitigating catastrophic forgetting in continual learning. In Conference on Lifelong Learning Agents (CoLLAs), 2022.
  6. 6.Prashant Shivaram Bhat, Bahram Zonooz, and Elahe Arani. Task-aware information routing from common representation space in lifelong learning. In International Conference on Learning Representations (ICLR), 2023.
  7. 7.Francisco M Castro, Manuel J Marín-Jiménez, Nicolás Guil, Cordelia Schmid, and Karteek Alahari. End-to-end incremental learning. In European Conference on Computer Vision (ECCV), 2018.
  8. 8.Zhixiang Chi, Li Gu, Huan Liu, Yang Wang, Yuanhao Yu, and Jin Tang. Metafscil: a meta-learning approach for few-shot class incremental learning. In Conference on Computer Vision and Pattern Recognition (CVPR), 2022.
  9. 9.Matthias De Lange, Rahaf Aljundi, Marc Masana, Sarah Parisot, Xu Jia, Aleš Leonardis, Gregory Slabaugh, and Tinne Tuytelaars. A continual learning survey: Defying forgetting in classification tasks. Transactions on Pattern Analysis and Machine Intelligence (T-PAMI), 2021.
  10. 10.Matthias De Lange and Tinne Tuytelaars. Continual prototype evolution: Learning online from non-stationary data streams. In International Conference on Computer Vision (ICCV), 2021.
  11. 11.Roy De Maesschalck, Delphine Jouan-Rimbaud, and Désiré L Massart. The mahalanobis distance. Chemometrics and intelligent laboratory systems, 2000.
  12. 12.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In Conference on Computer Vision and Pattern Recognition (CVPR), 2009.
  13. 13.Akshay Raj Dhamija, Touqeer Ahmad, Jonathan Schwan, Mohsen Jafarzadeh, Chunchun Li, and Terrance E Boult. Self-supervised features improve open-world learning. arXiv preprint arXiv:2102.07848, 2021.
  14. 14.Prithviraj Dhar, Rajat Vikram Singh, Kuan-Chuan Peng, Ziyan Wu, and Rama Chellappa. Learning without memorizing. In Conference on Computer Vision and Pattern Recognition (CVPR), 2019.
  15. 15.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition at scale. In International Conference on Learning Representations (ICLR), 2021.
  16. 16.Arthur Douillard, Matthieu Cord, Charles Ollion, Thomas Robert, and Eduardo Valle. Podnet: Pooled outputs distillation for small-tasks incremental learning. In European Conference on Computer Vision (ECCV), 2020.
  17. 17.Samantha Guerriero, Barbara Caputo, and Thomas Mensink. Deepncm: Deep nearest class mean classifiers. International Conference on Learning Representations Workshop (ICLR-W), 2018.
  18. 18.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Conference on Computer Vision and Pattern Recognition (CVPR), 2016.
  19. 19.Dan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath, Frank Wang, Evan Dorundo, Rahul Desai, Tyler Zhu, Samyak Parajuli, Mike Guo, et al. The many faces of robustness: A critical analysis of out-of-distribution generalization. In International Conference on Computer Vision (ICCV), 2021.
  20. 20.Geoffrey Hinton, Oriol Vinyals, Jeff Dean, et al. Distilling the knowledge in a neural network. arXiv preprint arXiv:1503.02531, 2015.
  21. 21.Saihui Hou, Xinyu Pan, Chen Change Loy, Zilei Wang, and Dahua Lin. Learning a unified classifier incrementally via rebalancing. In Conference on Computer Vision and Pattern Recognition (CVPR), 2019.
  22. 22.Paul Janson, Wenxuan Zhang, Rahaf Aljundi, and Mohamed Elhoseiny. A simple baseline that questions the use of pretrained-models in continual learning. In NeurIPS 2022 Workshop on Distribution Shifts: Connecting Methods and Applications, 2022.
  23. 23.Kishaan Jeeveswaran, Prashant Bhat, Bahram Zonooz, and Elahe Arani. Birt: Bio-inspired replay in vision transformers for continual learning. In International Conference on Machine Learning (ICML), 2023.
  24. 24.Gyuhak Kim, Bing Liu, and Zixuan Ke. A multi-head model for continual learning via out-of-distribution replay. In Conference on Lifelong Learning Agents (CoLLAs), 2022.
  25. 25.James Kirkpatrick, Razvan Pascanu, Neil Rabinowitz, Joel Veness, Guillaume Desjardins, Andrei A Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, et al. Overcoming catastrophic forgetting in neural networks. Proceedings of the National Academy of Sciences (PNAS), 2017.
  26. 26.Alex Krizhevsky, Geoffrey Hinton, et al. Learning multiple layers of features from tiny images. 2009.
  27. 27.Shakti Kumar and Hussain Zaidi. Gdc-generalized distribution calibration for few-shot learning. arXiv preprint arXiv:2204.05230, 2022.
  28. 28.Ya Le and Xuan Yang. Tiny imagenet visual recognition challenge. CS 231N, 2015.
  29. 29.Kimin Lee, Kibok Lee, Honglak Lee, and Jinwoo Shin. A simple unified framework for detecting out-of-distribution samples and adversarial attacks. Advances in Neural Information Processing Systems (NeurIPS), 2018.
  30. 30.Zhizhong Li and Derek Hoiem. Learning without forgetting. Transactions on Pattern Analysis and Machine Intelligence (T-PAMI), 2017.
  31. 31.Huan Liu, Li Gu, Zhixiang Chi, Yang Wang, Yuanhao Yu, Jun Chen, and Jin Tang. Few-shot class-incremental learning via entropy-regularized data-free replay. In European Conference on Computer Vision (ECCV), 2022.
  32. 32.Yu Liu, Sarah Parisot, Gregory Slabaugh, Xu Jia, Ales Leonardis, and Tinne Tuytelaars. More classifiers, less forgetting: A generic multi-classifier paradigm for incremental learning. In European Conference on Computer Vision (ECCV), 2020.
  33. 33.Yuyang Liu, Yang Cong, Dipam Goswami, Xialei Liu, and Joost van de Weijer. Augmented box replay: Overcoming foreground shift for incremental object detection. In International Conference on Computer Vision (ICCV), 2023.
  34. 34.Vincenzo Lomonaco and Davide Maltoni. Core50: a new dataset and benchmark for continuous object recognition. In Conference on Robot Learning (CoRL), 2017.
  35. 35.Chunwei Ma, Zhanghexuan Ji, Ziyun Huang, Yan Shen, Mingchen Gao, and Jinhui Xu. Progressive voronoi diagram subdivision enables accurate data-free class-incremental learning. In International Conference on Learning Representations (ICLR), 2023.
  36. 36.Marc Masana, Xialei Liu, Bartlomiej Twardowski, Mikel Menta, Andrew D Bagdanov, and Joost van de Weijer. Class-incremental learning: survey and performance evaluation. Transactions on Pattern Analysis and Machine Intelligence (T-PAMI), 2022.
  37. 37.Michael McCloskey and Neal J Cohen. Catastrophic interference in connectionist networks: The sequential learning problem. In Psychology of learning and motivation. Elsevier, 1989.
  38. 38.Thomas Mensink, Jakob Verbeek, Florent Perronnin, and Gabriela Csurka. Distance-based image classification: Generalizing to new classes at near-zero cost. IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), 2013.
  39. 39.Tsendsuren Munkhdalai and Hong Yu. Meta networks. In International Conference on Machine Learning (ICML), 2017.
  40. 40.Cuong V Nguyen, Yingzhen Li, Thang D Bui, and Richard E Turner. Variational continual learning. In International Conference on Learning Representations (ICLR), 2018.
  41. 41.Aristeidis Panos, Yuriko Kobe, Daniel Olmeda Reino, Rahaf Aljundi, and Richard E. Turner. First session adaptation: A strong replay-free baseline for class-incremental learning. In International Conference on Computer Vision (ICCV), 2023.
  42. 42.Can Peng, Kun Zhao, Tianren Wang, Meng Li, and Brian C Lovell. Few-shot class-incremental learning from an open-set perspective. In European Conference on Computer Vision (ECCV), 2022.
  43. 43.Grégoire Petit, Adrian Popescu, Hugo Schindler, David Picard, and Bertrand Delezoide. Fetril: Feature translation for exemplar-free class-incremental learning. In Winter Conference on Applications of Computer Vision (WACV), 2023.
  44. 44.Sylvestre-Alvise Rebuffi, Alexander Kolesnikov, Georg Sperl, and Christoph H Lampert. icarl: Incremental classifier and representation learning. In Conference on Computer Vision and Pattern Recognition (CVPR), 2017.
  45. 45.Tal Ridnik, Emanuel Ben-Baruch, Asaf Noy, and Lihi Zelnik-Manor. Imagenet-21k pretraining for the masses. In Conference on Neural Information Processing Systems Datasets and Benchmarks Track, 2021.
  46. 46.Anthony Robins. Catastrophic forgetting, rehearsal and pseudorehearsal. Connection Science, 1995.
  47. 47.Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, et al. Imagenet large scale visual recognition challenge. International journal of computer vision, 2015.
  48. 48.Andrei A Rusu, Neil C Rabinowitz, Guillaume Desjardins, Hubert Soyer, James Kirkpatrick, Koray Kavukcuoglu, and Raia Hadsell. Meta-learning with latent embedding optimization. In International Conference on Learning Representations (ICLR), 2019.
  49. 49.Ali Sharif Razavian, Hossein Azizpour, Josephine Sullivan, and Stefan Carlsson. Cnn features off-the-shelf: an astounding baseline for recognition. In Conference on Computer Vision and Pattern Recognition (CVPR) workshops, 2014.
  50. 50.Christian Simon, Masoud Faraki, Yi-Hsuan Tsai, Xiang Yu, Samuel Schulter, Yumin Suh, Mehrtash Harandi, and Manmohan Chandraker. On generalizing beyond domains in cross-domain continual learning. In Conference on Computer Vision and Pattern Recognition (CVPR), 2022.
  51. 51.James Smith, Yen-Chang Hsu, Jonathan Balloch, Yilin Shen, Hongxia Jin, and Zsolt Kira. Always be dreaming: A new approach for data-free class-incremental learning. In International Conference on Computer Vision (ICCV), 2021.
  52. 52.Andreas Peter Steiner, Alexander Kolesnikov, Xiaohua Zhai, Ross Wightman, Jakob Uszkoreit, and Lucas Beyer. How to train your vit? data, augmentation, and regularization in vision transformers. Transactions on Machine Learning Research, 2022.
  53. 53.Xiaoyu Tao, Xiaopeng Hong, Xinyuan Chang, Songlin Dong, Xing Wei, and Yihong Gong. Few-shot class-incremental learning. In Conference on Computer Vision and Pattern Recognition (CVPR), 2020.
  54. 54.Eleni Triantafillou, Tyler Zhu, Vincent Dumoulin, Pascal Lamblin, Utku Evci, Kelvin Xu, Ross Goroshin, Carles Gelada, Kevin Swersky, Pierre-Antoine Manzagol, and Hugo Larochelle. Meta-dataset: A dataset of datasets for learning to learn from few examples. In International Conference on Learning Representations (ICLR), 2020.
  55. 55.John W. Tukey. Exploratory data analysis. Addison-Wesley series in behavioral science : quantitative methods. Addison-Wesley, 1977.
  56. 56.Gido M Van de Ven and Andreas S Tolias. Three scenarios for continual learning. arXiv preprint arXiv:1904.07734, 2019.
  57. 57.John Van Ness. On the dominance of non-parametric Bayes rule discriminant algorithms in high dimensions. Pattern Recognition, 1980.
  58. 58.Oriol Vinyals, Charles Blundell, Timothy Lillicrap, Daan Wierstra, et al. Matching networks for one shot learning. Advances in Neural Information Processing Systems (NeurIPS), 2016.
  59. 59.Catherine Wah, Steve Branson, Peter Welinder, Pietro Perona, and Serge Belongie. The caltech-ucsd birds-200-2011 dataset. 2011.
  60. 60.Fu-Yun Wang, Da-Wei Zhou, Han-Jia Ye, and De-Chuan Zhan. Foster: Feature boosting and compression for class-incremental learning. In European Conference on Computer Vision (ECCV), 2022.
  61. 61.Liyuan Wang, Xingxing Zhang, Hang Su, and Jun Zhu. A comprehensive survey of continual learning: Theory, method and application. arXiv preprint arXiv:2302.00487, 2023.
  62. 62.Zifeng Wang, Zizhao Zhang, Sayna Ebrahimi, Ruoxi Sun, Han Zhang, Chen-Yu Lee, Xiaoqi Ren, Guolong Su, Vincent Perot, Jennifer Dy, et al. Dualprompt: Complementary prompting for rehearsal-free continual learning. In European Conference on Computer Vision (ECCV), 2022.
  63. 63.Zifeng Wang, Zizhao Zhang, Chen-Yu Lee, Han Zhang, Ruoxi Sun, Xiaoqi Ren, Guolong Su, Vincent Perot, Jennifer Dy, and Tomas Pfister. Learning to prompt for continual learning. In Conference on Computer Vision and Pattern Recognition (CVPR), 2022.
  64. 64.Yue Wu, Yinpeng Chen, Lijuan Wang, Yuancheng Ye, Zicheng Liu, Yandong Guo, and Yun Fu. Large scale incremental learning. In Conference on Computer Vision and Pattern Recognition (CVPR), 2019.
  65. 65.Shiming Xiang, Feiping Nie, and Changshui Zhang. Learning a mahalanobis distance metric for data clustering and classification. Pattern recognition, 2008.
  66. 66.Shipeng Yan, Jiangwei Xie, and Xuming He. Der: Dynamically expandable representation for class incremental learning. In Conference on Computer Vision and Pattern Recognition (CVPR), 2021.
  67. 67.Shuo Yang, Lu Liu, and Min Xu. Free lunch for few-shot learning: Distribution calibration. In International Conference on Learning Representations (ICLR), 2021.
  68. 68.Lu Yu, Bartlomiej Twardowski, Xialei Liu, Luis Herranz, Kai Wang, Yongmei Cheng, Shangling Jui, and Joost van de Weijer. Semantic drift compensation for class-incremental learning. In Conference on Computer Vision and Pattern Recognition (CVPR), 2020.
  69. 69.Sergey Zagoruyko and Nikos Komodakis. Wide residual networks. In British Machine Vision Conference (BMVC), 2016.
  70. 70.Chi Zhang, Nan Song, Guosheng Lin, Yun Zheng, Pan Pan, and Yinghui Xu. Few-shot incremental learning with continually evolved classifiers. In Conference on Computer Vision and Pattern Recognition (CVPR), 2021.
  71. 71.Bowen Zhao, Xi Xiao, Guojun Gan, Bin Zhang, and Shu-Tao Xia. Maintaining discrimination and fairness in class incremental learning. In Conference on Computer Vision and Pattern Recognition (CVPR), 2020.
  72. 72.Da-Wei Zhou, Fu-Yun Wang, Han-Jia Ye, Liang Ma, Shiliang Pu, and De-Chuan Zhan. Forward compatible few-shot class-incremental learning. In Conference on Computer Vision and Pattern Recognition (CVPR), 2022.
  73. 73.Da-Wei Zhou, Fu-Yun Wang, Han-Jia Ye, and De-Chuan Zhan. Pycil: a python toolbox for class-incremental learning. SCIENCE CHINA Information Sciences, 2023.
  74. 74.Da-Wei Zhou, Qi-Wei Wang, Zhi-Hong Qi, Han-Jia Ye, De-Chuan Zhan, and Ziwei Liu. Deep class-incremental learning: A survey. arXiv preprint arXiv:2302.03648, 2023.
  75. 75.Da-Wei Zhou, Qi-Wei Wang, Han-Jia Ye, and De-Chuan Zhan. A model or 603 exemplars: Towards memory-efficient class-incremental learning. In International Conference on Learning Representations (ICLR), 2022.
  76. 76.Da-Wei Zhou, Han-Jia Ye, Liang Ma, Di Xie, Shiliang Pu, and De-Chuan Zhan. Few-shot class-incremental learning by sampling multi-phase tasks. IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), 2022.
  77. 77.Da-Wei Zhou, Han-Jia Ye, and De-Chuan Zhan. Co-transport for class-incremental learning. In ACM International Conference on Multimedia, 2021.
  78. 78.Fei Zhu, Zhen Cheng, Xu-Yao Zhang, and Cheng-lin Liu. Class-incremental learning via dual augmentation. Advances in Neural Information Processing Systems (NeurIPS), 2021.
  79. 79.Fei Zhu, Xu-Yao Zhang, Chuang Wang, Fei Yin, and Cheng-Lin Liu. Prototype augmentation and self-supervision for incremental learning. In Conference on Computer Vision and Pattern Recognition (CVPR), 2021.
  80. 80.Kai Zhu, Wei Zhai, Yang Cao, Jiebo Luo, and Zheng-Jun Zha. Self-sustaining representation expansion for non-exemplar class-incremental learning. In Conference on Computer Vision and Pattern Recognition (CVPR), 2022.

Citation

MLA
Goswami, D., et al. “FeCAM: Exploiting the Heterogeneity of Class Distributions in Exemplar-Free Continual Learning”. arXiv, 2023, http://arxiv.org/abs/2309.14062v3.
APA
Goswami, D., Liu, Y., Twardowski, B., & Weijer, J. van . de . (2023). FeCAM: Exploiting the Heterogeneity of Class Distributions in Exemplar-Free Continual Learning. arXiv. http://arxiv.org/abs/2309.14062v3
Chicago
Goswami, D., Y. Liu, B. Twardowski, and J. van . de . Weijer. 2023. “FeCAM: Exploiting the Heterogeneity of Class Distributions in Exemplar-Free Continual Learning”. arXiv. http://arxiv.org/abs/2309.14062v3.
Harvard
Goswami, D. et al. (2023) “FeCAM: Exploiting the Heterogeneity of Class Distributions in Exemplar-Free Continual Learning”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2309.14062v3.
Vancouver
1. Goswami D, Liu Y, Twardowski B, Weijer J van de (2023) FeCAM: Exploiting the Heterogeneity of Class Distributions in Exemplar-Free Continual Learning. arXiv

BibTeX

@article{goswami2023fecam,
  title = {FeCAM: Exploiting the Heterogeneity of Class Distributions in Exemplar-Free Continual Learning},
  author = {Goswami, Dipam and Liu, Yuyang and Twardowski, Bartłomiej and Weijer, Joost van de},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2309.14062v3},
  eprint = {2309.14062}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors