Probing Representation Forgetting in Supervised and Unsupervised Continual Learning

MohammadReza DavariNader AsadiSudhir P. MudurRahaf AljundiEugene Belilovsky

article2022CVPR94 citations

Reveals through linear probing that neural networks retain substantially more past-task information during continual learning than standard accuracy metrics suggest, enabling a competitive rehearsal-free method based on supervised contrastive learning and class prototypes.

Listen

Deploying artificial intelligence systems across evolving real-world tasks often triggers catastrophic forgetting, where deep neural networks abruptly lose previously acquired knowledge when trained on new data. Organizations navigating dynamic environments must frequently update models, but traditional approaches assume that a drop in final classification accuracy means the underlying learned features are permanently destroyed. This assumption forces practitioners to rely on complex, computationally expensive, and memory-heavy continual learning algorithms to preserve model performance.

The article evaluates whether the internal feature representations of neural networks actually degrade as severely as end-task accuracy metrics suggest during sequential learning. The main objective is to measure true representation forgetting across standard continual learning benchmarks and establish how factors like model capacity, loss functions, and evaluation methods affect knowledge retention.

To conduct this evaluation, the analysis implements linear probing—a technique that trains an optimal linear classifier on frozen intermediate activations—to directly quantify the quality of preserved representations across sequential tasks. The researchers benchmark standard continual learning strategies against naive finetuning without forgetting controls across image datasets, including CIFAR and ImageNet variants, across sequences spanning up to 200 tasks. They test both supervised loss functions (standard cross-entropy and supervised contrastive learning) and unsupervised loss functions (such as SimCLR), while varying neural network architectures across different widths and depths in offline and online training settings.

The findings reveal that naive finetuning experiences far less representation forgetting than conventional task accuracy metrics indicate. When measured by linear probes, standard finetuning retains rich feature structures that can match or exceed specialized continual learning algorithms like Learning without Forgetting, especially in long task sequences. Second, training with supervised contrastive loss preserves representations remarkably well, showing stable performance or even positive transfer to early tasks over extended sequences. Third, increasing model capacity—particularly network width—significantly reduces representation forgetting, contrary to prior assumptions that scaling model size without pre-training offers no benefit. Finally, depth-wise evaluations demonstrate that representation forgetting is concentrated almost entirely in the uppermost classification layers, while lower network blocks remain virtually intact and retain their general utility.

These results show that neural network representations are fundamentally more robust to changing data streams than widely believed. Traditional task accuracy conflates superficial feature transformations, such as weight permutations, with actual information loss. Consequently, engineering organizations can avoid the steep compute and memory overhead of complex replay buffers by separating internal representation learning from task classification. A practical, low-cost approach combining supervised contrastive finetuning with nearest-mean-of-exemplar class prototypes achieves competitive recovery on past tasks using only five samples per class at test time, cutting computational training overhead by approximately half compared to standard rehearsal methods.

Decision-makers should re-evaluate their continual learning workflows by tracking internal representation quality via linear probes rather than relying exclusively on observed end-task accuracy. Teams deploying continually updated vision models should consider adopting contrastive representation learning paired with lightweight prototype classifiers for fast recovery on previous tasks. Because this investigation focuses on task-incremental scenarios with defined boundaries, teams should pilot these contrastive strategies on domain-specific data and conduct further testing in boundary-free, class-incremental settings before full-scale operational rollout.

Cover for Probing Representation Forgetting in Supervised and Unsupervised Continual Learning

Abstract

Continual Learning (CL) research typically focuses on tackling the phenomenon of catastrophic forgetting in neural networks. Catastrophic forgetting is associated with an abrupt loss of knowledge previously learned by a model when the task, or more broadly the data distribution, being trained on changes. In supervised learning problems this forgetting, resulting from a change in the model’s representation, is typically measured or observed by evaluating the decrease in old task performance. However, a model’s representation can change without losing knowledge about prior tasks. In this work we consider the concept of representation forgetting, observed by using the difference in performance of an optimal linear classifier before and after a new task is introduced. Using this tool we revisit a number of standard continual learning benchmarks and observe that, through this lens, model representations trained without any explicit control for forgetting often experience small representation forgetting and can sometimes be comparable to methods which explicitly control for forgetting, especially in longer task sequences. We also show that representation forgetting can lead to new insights on the effect of model capacity and loss function used in continual learning. Based on our results, we show that a simple yet competitive approach is to learn representations continually with standard supervised contrastive learning while constructing prototypes of class samples when queried on old samples.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Background and Methods
  • 3.1. Linear Probes for Representation Forgetting
  • 3.2. Centered Kernel Alignment
  • 3.3. Supervised and Unsupervised Contrastive Loss
  • 3.4. Exemplars and Fast Remembering
  • 4. Experiments
  • 4.1. Observed vs LP accuracy
  • 4.2. Effects of Increased Model Capacity
  • 4.3. Low-Cost Remembering with SupCon
  • 4.4. Depth-wise Probes and Comparison to CKA
  • 5. Conclusion
  • References

Knowls

  1. Knowl 1 — Linear Probe Evaluation of Representation Forgetting

    definition

    In continual learning (CL), catastrophic forgetting is traditionally quantified via observed accuracy, which evaluates a model's previous task performance using fixed, previously trained classification heads. However, changes in classification heads (e.g., coordinate permutations or weight drift) can cause severe observed accuracy degradation even when the underlying feature representation retains all task-relevant information.

    To decouple representation degradation from classifier mismatch, representation forgetting is defined as the change in the performance of an optimal linear classifier evaluated on frozen intermediate or final feature activations across task increments.

    Let fθif_{\theta_i} denote a feature extractor parameterized by θi\theta_i obtained after training step ii in a task sequence. For task ii with training dataset (Xi,Yi)(X_i, Y_i), the optimal linear classifier weights Wi∗W_i^* are obtained via convex optimization:

    Wi∗=arg⁡min⁡WL(Wfθi(Xi),Yi)W_i^* = \arg\min_W \mathcal{L}\left(W f_{\theta_i}(X_i), Y_i\right)

    where L\mathcal{L} is a standard classification loss (such as cross-entropy). To measure representation forgetting on task aa between training state θa\theta_a and a subsequent state θb\theta_b (b>ab > a), the metric compares the performance TT (e.g., test accuracy) of the optimal linear probe trained on the respective frozen representations:

    ΔTa→b=T(Wa∗fθa(Xa),Ya)−T(Wb∗fθb(Xa),Ya)\Delta T_{a \to b} = T\left(W_a^* f_{\theta_a}(X_a), Y_a\right) - T\left(W_b^* f_{\theta_b}(X_a), Y_a\right)

    A drop ΔTa→b>0\Delta T_{a \to b} > 0 indicates true loss of linearly separable features for task aa, whereas ΔTa→b≤0\Delta T_{a \to b} \le 0 indicates feature preservation or positive backward transfer.

  2. Knowl 2 — Representation Forgetting Resistance and Positive Backward Transfer in Contrastive Loss Functions

    empirical result

    When neural networks are continually trained across long task sequences without explicit mechanisms to mitigate forgetting (e.g., naive sequential finetuning), contrastive loss functions exhibit markedly lower representation forgetting than standard cross-entropy (CE) minimization, and can even exhibit positive backward transfer.

    Across multiple task-incremental benchmarks:

    1. SplitCIFAR100 (10 tasks, 10 classes each): Naive sequential training with Supervised Contrastive loss (SupCon) initially shows an accuracy drop after task 1, but its linear probe (LP) accuracy on task 1 stabilizes and increases over later tasks, outperforming regularization methods like Learning without Forgetting (LwF) and rivaling rehearsal methods (Experience Replay with 5 samples/class) without storing old training data.
    2. SplitMiniImageNet (20 tasks) & 200-Task SplitImageNet32 (5 classes per task): As the task sequence lengthens, the LP accuracy of Experience Replay (ER, buffer size M=20M=20) degrades over time. In contrast, SupCon finetuning remains flat or improves on task 1 representations throughout all 200 tasks, ultimately outperforming ER (M=5M=5).
    3. Unsupervised SimCLR: In purely self-supervised sequential learning with SimCLR, LP accuracy drops after the immediate first task boundary but remains flat and resilient across all subsequent tasks, demonstrating that contrastive instance discrimination preserves historical feature spaces.
  3. Knowl 3 — Discrepancy Between Layer-Wise Linear Probes and Centered Kernel Alignment (CKA)

    data/table

    On a 2-task SplitCIFAR10 benchmark, evaluating representations with layer-wise optimal linear probes (LP) reveals that feature forgetting is concentrated exclusively at the highest network layers, while early and intermediate representations remain intact or improve. This directly contrasts with Centered Kernel Alignment (CKA) metrics, which show dramatic drops across network depth and fail to differentiate benign coordinate rotations or positive backward transfer from catastrophic information loss.

    The table below shows the linear probe classification accuracy on Task 1 data evaluated on representations from each residual block (B-0 through B-5 for ResNet, B-0 through B-3 for VGG) immediately after training on Task 1 (extPostT1 ext{Post } T_1) versus after sequentially training on Task 2 (extPostT2 ext{Post } T_2), alongside linear CKA similarity between the two states:

    ResNet Layer LP Acc. Post T1T_1 LP Acc. Post T2T_2 Δ\Delta Acc. CKA
    B-0 63.54% 64.62% +1.08% 0.97
    B-1 68.24% 69.50% +1.26% 0.93
    B-2 71.62% 71.34% -0.28% 0.88
    B-3 77.64% 76.52% -1.12% 0.78
    B-4 80.06% 78.98% -1.08% 0.31
    B-5 85.82% 80.10% -5.72% 0.22
    VGG Layer LP Acc. Post T1T_1 LP Acc. Post T2T_2 Δ\Delta Acc. CKA
    B-0 67.94% 66.86% -1.08% 0.95
    B-1 73.60% 72.52% -1.08% 0.93
    B-2 78.58% 75.68% -2.90% 0.85
    B-3 81.54% 75.48% -6.06% 0.66

    Although the network's observed Task 1 test accuracy drops from 85%85\% to 63.64%63.64\% (ResNet) and 57.88%57.88\% (VGG), the optimal probe accuracy drops by only 5.72%5.72\% and 6.06%6.06\% at the final block, while blocks B-0 and B-1 on ResNet experience positive accuracy gains (+1.08%+1.08\% and +1.26%+1.26\%).

  4. Knowl 4 — Effect of Model Width and Depth on Representation Forgetting in Offline Continual Learning

    data/table

    In offline 10-task SplitCIFAR100, evaluating model capacity scaling using observed classification accuracy suggests that larger models do not alleviate catastrophic forgetting in naive finetuning or regularization methods. However, evaluating the representations with Linear Probing (LP) shows that increasing network width significantly reduces representation forgetting.

    The following table reports the Task 1 initial accuracy, Task 1 observed accuracy after task 10 (Obs. Acc. Task 1 at T=10T=10), Task 1 LP accuracy at T=10T=10, average LP accuracy across all 10 tasks at T=10T=10 (LP Acc. All T=10T=10), and average observed accuracy across all tasks (Avg. Obs. Acc.) using ResNet-18 (RN18) with base width 32 vs. expanded width 128, and ResNet-101 (RN101, depth 101, width 32):

    Method Architecture Task 1 Acc. Task 1 Obs. (T=10T=10) Task 1 LP (T=10T=10) LP All (T=10T=10) Avg. Obs.
    Finetuning (CE) RN18, Width=32 82.2% 20.8% 64.8% 70.8% 35.5%
    RN18, Width=128 83.3% 21.2% 70.5% 74.2% 36.8%
    RN101, Width=32 82.9% 19.8% 67.9% 72.4% 35.9%
    ER (M=5M=5) RN18, Width=32 82.7% 52.1% 74.2% 75.7% 54.8%
    RN18, Width=128 83.6% 54.8% 75.6% 77.3% 55.4%
    RN101, Width=32 83.0% 50.9% 74.5% 76.1% 51.6%
    ER (M=20M=20) RN18, Width=32 82.4% 61.3% 76.0% 76.4% 65.2%
    RN18, Width=128 83.2% 63.5% 78.8% 80.1% 67.0%
    RN101, Width=32 82.9% 60.7% 77.1% 77.5% 63.9%
    LwF RN18, Width=32 82.1% 36.2% 70.1% 73.4% 47.7%
    RN18, Width=128 83.9% 37.7% 74.8% 76.7% 49.1%
    RN101, Width=32 82.5% 35.5% 71.0% 74.6% 46.3%

    While observed accuracy on Task 1 for naive finetuning remains around 20%–21%20\%\text{--}21\% across all capacities, LP accuracy improves by +5.7%+5.7\% (from 64.8%64.8\% to 70.5%70.5\%) with increased width. Furthermore, a wide model with non-rehearsal LwF achieves 76.7%76.7\% average LP accuracy, matching or exceeding rehearsal-based ER (M=5M=5, 75.7%75.7\%).

  5. Knowl 5 — Effect of Model Capacity on Online Task-Incremental Representation Learning

    data/table

    In the online task-incremental setting (single-pass streaming data, learning rate 0.010.01, no momentum), observed classification accuracy heavily underestimates the benefits of model width and overestimates the utility of regularization methods like LwF.

    The table below details final metrics on 10-task SplitCIFAR100 in the online learning regime:

    Method Architecture Task 1 Acc. Task 1 Obs. (T=10T=10) Task 1 LP (T=10T=10) LP All (T=10T=10) Avg. Obs.
    Finetuning (CE) RN18, Width=32 18.6% 12.2% 39.8% 36.4% 22.3%
    RN18, Width=128 19.4% 12.7% 42.3% 41.7% 19.8%
    RN101, Width=32 14.6% 11.8% 28.2% 29.4% 14.5%
    ER (M=5M=5) RN18, Width=32 18.8% 27.3% 36.0% 40.1% 33.8%
    RN18, Width=128 19.5% 28.9% 54.7% 47.9% 31.6%
    RN101, Width=32 15.0% 24.7% 37.1% 30.4% 24.3%
    ER (M=20M=20) RN18, Width=32 18.4% 32.0% 46.8% 43.5% 34.7%
    RN18, Width=128 20.0% 31.8% 51.2% 50.7% 32.5%
    RN101, Width=32 14.5% 25.4% 36.5% 33.9% 24.3%
    LwF RN18, Width=32 18.5% 13.4% 29.5% 36.0% 22.7%
    RN18, Width=128 19.7% 18.3% 34.6% 39.1% 22.1%
    RN101, Width=32 14.8% 11.1% 25.4% 22.8% 16.8%

    Increasing network width from 32 to 128 substantially boosts feature extraction quality across all methods (e.g., ER M=5M=5 Task 1 LP jumps from 36.0%36.0\% to 54.7%54.7\%), despite average observed accuracy dropping from 33.8%33.8\% to 31.6%31.6\%. Conversely, increasing depth (RN101) degrades representation quality in online streams (28.2%28.2\% vs 39.8%39.8\% for Finetuning). LwF performs poorly in online streams (16.8%–22.7%16.8\%\text{--}22.7\% observed accuracy), showing that naive finetuning yields superior representation learning compared to distillation-based regularization under streaming constraints.

  6. Knowl 6 — Fast Remembering via Supervised Contrastive Learning with Nearest Mean of Exemplars

    model/method

    Because continual representation learning with Supervised Contrastive loss (SupCon) maintains high-quality linear feature separability on past tasks without memory buffers during training, rapid recovery ("fast remembering") can be achieved using class prototypes computed at inference or task re-encounter.

    Rather than executing continuous gradient-based rehearsal during training, the framework operates as follows:

    1. Sequential Feature Learning: The feature extractor fθf_\theta is updated on task tt using the standard SupCon loss over the current task minibatch XX: LSupCon=∑xi∈X−1∣P(i)∣∑xp∈P(i)log⁡exp⁡(fθ(xp)Tfθ(xi)τ∥fθ(xp)∥∥fθ(xi)∥)∑xa∈X∖{xi}exp⁡(fθ(xa)Tfθ(xi)τ∥fθ(xa)∥∥fθ(xi)∥)\mathcal{L}_{\text{SupCon}} = \sum_{x_i \in X} \frac{-1}{|P(i)|} \sum_{x_p \in P(i)} \log \frac{\exp\left(\frac{f_\theta(x_p)^T f_\theta(x_i)}{\tau \|f_\theta(x_p)\| \|f_\theta(x_i)\|}\right)}{\sum_{x_a \in X \setminus \{x_i\}} \exp\left(\frac{f_\theta(x_a)^T f_\theta(x_i)}{\tau \|f_\theta(x_a)\| \|f_\theta(x_i)\|}\right)} where P(i)P(i) contains indices of samples sharing class label with xix_i, and τ\tau is the temperature parameter.
    2. Prototype Construction (NME): At inference, or upon re-encountering an old task, a small set Ec\mathcal{E}_c of MM exemplars per class cc (e.g., M=5M=5) is passed through the frozen feature extractor fθf_\theta to compute class mean prototypes: μc=1∣Ec∣∑x∈Ecfθ(x)∥fθ(x)∥\mu_c = \frac{1}{|\mathcal{E}_c|} \sum_{x \in \mathcal{E}_c} \frac{f_\theta(x)}{\|f_\theta(x)\|}
    3. Classification: A query sample x∗x^* is classified using Nearest Mean of Exemplars: y^=arg⁡max⁡c(fθ(x∗)Tμc∥fθ(x∗)∥)\hat{y} = \arg\max_c \left( \frac{f_\theta(x^*)^T \mu_c}{\|f_\theta(x^*)\|} \right)

    This method incurs no replay computational overhead during training, requiring approximately half the training compute of Experience Replay.

  7. Knowl 7 — Empirical Performance of SupCon + NME Fast Remembering

    data/table

    On the 10-task SplitCIFAR100 benchmark, combining non-rehearsal SupCon finetuning with a post-hoc Nearest Mean of Exemplars (NME) classifier using M=5M=5 randomly selected exemplars per class outperforms regularization baselines and rivals Experience Replay (ER) without storing or iterating over rehearsal buffers during training.

    Method Obs. Acc. Task 1 at T=10T=10 Avg. Obs. Acc.
    Finetuning (CE) 20.8% 35.5%
    ER (M=5M=5) 52.1% 54.8%
    ER (M=20M=20) 61.3% 65.2%
    LwF 36.2% 47.7%
    Finetune (SupCon) + NME-M5M5 48.0% 53.9%

    Finetune (SupCon) + NME-M5M5 achieves 53.9%53.9\% average accuracy across all 10 tasks, surpassing LwF (47.7%47.7\%) by +6.2%+6.2\% and approaching ER-M5M5 (54.8%54.8\%), while eliminating the gradient compute and memory access overhead required by ER during sequential model training.

  8. Knowl 8 — Trade-Offs on Large-Scale ImageNet Sequential Transfer

    data/table

    In a 4-task sequential transfer setting comprising ImageNet →\to MIT Scenes →\to CUB Birds →\to Oxford Flowers using ResNet-18, naive SupCon training maintains high linear probe accuracy on initial ImageNet representations while simultaneously achieving competitive or superior performance on downstream target tasks compared to cross-entropy finetuning and regularization methods.

    The table below shows the observed classification accuracy on each subsequent task in the sequence:

    Method Acc. Scenes Acc. CUB Acc. Flowers
    Finetuning (CE) 56.9% ±\pm 1.1 54.5% ±\pm 2.6 89.3% ±\pm 1.1
    LwF 57.6% ±\pm 1.5 43.1% ±\pm 2.9 85.3% ±\pm 0.5
    EWC (λ=0.5k\lambda = 0.5k) 52.5% ±\pm 1.1 47.8% ±\pm 2.5 85.9% ±\pm 1.6
    EWC (λ=8k\lambda = 8k) 42.1% ±\pm 1.5 38.3% ±\pm 0.9 79.1% ±\pm 1.0
    Finetuning (SupCon) 57.1% ±\pm 1.2 50.4% ±\pm 1.0 85.3% ±\pm 0.9

    Although heavily regularized EWC (λ=8k\lambda = 8k) retains the highest LP and observed accuracy on old ImageNet data, its current-task performance is heavily impaired (e.g., 42.1%42.1\% on Scenes vs 57.1%57.1\% for SupCon). Naive SupCon finetuning provides strong plasticity on incoming tasks while outperforming LwF on old-task LP accuracy without explicit forgetting regularization.

  9. Knowl 9 — Limitations of Linear Probe Representation Analysis in Continual Learning

    limitation

    The analysis and evaluation methodology presented have three primary limitations:

    1. Linear Separability as a Single Knowledge Proxy: Evaluating representation preservation via an optimal linear probe measures the availability of linearly readable features. However, task accuracy under an optimal linear readout may not encompass all facets of knowledge retention (e.g., nonlinear interactions, calibration, or generative properties).
    2. Task-Incremental Scope: The experiments and linear probe formulations are conducted in the task-incremental setting (where task identities or separate task evaluation contexts are known). The findings are not evaluated in the class-incremental setting where cross-task class competition occurs.
    3. Domain Similarity: While evaluated across varied natural image datasets (ImageNet, MIT Scenes, CUB Birds, Oxford Flowers, CIFAR-100), the task transitions remain within natural photographic domains and do not cover extreme distribution shifts (e.g., transfer to sketches or highly stylized domains).

Coverage note — Deliberately omitted qualitative hyperparameter tables and routine architecture descriptions from the appendix to focus on core methodological formulations, empirical linear probe findings, scaling dynamics, and fast-remembering algorithms.

References

  1. 1.Rahaf Aljundi, Francesca Babiloni, Mohamed Elhoseiny, Marcus Rohrbach, and Tinne Tuytelaars. Memory aware synapses: Learning what (not) to forget. In ECCV 2018, 2018. 2
  2. 2.Rahaf Aljundi, Punarjay Chakravarty, and Tinne Tuytelaars. Expert gate: Lifelong learning with a network of experts. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016. 2
  3. 3.Gaurav Arora, Afshin Rahimi, and Timothy Baldwin. Does an LSTM forget more than a CNN? an empirical study of catastrophic forgetting in NLP. In Proceedings of the The 17th Annual Workshop of the Australasian Language Technology Association, pages 77–86, Sydney, Australia, 4–6 Dec. 2019. Australasian Language Technology Association. 2
  4. 4.Yoshua Bengio, Aaron Courville, and Pascal Vincent. Representation learning: A review and new perspectives. IEEE transactions on pattern analysis and machine intelligence, 35(8):1798–1828, 2013. 1
  5. 5.Lucas Caccia, Rahaf Aljundi, Tinne Tuytelaars, Joelle Pineau, and Eugene Belilovsky. Reducing representation drift in online continual learning. arXiv preprint arXiv:2104.05025, 2021. 1, 3
  6. 6.Rich Caruana, Steve Lawrence, and Lee Giles. Overfitting in neural nets: Backpropagation, conjugate gradient, and early stopping. Advances in neural information processing systems, pages 402–408, 2001. 11
  7. 7.Arslan Chaudhry, Marcus Rohrbach, Mohamed Elhoseiny, Thalaiyasingam Ajanthan, Puneet K Dokania, Philip HS Torr, and Marc’Aurelio Ranzato. Continual learning with tiny episodic memories. arXiv preprint arXiv:1902.10486, 2019. 2, 3, 5, 7
  8. 8.Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. A simple framework for contrastive learning of visual representations. arXiv preprint arXiv:2002.05709, 2020. 2, 3
  9. 9.Patryk Chrabaszcz, Ilya Loshchilov, and Frank Hutter. A downsampled variant of imagenet as an alternative to the cifar datasets. arXiv preprint arXiv:1707.08819, 2017. 4
  10. 10.Ian Goodfellow, Yoshua Bengio, and Aaron Courville. Deep learning. Book in preparation for MIT Press, 2016. 1
  11. 11.Ian J Goodfellow, Mehdi Mirza, Da Xiao, Aaron Courville, and Yoshua Bengio. An empirical investigation of catastrophic forgetting in gradient-based neural networks. arXiv preprint arXiv:1312.6211, 2013. 1
  12. 12.Florian Graf, Christoph Hofer, Marc Niethammer, and Roland Kwitt. Dissecting supervised constrastive learning. In International Conference on Machine Learning, pages 3821–3830. PMLR, 2021. 3
  13. 13.David Ha and Douglas Eck. A neural representation of sketch drawings. arXiv preprint arXiv:1704.03477, 2017. 8
  14. 14.Raia Hadsell, Dushyant Rao, Andrei A Rusu, and Razvan Pascanu. Embracing change: Continual learning in deep neural networks. Trends in cognitive sciences, 2020. 1
  15. 15.Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick. Momentum contrast for unsupervised visual representation learning. arXiv preprint arXiv:1911.05722, 2019. 2, 3
  16. 16.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016. 11
  17. 17.Xu He, Jakub Sygnowski, Alexandre Galashov, Andrei A Rusu, Yee Whye Teh, and Razvan Pascanu. Task agnostic continual learning via meta learning. arXiv preprint arXiv:1906.05201, 2019. 1
  18. 18.Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. Distilling the knowledge in a neural network. arXiv preprint arXiv:1503.02531, 2015. 2
  19. 19.Dapeng Hu, Qizhengqiu Lu, Lanqing Hong, Hailin Hu, Yifan Zhang, Zhenguo Li, Alfred Shen, and Jiashi Feng. How well self-supervised pre-training performs with streaming data? arXiv preprint arXiv:2104.12081, 2021. 2
  20. 20.Ferenc Huszar. On quadratic penalties in elastic weight con- ´ solidation. arXiv preprint arXiv:1712.03847, 2017. 2, 5
  21. 21.Prannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna, Yonglong Tian, Phillip Isola, Aaron Maschinot, Ce Liu, and Dilip Krishnan. Supervised contrastive learning. In H. Larochelle, M. Ranzato, R. Hadsell, M. F. Balcan, and H. Lin, editors, Advances in Neural Information Processing Systems, volume 33, pages 18661–18673. Curran Associates, Inc., 2020. 2, 3, 4, 5
  22. 22.James Kirkpatrick, Razvan Pascanu, Neil Rabinowitz, Joel Veness, Guillaume Desjardins, Andrei A Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka GrabskaBarwinska, et al. Overcoming catastrophic forgetting in neural networks. arXiv preprint arXiv:1612.00796, 2016. 2
  23. 23.Simon Kornblith, Mohammad Norouzi, Honglak Lee, and Geoffrey Hinton. Similarity of neural network representations revisited. In International Conference on Machine Learning, pages 3519–3529. PMLR, 2019. 2, 3, 8
  24. 24.Alex Krizhevsky, Geoffrey Hinton, et al. Learning multiple layers of features from tiny images. Technical report, University of Toronto, 2009. 4
  25. 25.Xilai Li, Yingbo Zhou, Tianfu Wu, Richard Socher, and Caiming Xiong. Learn to grow: A continual structure learning framework for overcoming catastrophic forgetting. In International Conference on Machine Learning, pages 3925– 3934. PMLR, 2019. 2
  26. 26.Zhizhong Li and Derek Hoiem. Learning without forgetting. IEEE Transactions on Pattern Analysis and Machine Intelligence, 40(12):2935–2947, 2017. 2, 4, 5, 6, 11, 12, 13
  27. 27.David Lopez-Paz et al. Gradient episodic memory for continual learning. In Advances in Neural Information Processing Systems, pages 6467–6476, 2017. 2, 5, 7
  28. 28.Ilya Loshchilov and Frank Hutter. Sgdr: Stochastic gradient descent with warm restarts. arXiv preprint arXiv:1608.03983, 2016. 11
  29. 29.Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101, 2017. 11
  30. 30.Zheda Mai, Ruiwen Li, Hyunwoo Kim, and Scott Sanner. Supervised contrastive replay: Revisiting the nearest class mean classifier in online class-incremental continual learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3589–3599, 2021. 2, 3
  31. 31.Michael McCloskey and Neal J Cohen. Catastrophic interference in connectionist networks: The sequential learning problem. Psychology of learning and motivation, 24:109– 165, 1989. 1
  32. 32.In Jae Myung. Tutorial on maximum likelihood estimation. Journal of mathematical Psychology, 47(1):90–100, 2003. 2
  33. 33.Cuong V Nguyen, Alessandro Achille, Michael Lam, Tal Hassner, Vijay Mahadevan, and Stefano Soatto. Toward understanding catastrophic forgetting in continual learning. arXiv preprint arXiv:1908.01091, 2019. 2
  34. 34.Cuong V Nguyen, Yingzhen Li, Thang D Bui, and Richard E Turner. Variational continual learning. arXiv preprint arXiv:1710.10628, 2017. 2
  35. 35.M-E. Nilsback and A. Zisserman. Automated flower classification over a large number of classes. In Proceedings of the Indian Conference on Computer Vision, Graphics and Image Processing, Dec 2008. 5
  36. 36.Edouard Oyallon. Building a regular decision boundary with deep networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 5106– 5114, 2017. 2
  37. 37.Ariadna Quattoni and Antonio Torralba. Recognizing indoor scenes. In Computer Vision and Pattern Recognition, 2009. CVPR 2009. IEEE Conference on, pages 413–420. IEEE, 2009. 4, 5, 11
  38. 38.Vinay V Ramasesh, Ethan Dyer, and Maithra Raghu. Anatomy of catastrophic forgetting: Hidden representations and task semantics. arXiv preprint arXiv:2007.07400, 2020. 2, 3, 4, 5, 8, 11
  39. 39.Vinay Venkatesh Ramasesh, Aitor Lewkowycz, and Ethan Dyer. Effect of scale on catastrophic forgetting in neural networks. In International Conference on Learning Representations, 2022. 2, 6
  40. 40.Sylvestre-Alvise Rebuffi, Alexander Kolesnikov, Georg Sperl, and Christoph H Lampert. icarl: Incremental classifier and representation learning. In Proc. CVPR, 2017. 2, 3, 5
  41. 41.Amir Rosenfeld and John K Tsotsos. Incremental learning through deep adaptation. IEEE transactions on pattern analysis and machine intelligence, 42(3):651–663, 2018. 2
  42. 42.Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei. ImageNet Large Scale Visual Recognition Challenge. International Journal of Computer Vision (IJCV), 115(3):211–252, 2015. 1, 4, 5, 11
  43. 43.Andrei A Rusu, Neil C Rabinowitz, Guillaume Desjardins, Hubert Soyer, James Kirkpatrick, Koray Kavukcuoglu, Razvan Pascanu, and Raia Hadsell. Progressive neural networks. arXiv preprint arXiv:1606.04671, 2016. 2
  44. 44.Patsorn Sangkloy, Nathan Burnell, Cusuh Ham, and James Hays. The sketchy database: learning to retrieve badly drawn bunnies. ACM Transactions on Graphics (TOG), 35(4):1–12, 2016. 8
  45. 45.Karen Simonyan and Andrew Zisserman. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556, 2014. 11
  46. 46.Catherine Wah, Steve Branson, Peter Welinder, Pietro Perona, and Serge Belongie. The Caltech-UCSD Birds-200- 2011 Dataset. Technical Report CNS-TR-2011-001, California Institute of Technology, 2011. 5, 11
  47. 47.P. Welinder, S. Branson, T. Mita, C. Wah, F. Schroff, S. Belongie, and P. Perona. Caltech-UCSD Birds 200. Technical Report CNS-TR-2010-001, California Institute of Technology, 2010. 4
  48. 48.Matthew D Zeiler and Rob Fergus. Visualizing and understanding convolutional networks. In European conference on computer vision, pages 818–833. Springer, 2014. 2, 3
  49. 49.Friedemann Zenke, Ben Poole, and Surya Ganguli. Improved multitask learning through synaptic intelligence. In Proceedings of the International Conference on Machine Learning (ICML), 2017. 2

Citation

MLA
Davari, M., et al. “Probing Representation Forgetting in Supervised and Unsupervised Continual Learning”. arXiv, 2022, http://arxiv.org/abs/2203.13381v2.
APA
Davari, M., Asadi, N., Mudur, S., Aljundi, R., & Belilovsky, E. (2022). Probing Representation Forgetting in Supervised and Unsupervised Continual Learning. arXiv. http://arxiv.org/abs/2203.13381v2
Chicago
Davari, M., N. Asadi, S. Mudur, R. Aljundi, and E. Belilovsky. 2022. “Probing Representation Forgetting in Supervised and Unsupervised Continual Learning”. arXiv. http://arxiv.org/abs/2203.13381v2.
Harvard
Davari, M. et al. (2022) “Probing Representation Forgetting in Supervised and Unsupervised Continual Learning”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2203.13381v2.
Vancouver
1. Davari M, Asadi N, Mudur S, Aljundi R, Belilovsky E (2022) Probing Representation Forgetting in Supervised and Unsupervised Continual Learning. arXiv

BibTeX

@article{davari2022probing,
  title = {Probing Representation Forgetting in Supervised and Unsupervised Continual Learning},
  author = {Davari, MohammadReza and Asadi, Nader and Mudur, Sudhir and Aljundi, Rahaf and Belilovsky, Eugene},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2203.13381v2},
  eprint = {2203.13381}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE